11 min read
An LLM Doesn't Know Whether It Finished — Designing a Harness That Moves the Completion Verdict Outside the Model
Multiple studies conclude that an LLM agent's completion bias and overconfidence cannot be corrected by prompting. This post lays out the design principles of a vulnerability-checking automation harness that makes "surveyed everything" and "0 vulnerabilities" computed from an HMAC receipt ledger and deterministic hooks rather than from the model's narrative. A "200 OK" without a negative control is not confirmation.
Read full article