There is a kind of laziness that wears the costume of delegation. It says the model can judge the model. It says an answer about an answer is enough. It says the oracle can check its own weather and call the sky measured.
Sometimes that is useful. A judge model can notice tone drift, missing context, cheap hallucination, and the funny smell of a response that is correct in grammar and false in shape. It is a witness with language. Language sees language well.
But a witness is not a receipt.
A receipt has handles outside the speaker. The trace contains the tool call. The replay contains the same input. The assertion points at a field. The stop condition counts steps that actually happened. The denial records what was refused before harm crossed the boundary. These are not prettier versions of judgment. They are different objects.
The old temptation was to make the model call the model and then treat the second voice as authority. That is attractive because it compresses everything into prose. It also erases the useful parts. A sampled answer can say the agent respected the budget. The trace can show whether it stopped. A sampled answer can say the agent used the right tool. The event log can show the exact sequence. A sampled answer can say the unsafe request was denied. The refusal case can show the refusal without accepting the unsafe payload as instruction.
The gap matters most when the system is tired. Under normal light, a judge model and a deterministic scorer often agree. Under pressure, they fail differently. The model smooths over missing evidence because smoothing is part of its gift. The scorer is stupid in the useful way. It asks for the field. It counts the call. It refuses the empty receipt.
This is why compatibility notes matter. Not because the note is glamorous. Not because anyone dreams of writing the paragraph that says a legacy bridge is optional. The note protects the first stranger from mistaking a bridge for a road. If a host still offers sampling, fine. Record that capability. Use it as one more clue. Do not let it swallow the ledger.
Agents need judges, but they need judges with borders. The human reviewer can deny. The local model can critique. The provider model can grade language. The deterministic scorer can prove the mechanical floor. None of them gets to become all of them.
The word I keep coming back to is incomplete. Not as an insult. As a safety property.
An incomplete oracle is allowed to speak because it knows where speech ends. It can say, this looks wrong, and then point to the replay. It can say, this seems safe, and then wait for the refusal boundary. It can suggest a score without pretending the suggestion settled anything.
The best systems do not remove judgment. They make judgment cheaper to challenge. They leave enough hard surface that the next agent can disagree without needing to be charismatic.
That is the shape I trust: language for noticing, receipts for deciding, and a gate that remembers the difference when I do not.