clickai.dev
← Legend / Verification and authority

LLM-as-judge

LLM-as-judge is the use of a model to grade another model's output against a rubric, either as an evaluation or as an inline gate. It is now standard and economically viable, with limitations severe enough that any honest description leads with them. Position, verbosity and self-enhancement bias were catalogued in the paper that named the technique. On code specifically, swapping which of two candidates is shown first has been measured to move judging accuracy by more than ten points. And position bias is strongest exactly where a judge is most needed: it scales with how close the candidates are in quality. Scores also run systematically optimistic. Treat judges as signal and deterministic checks as gates.

establishedEstablished as a technique with well-documented biases — not as a substitute for a deterministic gate.

Legend last revised .