LLM-as-judge
establishedEstablished as a technique with well-documented biases — not as a substitute for a deterministic gate.
Using a model to grade another model's output against a rubric, either as an evaluation or as an inline gate. Now standard and economically viable, with limitations severe enough that any honest description leads with them: position and presentation bias can move accuracy by more than ten points on code evaluation specifically, and it is worst exactly when candidates are close in quality; scores run systematically optimistic. Treat judges as signal and deterministic checks as gates.