Scoring open-ended model output by hand does not scale, so most teams now use a language model as the judge. It is a practical approach, and it is also easy to misuse. A judge score is a measurement, and a measurement needs to be checked before it is believed.
Zheng and colleagues studied this in "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (NeurIPS 2023, Datasets and Benchmarks track). They found that a strong judge model can reach over 80% agreement with human preferences, which is about the level at which humans agree with each other. They also documented systematic weaknesses. Agreement on average does not mean the judge is unbiased on any given dataset, and that is the point of calibrating it.
When a judge compares two answers, it can favour whichever one appears first (or second), regardless of quality. Control: run every comparison twice with the order swapped, and count a verdict only when both runs agree. Treat inconsistent results as ties, and report how often they occur.
Judges can prefer longer answers even when the extra length adds nothing. This matters most when you are comparing models that differ in how wordy they are. Control: score against explicit criteria rather than an overall impression, penalize padding in the rubric, and check whether your scores correlate with answer length.
The same paper discusses a tendency for a judge to favour answers written by itself or by similar models, with evidence that is suggestive rather than conclusive. The risk is practical: if the same model family both generates and judges, errors can go unnoticed. Control: use a judge from a different model family than the generator, and spot-check with human review.
The same study notes that judges are less reliable when grading math and reasoning answers. Control: where a verified answer exists, give it to the judge as a reference, or check the final answer programmatically instead of by judgment.
LLM judges are a useful tool, not an oracle. Calibrate against humans, test for the known biases, and publish the evidence next to the score. A benchmark result you cannot audit is an opinion.
Synthetic datasets and evaluation benchmarks for teams shipping LLMs and AI agents, delivered with the evidence.