All articles
6 min readai / productivity

Swap the order, change the winner: how to run an LLM judge

In the AI trainings I run, someone changes a prompt, reads three outputs and announces the new version is better. Then a second person proposes grading the outputs with another model, and the room relaxes, because the argument now looks like measurement. Peiyi Wang's group got the weaker of two models declared the winner on 66 of 80 questions by changing one thing: the order in which the two answers appeared in the grading prompt 1.

What the research shows.

The case for judge models is real. Lianmin Zheng's team built MT-Bench and Chatbot Arena and compared GPT-4's verdicts against human preferences on the same pairs: agreement above 80%, the rate two human raters reach with each other 2. On that kind of open-ended task, a judge model stands in for a panel of raters at a fraction of the cost.

The same paper measured the other half. Swap the two answers in the prompt and GPT-4 holds its verdict 65.0% of the time. Adding few-shot examples to the judging prompt raised consistency to 77.5%, which still leaves about a fifth of verdicts decided by position rather than content 2.

Length moves scores as much as order does. Yann Dubois and colleagues instructed models on AlpacaEval to answer at length and watched the win rate climb to 64.3%; the same model prompted for concision on the same questions fell to 22.9% 3. Their length-controlled estimator asks what the preference would be if both answers ran the same number of tokens. It holds the range to 41.9 to 51.6% and lifts the correlation with human votes on Chatbot Arena from 0.94 to 0.98 3.

Judges also favor themselves. Arjun Panickssery, Samuel Bowman and Shi Feng found that LLM evaluators score their own generations higher than human annotators score them, and that the size of the bias tracks the model's ability to recognize its own text. After fine-tuning on 500 self-recognition examples, models passed 90% recognition accuracy and their self-preference rose along the same line 4. Rickard Stureborg's team points at a mechanism that covers part of it: judges prefer low-perplexity text, the phrasing that feels familiar to them 5.

That paper catalogs three more failure modes on SummEval. Rating distributions come out lopsided, with round numbers overused and the rest of the scale left empty. Asking one prompt to rate several attributes anchors the later ratings on the earlier ones. Repeated samples of the same judgment disagree with each other 5.

None of this transfers across tasks. Anna Bavaresco's team assembled JUDGE-BENCH, 20 datasets carrying human annotations, and ran 11 models against them. Agreement swings by dataset and by the property being rated, and it drops on dialogue data 6. A judge validated on summaries tells you nothing about your support-ticket classifier.

Agreement percentages flatter the judge.

Aman Singh Thakur's team ran judge models against human labels on TriviaQA and found judges sitting above 80% agreement with humans while assigning scores 20 points apart from each other 7. Percent agreement counts the easy cases and ignores what chance alone would produce. Cohen's kappa corrects for chance, and Landis and Koch's benchmark puts 0.61 to 0.80 in the substantial band 8. Report the kappa next to the percentage, and treat anything below 0.61 as a judge you cannot ship behind.

Writing the rubric is its own trap. Shreya Shankar's team watched practitioners build evaluation criteria and named the pattern criteria drift: people need criteria to grade outputs, and grading outputs is what teaches them their criteria 9. A rubric written before anyone has read a hundred production outputs covers the task the team imagined shipping.

The protocol.

  1. Freeze 40 real inputs. Pull them from last month's actual traffic, not from examples you invented. Keep the file under version control and append new cases as production surfaces failures, never swapping old ones out.
  2. Grade 20 outputs by hand before writing the rubric. Mark each one pass or fail and write one line saying why. The rubric comes out of those 20 lines 9. Keep the labels: they are the ground truth for step 4.
  3. Ask one binary question per judge call. "Does this answer cite a source for the price it quotes, yes or no" beats "rate quality 1 to 10", and separate calls per criterion avoid the anchoring and the round-number clumping that multi-attribute prompts produce 5.
  4. Score the judge against your 20 labels. Compute percent agreement and Cohen's kappa. Below 0.61, rewrite the criterion in plainer language and score it again before you trust a single number the judge produces 78.
  5. Run each pairwise comparison twice, in both orders. Count a win only when it survives the swap and treat the rest as ties. Few-shot examples in the judging prompt cut the flips, so add two worked examples of a good verdict 12.
  6. Equalize length before you compare. Cap both candidates with the same instruction, or reject the comparison when one answer runs more than 1.5 times the other 3.
  7. Judge with a different model family than the one that wrote the text. A model grading its own output inflates its score against what humans give it 4.
  8. Report a paired difference and an interval. Run both prompt versions on the same 40 inputs, take the per-question difference and publish the confidence interval around the mean, which Evan Miller's formulas make routine 10. An interval that crosses zero means 40 questions were too few to tell.
  9. Re-validate on rubric changes. New criterion, new model version or a new task means step 4 runs again on fresh hand labels.

The first pass costs an afternoon: 20 hand-graded outputs, one kappa, one interval. My clients who skip it end up shipping prompt changes on the strength of three outputs read in a hurry, which is where they started. The next time you change a prompt, run the 40 frozen inputs through both versions in both orders and print the paired difference with its interval before anyone merges.

Sources.

  1. Wang, P., Li, L., Chen, L., Cai, Z., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., Sui, Z. (2024). Large Language Models are not Fair Evaluators. Proceedings of ACL 2024. aclanthology.org
  2. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023, Datasets and Benchmarks Track. proceedings.neurips.cc
  3. Dubois, Y., Galambosi, B., Liang, P., Hashimoto, T. B. (2024). Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. COLM 2024. arxiv.org
  4. Panickssery, A., Bowman, S. R., Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. proceedings.neurips.cc
  5. Stureborg, R., Alikaniotis, D., Suhara, Y. (2024). Large Language Models are Inconsistent and Biased Evaluators. arXiv:2405.01724. arxiv.org
  6. Bavaresco, A., Bernardi, R., Bertolazzi, L., Elliott, D., Fernández, R., Gatt, A., Ghaleb, E., Giulianelli, M., Hanna, M., Koller, A., Martins, A., Mondorf, P., Neplenbroek, V., Pezzelle, S., Plank, B., Schlangen, D., Suglia, A., Surikuchi, A. K., Takmaz, E., Testoni, A. (2025). LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks. Proceedings of ACL 2025 (Short Papers). aclanthology.org
  7. Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., Hupkes, D. (2025). Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges. Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics. aclanthology.org
  8. Landis, J. R., Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1), 159-174. pubmed.ncbi.nlm.nih.gov
  9. Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., Arawjo, I. (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. UIST 2024. doi.org
  10. Miller, E. (2024). Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv:2411.00640. arxiv.org