Swap the order, change the winner: how to run an LLM judge
Reordering two answers handed the win to the weaker model on 66 of 80 questions. Ten studies on LLM judges, their biases, and a validation protocol.
Read the article →Research-backed, actionable articles on AI at work, endurance training and how people learn. Every claim has a source.
Reordering two answers handed the win to the weaker model on 66 of 80 questions. Ten studies on LLM judges, their biases, and a validation protocol.
Read the article →Eight studies on the security of AI-written code: the failure rates by language, the confidence gap it creates, and the review routine that catches the weaknesses.
Read the article →Human and AI teams scored below the better of the two alone across 106 experiments. Ten studies on why we accept wrong answers, and a protocol that puts yours first.
Read the article →Ten studies on prompting: personas, tips and threats change nothing, while formatting, example order and the position of key facts move accuracy by double digits.
Read the article →Five controlled studies on AI and productivity, from GitHub Copilot to BCG consultants, and a method to know when AI helps you and when it makes you worse.
Read the article →