Swap the order, change the winner: how to run an LLM judge
Reordering two answers handed the win to the weaker model on 66 of 80 questions. Ten studies on LLM judges, their biases, and a validation protocol.
Read the article →Research-backed, actionable articles on AI at work, endurance training and how people learn. Every claim has a source.
Reordering two answers handed the win to the weaker model on 66 of 80 questions. Ten studies on LLM judges, their biases, and a validation protocol.
Read the article →Eight studies on the security of AI-written code: the failure rates by language, the confidence gap it creates, and the review routine that catches the weaknesses.
Read the article →Students predicted 33.9 days for a thesis and took 55.5. Ten sources on the planning fallacy, where it reverses, and a forecasting routine built on finished projects.
Read the article →Human and AI teams scored below the better of the two alone across 106 experiments. Ten studies on why we accept wrong answers, and a protocol that puts yours first.
Read the article →Employees who wrote a date and a time got vaccinated 4.2 points more often. Nine studies on if-then plans, where the effect shrinks, and a format to copy.
Read the article →Ten studies on prompting: personas, tips and threats change nothing, while formatting, example order and the position of key facts move accuracy by double digits.
Read the article →An interruption of 2.8 seconds doubles error rates. Six studies on attention residue, batching and stress, plus a protocol to protect your focus this week.
Read the article →Five controlled studies on AI and productivity, from GitHub Copilot to BCG consultants, and a method to know when AI helps you and when it makes you worse.
Read the article →