Swap the order, change the winner: how to run an LLM judge
Reordering two answers handed the win to the weaker model on 66 of 80 questions. Ten studies on LLM judges, their biases, and a validation protocol.
Read the article →Research-backed, actionable articles on AI at work, endurance training and how people learn. Every claim has a source.
Reordering two answers handed the win to the weaker model on 66 of 80 questions. Ten studies on LLM judges, their biases, and a validation protocol.
Read the article →Harvard physics students learned more in the classes they rated lower. Six studies on why satisfaction scores miss learning, and a feedback form that does not.
Read the article →Two weeks out, cut volume 41 to 60% and leave the hard sessions alone. Nine studies on taper length, what to cut, and the overload block that backfires.
Read the article →Eight studies on the security of AI-written code: the failure rates by language, the confidence gap it creates, and the review routine that catches the weaknesses.
Read the article →Ten studies on caffeine and endurance: the dose that moves a time trial, why your daily coffee habit does not blunt it, and the 45 minutes of sleep it costs.
Read the article →Guessing before you study beat extra reading in all 5 founding experiments and 97 later ones. Eight papers on pretesting, where it stops working, and a protocol.
Read the article →Students predicted 33.9 days for a thesis and took 55.5. Ten sources on the planning fallacy, where it reverses, and a forecasting routine built on finished projects.
Read the article →Human and AI teams scored below the better of the two alone across 106 experiments. Ten studies on why we accept wrong answers, and a protocol that puts yours first.
Read the article →Employees who wrote a date and a time got vaccinated 4.2 points more often. Nine studies on if-then plans, where the effect shrinks, and a format to copy.
Read the article →Heavy resistance training beat jump work on running economy and time trials across 22 studies. Eight papers on the load, the frequency and the timing.
Read the article →Students who mixed problem types scored 61% against 38% for blocked practice. Eight studies on interleaving, where it fails, and how to rebuild a practice set.
Read the article →Ten studies on prompting: personas, tips and threats change nothing, while formatting, example order and the position of key facts move accuracy by double digits.
Read the article →Marathon runners take in 22 g of carbohydrate an hour against 60 to 90 g of evidence-based advice. Seven studies on fueling rates, gut training and the real ceiling.
Read the article →An interruption of 2.8 seconds doubles error rates. Six studies on attention residue, batching and stress, plus a protocol to protect your focus this week.
Read the article →Elite endurance athletes do 80% of their training at low intensity. Three studies explain why polarized training beats the comfortable middle, with a weekly protocol.
Read the article →Five controlled studies on AI and productivity, from GitHub Copilot to BCG consultants, and a method to know when AI helps you and when it makes you worse.
Read the article →A century of memory research: people forget most of a training within days. Retrieval practice and spacing are the two proven fixes, with numbers and a protocol.
Read the article →