Expert personas do nothing: what controlled tests say about prompting
In the AI trainings I run, someone asks for the magic sentence within the first hour. Tell the model it is a senior expert in the field. Offer it a tip. Threaten it. Research teams have now put those sentences on graduate-level benchmarks with enough repetitions to see past the noise, and they lose. The choices that move accuracy are the boring ones: where you place the key facts, how you format the prompt, and whether you ask for reasoning on a task that needs it.
The persona line buys nothing.
In December 2025, Savir Basil and colleagues at Wharton tested expert personas across six models on GPQA Diamond and MMLU-Pro, graduate-level questions in science, engineering and law. Assigning a persona matched to the domain had no measurable effect on accuracy, with Gemini 2.0 Flash as the single exception 1. Personas from the wrong domain sometimes lowered scores, and low-knowledge personas (layperson, young child, toddler) often did.
The same group tested the two other rituals people bring to my sessions: offering the model a tip and threatening it. Across GPQA Diamond and MMLU-Pro, neither moved benchmark averages 2. Individual questions did swing, in both directions, which explains why the folklore survives. Someone tries "I'll tip you $200", sees one answer improve, and tells the story for a year.
That per-question variance is the reason the Wharton team runs each question 25 times per condition 3. Your own before-and-after test on a single prompt cannot separate a real gain from a coin flip.
Chain of thought earns its cost on math.
The habit came from a real result. In 2022, Takeshi Kojima and co-authors appended "Let's think step by step" to prompts and pushed text-davinci-002 from 17.7% to 78.7% on MultiArith and from 10.4% to 40.7% on GSM8K 4. Those numbers deserved the attention they got.
The scope narrowed later. Zayne Sprague and colleagues ran a meta-analysis of more than 100 papers plus their own evaluation of 20 datasets across 14 models, published at ICLR 2025. Chain-of-thought gains concentrate on math and symbolic reasoning. On MMLU, as much as 95% of the improvement came from questions where an equals sign appeared in the question or in the model's answer 5. On everything else, asking the model to reason out loud changed little.
Lennart Meincke, Ethan Mollick, Lilach Mollick and Dan Shapiro measured the bill. For reasoning models, adding chain-of-thought instructions produced marginal accuracy gains while raising response time by 20% to 80%. For non-reasoning models, average performance improved a little and answer variability rose, including fresh errors on questions the model had been getting right 3.
Structure moves accuracy more than wording.
Melanie Sclar and colleagues varied the surface format of few-shot prompts, separators, casing, spacing, choices that leave the meaning untouched. On LLaMA-2-13B, performance across those equivalent formats spread by up to 76 accuracy points 6. Model size, more examples and instruction tuning all failed to remove the sensitivity.
Order carries the same weight. Yao Lu and co-authors showed that permuting few-shot examples moves GPT-family models between near state-of-the-art and random guessing on eleven classification tasks, and that a good permutation for one model does not transfer to another 7.
Position matters inside long context too. Nelson Liu and colleagues found a U-shaped curve: models use information best at the beginning or the end of the input. With 20 documents, GPT-3.5-Turbo answered worse when the relevant document sat in the middle than when it received no documents at all, a closed-book score of 56.1% 8. Paste 40 pages, bury the decisive paragraph at page 20, and you have handed the model a distraction.
Reasoning models want shorter prompts.
The DeepSeek team states it plainly in the R1 paper: the model is sensitive to prompts, few-shot prompting consistently degrades its performance, and users should describe the problem and specify the output format in a zero-shot setting 9. Read alongside the Wharton timing result 3, the guidance for reasoning models inverts a decade of prompt advice. Fewer examples, no step-by-step instruction, more precision about the output you want.
The protocol.
1. Cut the persona line, keep what it was standing in for. Replace "you are a senior financial analyst" with the constraint you meant: audience, depth, vocabulary, length. "Write 200 words for a CFO with 3 minutes, no jargon" gives the model something to act on 1.
2. Put decisive material first or last. When you paste more than a few pages, open with the question, then the documents, then repeat the question. Never leave the one paragraph that matters in the middle of a long dump 8.
3. Ask for step-by-step reasoning on math, unit conversions, date arithmetic, code traces and logic puzzles. On summarizing, drafting, classification and extraction, drop it and take the latency back 53.
4. On a reasoning model, delete the examples and specify the format. Give the schema, the length and the constraints, then let it work 9.
5. Freeze a prompt once it works. Keep the separators, the casing and the example order as they are, and store the template in a file rather than retyping it. A rewrite you consider cosmetic can cost double-digit accuracy 67.
6. Run high-stakes questions three times in separate chats and compare. Xuezhi Wang and colleagues formalized this as self-consistency: sampling several reasoning paths and keeping the majority answer added 17.9% on GSM8K over greedy decoding 10. Three divergent answers tell you the model is guessing.
7. Test prompt changes with 5 runs per version, not one. Count the failures on each side. Below that, per-question noise will sell you a technique that does nothing 23.
Open the prompt you reuse most this week. Delete its first line if it starts with "you are an expert", move the question above the pasted material, then run the old and the new version 5 times each on the same task and count how often the output survives your review.
Sources.
- Basil, S., Shapiro, I., Shapiro, D., Mollick, E. R., Mollick, L., Meincke, L. (2025). Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy. arXiv:2512.05858
- Meincke, L., Mollick, E. R., Mollick, L., Shapiro, D. (2025). Prompting Science Report 3: I'll pay you or I'll kill you, but will you care? arXiv:2508.00614
- Meincke, L., Mollick, E. R., Mollick, L., Shapiro, D. (2025). Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting. arXiv:2506.07142
- Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., Iwasawa, Y. (2022). Large Language Models are Zero-Shot Reasoners. NeurIPS 2022. arXiv:2205.11916
- Sprague, Z., Yin, F., Rodriguez, J. D., Jiang, D., Wadhwa, M., Singhal, P., Zhao, X., Ye, X., Mahowald, K., Durrett, G. (2025). To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning. ICLR 2025. arXiv:2409.12183
- Sclar, M., Choi, Y., Tsvetkov, Y., Suhr, A. (2024). Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design. ICLR 2024. arXiv:2310.11324
- Lu, Y., Bartolo, M., Moore, A., Riedel, S., Stenetorp, P. (2022). Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. ACL 2022. aclanthology.org/2022.acl-long.556
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. TACL, 12, 157-173. doi:10.1162/tacl_a_00638
- DeepSeek-AI (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature. doi:10.1038/s41586-025-09422-z
- Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., Zhou, D. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR 2023. arXiv:2203.11171