All articles
5 min readlearning / training

0.46 SD more learning, 0.56 SD worse ratings: fix your feedback form

Every training day I run ends with a rating, and the case study on this site quotes one of them: NPS 9.2 out of 10. That number decides whether a client books the next cohort. A randomized experiment in Harvard physics classes found students scoring 0.46 standard deviations higher after the sessions they rated 0.56 standard deviations lower, so my best day and my highest score can land on different afternoons.

What the studies measured.

Scott Freeman's group pooled 225 comparisons of undergraduate STEM courses taught by lecture against courses taught with active formats. Examination performance rose by 0.47 standard deviations under active learning, worth around 6% on a typical exam. Failure rates fell from 33.8% under lecturing to 21.8% 1. The format gap is wide enough to move a letter grade.

Louis Deslauriers built the experiment that separates learning from the sense of learning. In a large introductory physics course at Harvard, two instructors taught the same content from the same handouts to randomly assigned halves of the class. In the first session, on static equilibrium, instructor A ran an active format while instructor B lectured. In the second session, on fluids, they swapped. Every student took a 12-question test at the end of the hour and filled in a survey with items such as "I feel like I learned a lot from this lecture". The active groups scored 0.46 standard deviations higher on the test and reported a feeling of learning 0.56 standard deviations lower 2. For the same students in the same hour, the two measures pointed in opposite directions.

Shana Carpenter's team isolated the delivery on its own. They filmed a 31-minute science lecture twice with the same script. In one take the instructor stood upright, held eye contact and spoke without notes. In the other he slumped, looked away and read haltingly. Viewers of the polished version predicted they would recall far more of the content and scored the instructor higher on standard evaluation items. Both groups then recalled about 25% of the material 3. The delivery moved the evaluation and left recall where it started.

The forms we hand out at the end have their own literature. George Alliger pooled 34 training studies and 115 correlations: affective reactions, the "I enjoyed this" family of questions, correlated 0.02 with immediate learning, while utility reactions, the "I can use this in my job" family, reached 0.26 4. Bob Uttl went back over the multisection literature on university teaching evaluations and put the ratings-to-learning correlation at 0.08, falling to about zero once prior ability is accounted for 5. A room that puts you at the top of the scale can still walk out without the material.

The perception penalty is not automatic.

Peter Boedeker ran a two-day randomized cross-over trial with 146 second-year medical students working the same clinical cases in both conditions: a lecture with minimal interaction, and a large-group session where teams worked the cases themselves. The interactive format produced test scores 0.27 standard deviations higher (p = 0.010), and the feeling of learning came out 0.56 standard deviations higher as well 6. Students in the lower half of prior achievement gained the most.

Both directions have now been measured, so the reversal Deslauriers found is a risk rather than a rule. Expectation looks like the moving part: medical students arrive trained on case work, physics undergraduates arrive expecting a lecture. Deslauriers acted on that reading in the following term by opening the semester with a 20-minute explanation of these results, and students reported that the framing helped them read their own effort correctly 2.

The protocol.

  1. Replace the satisfaction question with a utility question. Drop "how satisfied were you" from the top of the form. Ask two items instead: "Name one thing from today you will use in the next two weeks" and "How likely are you to apply it on your next task, 0 to 10". Utility items carry a 0.26 correlation with learning against 0.02 for enjoyment 4.
  2. Score the room before it leaves. Ten minutes, 8 to 12 closed-book questions on the day's material, results reported to the client next to the rating. Without that column you have a mood reading and nothing else.
  3. Open with the trade. In the first 15 minutes, tell the room that the work will feel harder than a lecture and that the difficulty is what moves the score. Give them the two numbers from the Harvard experiment. Deslauriers found the framing changed how students received active instruction 2.
  4. Keep the exercise you were about to cut. A day gets pleasant by removing the parts where people struggle in front of colleagues, and those are the parts holding the 0.47 standard deviations.
  5. Rehearse for clarity, then step back. Your fluency raises the evaluation without raising recall 3. Polish the 10 minutes of explanation, then hand the next 40 to the room.
  6. Ask for the 8-week measure. Write into the proposal that the client reports what changed in the work: tickets closed, documents produced, whichever artefact the training targeted. Utility reactions predict transfer to the job better than any learning measure taken on the day 4.

I still read the ratings. They tell me whether the room felt respected and whether the pacing worked, and a session people resent teaches nothing at all. For the question a client is paying to answer, the exit test and the 8-week report do the work. Run your next session with both columns filled and compare them.

Sources.

  1. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410-8415. doi.org/10.1073/pnas.1319030111
  2. Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences, 116(39), 19251-19257. doi.org/10.1073/pnas.1821936116
  3. Carpenter, S. K., Wilford, M. M., Kornell, N., & Mullaney, K. M. (2013). Appearances can be deceiving: instructor fluency increases perceptions of learning without increasing actual learning. Psychonomic Bulletin & Review, 20(6), 1350-1356. doi.org/10.3758/s13423-013-0442-z
  4. Alliger, G. M., Tannenbaum, S. I., Bennett, W., Traver, H., & Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology, 50(2), 341-358. doi.org/10.1111/j.1744-6570.1997.tb00911.x
  5. Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22-42. doi.org/10.1016/j.stueduc.2016.08.007
  6. Boedeker, P., Schlingmann, T., Kailin, J., Nair, A., Foldes, C., Rowley, D., Salciccioli, K., Maag, R., Moreno, N., & Ismail, N. (2025). Active versus passive learning in large-group sessions in medical school: a randomized cross-over trial investigating effects on learning and the feeling of learning. Medical Science Educator, 35(1), 459-467. doi.org/10.1007/s40670-024-02219-1