Research Note
Your post-mortem would not have qualified
By James Carter · July 2026
Part of a series on the five elements of the Flag Model. Earlier notes covered what decision paralysis actually costs and why mandating candor backfires.
There is one practice in the management literature with an unusually strong evidence base behind it. It is cheap, it takes under twenty minutes, and three separate meta-analyses agree it works. Almost no executive team runs it.
That gap is the subject of this note, and the interesting part is not that teams neglect a good idea. It is that most of them believe they already do it.
What independent research agrees on
Tannenbaum and Cerasoli (Human Factors, 2013) meta-analyzed debriefs and after-action reviews across 46 independent samples drawn from 31 studies, combined sample 2,136. Average improvement over control was roughly 25 percent, an effect size of d = 0.67, falling to d = 0.54 with three outliers removed. Effects held for teams and individuals, in simulated and real settings, in medical and nonmedical populations, and under both within-group and between-group designs. The average session lasted about eighteen minutes.
Keiser and Arthur (Journal of Applied Psychology, 2021) rebuilt the analysis on a larger base: 61 studies, 107 effect sizes, 915 teams and 3,499 individuals. They found d = 0.79, larger than the 2013 estimate and larger than some of the biggest training-method effects in the literature. Two characteristics consistently contributed: alignment of the review to the individual or the team, and use of objective rather than subjective performance-review media.
Keiser and Arthur (Journal of Business and Psychology, 2022) extended it again to 83 studies, 134 effect sizes, 955 teams and 4,684 individuals, finding an overall d = 0.92. This is the finding that matters most here, and it is not the number. They report that the AAR is most effective in task environments combining high complexity with ambiguity — environments that supply no intrinsic feedback about whether you did well. They identify those as most often project and decision-making tasks, and they note that such tasks are concentrated in industries that do not traditionally use after-action reviews.
Sull, Homkes and Sull (Harvard Business Review, 2015) surveyed 7,600 managers across 262 companies. Fewer than a third said they could have open and honest discussions about the most difficult issues. A third said many important issues were treated as off limits entirely.
Three independent meta-analyses, conducted by two different research groups nine years apart over substantially different bodies of primary studies, all point the same direction and none contradicts the others. That is a stronger evidentiary position than anything else written about executive teams. The effect estimates rise as the base widens, which is unusual and worth noting rather than exploiting: the honest reading is that the true effect sits somewhere in a range roughly bounded by d = 0.54 and d = 0.92, and that the practice works. Anyone quoting a single precise percentage from this literature has picked one study.
On segment, and this one matters. None of these studies examined an executive team directly. The underlying research is drawn from military, medical, aviation and educational settings, with average team sizes around five. That was the honest limitation of this note in its first version, and the 2022 finding narrows it without closing it. Keiser and Arthur did not study executives — but they did identify the task profile where AARs pay off most, and it is the profile executive work fits: complex, ambiguous, and lacking any built-in signal about whether the call was right. The inference from that task profile to your leadership team is ours. The task profile itself is theirs.
Sull’s sample had median annual sales around $430 million, which sits inside the mid-market band, alongside average headcount near 6,000 and a sector concentration that does not.
What the research names, and what it misses
The advice that follows from all of this is “run after-action reviews.” Most executive teams hear that and conclude they already do. They hold post-mortems. They run quarterly business reviews. They do 360s.
Here is the claim we make that the research does not.
Those are not after-action reviews, and the 2013 meta-analysis says so in its own methodology. Before Tannenbaum and Cerasoli could measure anything, they had to define a debrief tightly enough to decide what to exclude. They settled on four required elements.
Active self-discovery. Participants surface the lessons themselves, rather than passively receiving feedback.
Developmental, non-punitive intent. The purpose is learning, not evaluation or administration.
Focus on a specific event. A particular episode, not general performance.
Multiple information sources. Several accounts, not one person’s version.
Run your leadership team’s practices through those four.
The quarterly business review fails on intent. It is evaluative by design, and the paper explicitly excludes performance appraisals and reviews.
The post-mortem after a bad quarter usually fails on intent too, because everyone in the room knows it is partly about who is accountable. It often fails on specificity as well, drifting from a particular episode into general performance.
The 360 fails on specific events. The authors name this exclusion directly.
The version where the CEO explains what went wrong fails on active self-learning, and it fails on multiple information sources, because it is one account.
Most executive teams have never conducted a single session that would have qualified for inclusion in any of these three meta-analyses. Not one. That is checkable, which is the point of putting it in writing.
Why it isn’t a process problem
Look at the four elements again and notice what three of them have in common. Non-punitive intent. Multiple sources instead of the boss’s version. Self-discovery instead of being told. None is a technique. All three are conditions about whether the room is safe enough for people to say what they actually saw.
Edmondson (Administrative Science Quarterly, 1999) supplies the mechanism. Studying 51 work teams in a manufacturing company, she found that psychological safety predicted learning behavior, and that learning behavior mediated the relationship between safety and team performance. Team efficacy did not survive as a predictor once safety was controlled for. Set that against Sull’s numbers — where fewer than a third of managers report being able to discuss the hardest issues openly — and the picture resolves.
We flag this as our synthesis. Edmondson measured manufacturing teams and a specific construct; Sull measured self-reported candor in a different population. Our reading is that the reason your team does not run real after-action reviews is not that it lacks a protocol. Protocols are free and everywhere. It is that the protocol requires a room where being wrong out loud is survivable — and installing a protocol into a room that isn’t will produce a meeting that satisfies the form and none of the four elements.
This is why “we should do better retrospectives” reliably changes nothing. It treats a discipline as a calendar item.
Start with the quarter that went well
Ellis and Davidi (Journal of Applied Psychology, 2005) ran a quasi-field experiment with soldiers on repeated navigation exercises. Those who reviewed both successes and failures after each training day improved significantly more than those who reviewed only failures. Their second finding explains why: before any intervention, participants’ mental models of events that went badly were already richer than their models of events that went well.
The instinct is to examine failures. Reviewing what went wrong largely re-treads ground the team has covered privately for weeks. The win is the unexamined case.
There is a second reason to start there, and this one is ours. The successful quarter is the safest thing your team can put on the table. If the room cannot yet examine a failure without it becoming an accountability conversation, examining a success is the entry point that does not require the safety to exist first. It builds it.
The test
If the Learning discipline is what broke on your team, you would predict two things. The same category of mistake returns each year wearing a different name. And the team can describe what happened in detail but cannot name what it changed about how it operates.
So run the test. Take a decision from six months ago that went badly and ask what the team would do differently. If the answer is a list of what went wrong, that is description, not learning. Then run it again on something that went well. If the team can do the first exercise and not the second, your constraint is not process. It is safety, and no framework fixes that.
Which one broke
Learning is one of four disciplines that run off a central set of operating priorities, and execution fails when one of them goes.
The Flag Model sets out the five elements and what rebuilding each one requires. Eighteen minutes is not the hard part.
Sources
All primary. Every figure was verified against the original study text or the publishing journal’s record.
- Tannenbaum, S. I., & Cerasoli, C. P. (2013). Do Team and Individual Debriefs Enhance Performance? A Meta-Analysis. Human Factors, 55(1), 231–245. ↗
- Keiser, N. L., & Arthur, W., Jr. (2021). A Meta-Analysis of the Effectiveness of the After-Action Review (or Debrief) and Factors That Influence Its Effectiveness. Journal of Applied Psychology, 106(7), 1007–1032. ↗
- Keiser, N. L., & Arthur, W., Jr. (2022). A Meta-Analysis of Task and Training Characteristics That Contribute to or Attenuate the Effectiveness of the After-Action Review (or Debrief). Journal of Business and Psychology, 37, 953–976. ↗
- Ellis, S., & Davidi, I. (2005). After-Event Reviews: Drawing Lessons From Successful and Failed Experience. Journal of Applied Psychology, 90(5), 857–871. ↗
- Edmondson, A. C. (1999). Psychological Safety and Learning Behavior in Work Teams. Administrative Science Quarterly, 44(2), 350–383. ↗
- Sull, D., Homkes, R., & Sull, C. (2015). Why Strategy Execution Unravels — and What to Do About It. Harvard Business Review, March 2015. ↗
“James helped us turn seven groups of rivals into one leadership team. By the time US Holdings became Eagle Manufacturing Group and I moved from COO to CEO, we were no longer seven companies protecting our own territory — we were one company working toward the same outcome.”
Ronn Page · former CEO, Eagle Manufacturing Group
About the author
James Carter
Founder of Be Legendary and creator of the Flag Model™. Twenty-five years inside executive teams; co-author alongside Stephen Covey, Ken Blanchard, Deepak Chopra & Brian Tracy, and featured on CNN and in Business Insider. More about James →
See where your team breaks first.
A Calibration Call is 15 minutes — you leave with a concrete read on your team, whether or not we work together. It's a calibration, not a pitch.
Not ready to talk? Start free
Take the free Break-Point Self-Assessment
See which of the five components your team is most likely to lose first — in ~4 minutes. No email required to see your result.
Field notes, by email
Straight thinking on executive-team execution.
One short note, roughly monthly, on the disciplines that decide whether a leadership team executes. No fluff, no pitch. Unsubscribe anytime.