Skip to content

Research Note

Your post-mortem would not have qualified

By James Carter · July 2026

Second in a series. The first examined what decision paralysis actually costs.

There is one practice in the management literature with an unusually strong evidence base behind it. It is cheap, it takes under twenty minutes, and it works about as well for teams as for individuals. Almost no executive team runs it.

That gap is the subject of this note, and the interesting part is not that teams neglect a good idea. It is that most of them believe they already do it.

What independent research agrees on

Four bodies of work, built with genuinely different methods on genuinely different populations, converge.

Tannenbaum and Cerasoli (Human Factors, 2013) ran a meta-analysis of debriefs and after-action reviews across 46 independent samples drawn from 31 studies, with a combined sample of 2,136. Average improvement over control conditions was roughly 25 percent. Effects held for teams and individuals, in simulated and real settings, in medical and nonmedical populations, and under both within-group and between-group designs. The average session in the studies they reviewed lasted about eighteen minutes.

Ellis and Davidi (Journal of Applied Psychology, 2005) ran a quasi-field experiment with soldiers on repeated navigation exercises. Those who reviewed both their successes and their failures after each training day improved significantly more than those who reviewed only their failures. The authors also found something less expected: before any intervention, participants’ mental models of the events that went badly were already richer than their models of the events that went well.

Edmondson (Administrative Science Quarterly, 1999) studied 51 work teams inside a manufacturing company using a multimethod design and found that team psychological safety predicted learning behavior, and that learning behavior mediated the relationship between safety and team performance. Team efficacy did not survive as a predictor once psychological safety was controlled for.

Sull, Homkes and Sull (Harvard Business Review, 2015) surveyed 7,600 managers across 262 companies. Fewer than a third said they could have open and honest discussions about the most difficult issues. A third said many important issues were treated as off limits entirely.

A meta-analysis of quasi-experiments, a field experiment, a multimethod field study, and a large-sample manager survey. Soldiers, surgical teams, manufacturing crews, corporate managers. No shared instrument, no shared population, no shared decade. They were not built to confirm each other, which is why the agreement between them carries weight that any one of them would not.

On segment, and this one matters. Not one of these studies examined an executive team. Tannenbaum and Cerasoli drew primarily on military, medical, aviation and educational contexts, with an average team size of about five. Ellis and Davidi studied soldiers navigating terrain. Edmondson studied one manufacturing company. Sull’s sample had median annual sales around $430 million — inside the mid-market band — but average headcount near 6,000 and a heavy concentration in financial services, IT, telecom and oil and gas.

So the honest position is this: the best-evidenced practice in the management literature has essentially never been tested on the people running the company. We think it transfers. We are telling you plainly that transfer is an inference, not a finding.

And on the number. The figure you will see quoted everywhere is 25 percent. The authors’ own conservative estimate, with three outlier studies removed, is 21 percent. Studies using objective performance criteria produced smaller effects than those using subjective ratings, and even those landed near 20 percent. Most of the underlying research is quasi-experimental, and the authors say directly that causal inference should be drawn carefully. Take 20 percent, not 25.

What the research names, and what it misses

The advice that follows from all of this is “run after-action reviews.” Most executive teams hear that and conclude they already do. They hold post-mortems. They run quarterly business reviews. They do 360s.

Here is the claim we make that the research does not.

Those are not after-action reviews, and the meta-analysis says so in its own methodology. Before Tannenbaum and Cerasoli could measure anything, they had to define a debrief tightly enough to decide what to exclude. They settled on four required elements.

01

Active self-discovery. Participants surface the lessons themselves, rather than passively receiving feedback.

02

Developmental, non-punitive intent. The purpose is learning, not evaluation or administration.

03

Focus on a specific event. A particular episode, not general performance.

04

Multiple information sources. Several accounts, not one person’s version.

Run your leadership team’s practices through those four.

The quarterly business review fails on intent. It is evaluative by design, and the paper explicitly excludes performance appraisals and reviews.

The post-mortem after a bad quarter usually fails on intent too, because everyone in the room knows it is partly about who is accountable. It often fails on specificity as well, drifting from a particular episode into general performance.

The 360 fails on specific events. The authors name this exclusion directly.

The version where the CEO explains what went wrong fails on active self-learning, and it fails on multiple information sources, because it is one account.

Most executive teams have never conducted a single session that would have qualified for inclusion in this meta-analysis. Not one. That is checkable, which is the point of putting it in writing.

Why it isn’t a process problem

Look at the four elements again and notice what three of them have in common. Non-punitive intent. Multiple sources instead of the boss’s version. Self-discovery instead of being told. None of those is a technique. All three are conditions about whether the room is safe enough for people to say what they actually saw.

Edmondson’s finding supplies the mechanism. Learning behavior is what carries psychological safety through to performance. Without the safety, the behavior does not appear, and without the behavior, nothing transfers. Set that against Sull’s numbers — where fewer than a third of managers report being able to discuss the hardest issues openly — and the picture resolves.

We flag this as our synthesis rather than either study’s finding. Edmondson measured manufacturing teams and a specific construct; Sull measured self-reported candor in a different population. Neither set out to explain executive-team debriefs. Our reading is that the reason your team does not run real after-action reviews is not that it lacks a protocol. Protocols are free and everywhere. It is that the protocol requires a room where being wrong out loud is survivable — and installing a protocol into a room that isn’t will produce a meeting that satisfies the form and none of the four elements.

This is why “we should do better retrospectives” reliably changes nothing. It treats a discipline as a calendar item.

Start with the quarter that went well

Ellis and Davidi’s result points somewhere counterintuitive and useful. The instinct is to examine failures, and the teams that examined only failures learned less than the teams that examined both. Their second finding explains why. People already understand their failures in more detail than their successes. Reviewing what went wrong is largely re-treading ground the team has covered privately for weeks. The win is the unexamined case.

There is a second reason to start there, and this one is ours rather than theirs. The successful quarter is the safest thing your team can put on the table. If the room cannot yet examine a failure without it becoming an accountability conversation, examining a success is the entry point that does not require the safety to exist first. It builds it.

The test

If the Learning discipline is what broke on your team, you would predict two things. The same category of mistake returns each year wearing a different name. And the team can describe what happened in detail but cannot name what it changed about how it operates.

So run the test. Take a decision from six months ago that went badly and ask what the team would do differently. If the answer is a list of what went wrong, that is description, not learning. Then run it again on something that went well. If the team can do the first exercise and not the second, your constraint is not process. It is safety, and no framework fixes that.

Which one broke

Learning is one of four disciplines that hold a team upright, and execution fails when one of them goes.

The Flag Model sets out which four disciplines hold a team upright — the Decision, the Rhythm, the Standard, the Learning — and what rebuilding each one requires. Eighteen minutes is not the hard part.

Sources

All primary. Every figure was verified against the original study text.

“This truly resonated and will have a lasting influence on us as we work to create a technology organization where the best and brightest choose to be.”

Larry Quinlan · Global CIO, Deloitte

As featured in CNN CNN Money Business Insider Client results →
James Carter, founder of Be Legendary

About the author

James Carter

Founder of Be Legendary and creator of the Flag Model™. Twenty-five years inside executive teams; co-author alongside Stephen Covey, Ken Blanchard, Deepak Chopra & Brian Tracy, and featured on CNN and in Business Insider. More about James →

See where your team breaks first.

A Calibration Call is 15 minutes — you leave with a concrete read on your team, whether or not we work together. It's a calibration, not a pitch.

Book a Calibration Call

Not ready to talk? Start free

Take the free Break-Point Self-Assessment

See which of the five components your team is most likely to lose first — in ~4 minutes. No email required to see your result.

Start the assessment →

Field notes, by email

Straight thinking on executive-team execution.

One short note, roughly monthly, on the disciplines that decide whether a leadership team executes. No fluff, no pitch. Unsubscribe anytime.

Share this
LinkedIn Email