Skip to content

Research Note

Sprawl persists because it works at first

By James Carter · July 2026

Part of a series on the five elements of the Flag Model. Earlier notes covered decision paralysis, after-action reviews, and why mandating candor backfires.

Every executive team carrying a dozen half-finished initiatives has been told to focus. The advice is correct and useless, in that order. It names a destination and no mechanism, which is why teams agree with it, kill nothing, and add two more things by June.

A disclosure. This note is about the mechanism, and the evidence here is the weakest of the four disciplines we have written about. There is no meta-analysis of the mechanism itself. Not one study examined an executive team. One of the sources below is a simulation rather than a measurement. We are telling you that up front because the advice industry states this case with more confidence than the research supports — and because you should weight what follows accordingly.

What the research shows

Sull, Homkes and Sull (Harvard Business Review, 2015), surveying 7,600 managers across 262 companies, documented the conditions that let sprawl accumulate. Fewer than a third of managers believed their organizations reallocated funds quickly enough to the right places. Only one in five said their organization did a good job of moving people across units to support strategic priorities. Half of the middle managers surveyed believed they could secure significant resources for attractive opportunities that fell outside the company’s strategic objectives. And when asked what obstructed their understanding of strategy, middle managers were four times more likely to cite the sheer number of corporate priorities and initiatives than to cite a lack of clarity in how strategy was communicated.

KC and Terwiesch (Management Science, 2009) analyzed operational data from two hospital services and found something that complicates the standard advice. Workers speed up as load rises. A 10 percent increase in load reduced length of stay by two days in cardiothoracic surgery. But the acceleration does not hold. Sustained high load reversed it, with a 1 percent increase in their overwork measure adding roughly six hours to length of stay, and quality measures degraded alongside it.

Aral, Brynjolfsson and Van Alstyne (Information Systems Research, 2012) studied an executive recruiting firm using five years of project accounting data across more than 1,300 projects, 125,000 email messages, and survey data on the same workers. Their published finding is that more multitasking was associated with more project output, subject to diminishing marginal returns.

Repenning (Journal of Product Innovation Management, 2001) built a formal model of firefighting in a multi-project development environment. The model produces two results worth carrying: firefighting is self-reinforcing, and multi-project systems are considerably more vulnerable to it than practitioners assume. Past a tipping point, the system settles into a stable equilibrium where reacting to late-stage problems becomes the actual development process.

Microsoft (Work Trend Index, 2025) measured the environment sprawl produces, using aggregated Microsoft 365 telemetry through February 2025 alongside a global survey. In the survey, 48 percent of employees and 52 percent of leaders said their work felt chaotic and fragmented. Roughly a third of meetings now span multiple time zones, up 35 percent since 2021, and meetings after 8 p.m. are up 16 percent year over year.

On what these are and are not. A large-sample manager survey, an econometric study of hospital operations, an econometric panel study of one recruiting firm, a simulation, and product telemetry published by the company selling the remedy. Different methods, and no shared population, which is the honest form of convergence. But hospitals are not leadership teams, recruiters are not executives, a model demonstrates what follows from its assumptions rather than what happens in the world, and telemetry from one software vendor’s tenants is not a random sample of working life. Sull’s sample had median annual sales near $430 million, overlapping the mid-market band, with average headcount around 6,000 and a sector mix that does not.

Two corrections worth making, one of them on ourselves.

The most-quoted claim in this area is that multitasking has an inverted-U relationship with productivity, with completion rates declining beyond an optimum. That language appears in the 2007 NBER working paper version of the Aral study. The peer-reviewed version published in Information Systems Research states the weaker result: more output with diminishing returns. Diminishing returns and decline are different claims. We use the published one. If you have seen the inverted U cited, it was probably sourced from the working paper.

The second is the figure you have almost certainly seen quoted from the Microsoft research: that employees are interrupted every two minutes. Microsoft’s own methodology note states that this figure is based on the top 20 percent of users by ping volume received. It is not the average worker — it is the most-pinged fifth of them. That qualifier is stated plainly by Microsoft and dropped by nearly everyone who repeats the number. We are not using the two-minute figure, because the version of it in circulation is not what the data says.

What this actually explains

The research names the symptom and the advice industry supplies a target. Neither explains the thing CEOs actually find baffling, which is why intelligent people who agree that focus matters keep adding work anyway.

Here is the claim we make that the research does not.

Sprawl persists because it works at first.

Read KC and Terwiesch again. Load produced a real, measurable speedup before it produced the reversal. That is not a curiosity — it is the entire reason the pattern is stable. Add a priority to a loaded leadership team and the following weeks look good. People move faster, meetings get crisper, the sense of momentum is genuine rather than imagined. The decision to add appears to have been validated. Then the cost arrives one or two quarters later, arriving as slippage rather than as a signal, by which point it is attributable to a hire who did not work out, a competitor move, a vendor delay — anything but the initiative added in March.

That delay between cause and effect is what makes this feel unfixable. Repenning’s model gives it a name and a shape: a self-reinforcing loop with a tipping point past which the degraded mode becomes the operating mode. The team is not failing to see the problem. It is seeing a problem, correctly, and misattributing it, because the evidence available in the moment genuinely points somewhere else.

We flag this as our synthesis. KC and Terwiesch measured hospital workers, not leadership teams, and made no claim about how managers interpret the delay. Repenning modeled product development. Connecting the load reversal to the misattribution is our reading of what the two imply together.

This is why the Rhythm rebuild is a limit rather than a priority list. A priority list is a statement of intent that survives contact with nothing. A limit is a rule that forces the trade to happen at the moment of decision, when the cost is still visible, rather than two quarters later when it is not. One in, one out means the price of the new thing is paid in the room where the new thing is proposed.

We should be clear that the specific prescription is not directly evidenced. No study we found tested initiative limits on an executive team. The case for a limit is an inference from queueing logic and from the load research, and it is offered as such.

There is also one body of evidence that complicates it, and we would rather raise it ourselves. Leicht-Deobald and colleagues (Journal of Management, 2025) meta-analyzed thirty years of research on team boundary management across 85 studies and 10,848 teams. Boundary management was positively associated with team performance overall. But the association was stronger for boundary spanning — coordination, representation and information search — than for boundary strengthening — buffering, guarding and shielding the team from outside demands.

Read carefully, that is a complication rather than a contradiction, and the distinction matters. Boundary strengthening was still positively associated with performance. It was weaker, and it rested on nine studies and 842 teams with a confidence interval running from .06 to .45 and a credibility interval that crosses zero. So the honest statement is that protective activity appears to help less than outward-facing activity, on thin and imprecise data — not that protection fails.

The inferential distance also runs against overreading it. A limit on how many things a leadership team carries is not the same construct as buffering a team from external parties, which is what this literature measures. We raise it because a limit is a protective move, and the best available evidence on protective moves says they are the weaker half of the picture. If your leadership team institutes a limit and does nothing about coordination with the functions around it, this meta-analysis suggests you have chosen the smaller of two available gains.

The test

If Rhythm is what broke on your team, you would predict a specific timing pattern rather than a general sense of overload.

Take the last significant initiative your team added, and find the date. Then look at what happened to delivery in the six to ten weeks that followed. If throughput improved, note that, because it is the part that fooled you. Now look at the two quarters after that, and at what you attributed the slowdown to at the time.

If the degradation began before the event you blamed it on, you have your answer, and no amount of prioritization language will fix it. If throughput never improved after the add, something else is wrong and this is not your discipline.

Closing the set

Four disciplines run off a central set of operating priorities, and execution fails when one of them goes. The Flag Model sets out all five and what rebuilding each one requires.

A final word on the evidence. It is not equally strong across the five. The Learning discipline has three meta-analyses behind it. The Standard has two that disagree and a study of 70 top management teams that resolves them. The Decision has longitudinal and case evidence but no price tag anyone can defend. The Flag has the largest samples in the set. Rhythm has what you have just read, which is thinner than any of it.

We would rather tell you that than have you find out from someone else.

Sources

All primary. Every figure was verified against the original study or article text.

“James helped us turn seven groups of rivals into one leadership team. By the time US Holdings became Eagle Manufacturing Group and I moved from COO to CEO, we were no longer seven companies protecting our own territory — we were one company working toward the same outcome.”

Ronn Page · former CEO, Eagle Manufacturing Group

As featured in CNN CNN Money Business Insider Client results →
James Carter, founder of Be Legendary

About the author

James Carter

Founder of Be Legendary and creator of the Flag Model™. Twenty-five years inside executive teams; co-author alongside Stephen Covey, Ken Blanchard, Deepak Chopra & Brian Tracy, and featured on CNN and in Business Insider. More about James →

See where your team breaks first.

A Calibration Call is 15 minutes — you leave with a concrete read on your team, whether or not we work together. It's a calibration, not a pitch.

Book a Calibration Call

Not ready to talk? Start free

Take the free Break-Point Self-Assessment

See which of the five components your team is most likely to lose first — in ~4 minutes. No email required to see your result.

Start the assessment →

Field notes, by email

Straight thinking on executive-team execution.

One short note, roughly monthly, on the disciplines that decide whether a leadership team executes. No fluff, no pitch. Unsubscribe anytime.

Share this
LinkedIn Email