The pitch for this year is that one person can run a swarm. Orchestrate a fleet of AI agents, the story goes, and a single engineer outproduces a whole team: dozens of agents in flight, then hundreds, with serious talk of thousands, and a productivity ceiling far above the old 10x.

Some of it is real. A majority of companies now run agents in production. But of the enterprises that tried to scale those fleets, by early 2026 only about one in eight reached durable value, and roughly seven in eight stalled. The reason is older than software, and an army could have told you in advance: nobody commands a thousand of anything.

The constraint has a name, span of control, and the people who got it wrong paid in blood. They settled it the same way everywhere, a long time ago.

Four armies, one shape

The Roman legion built up from the contubernium, eight men who shared a tent, in steps of six to ten to a legion of five thousand. The Mongols, twelve centuries later with no Roman manual to copy, ran on strict tens, the arban of ten up to the tumen of ten thousand. The Inca, an ocean away with no contact with either, landed on almost the same decimal scheme. A NATO army, the structure Ukraine drilled in after 2014, branches three or four at a time from squad up to brigade.

The base unit is always a single-digit handful. The branching factor from each tier to the next is always small, somewhere between three and ten, never dozens. And the whole thing scales to hundreds of thousands without any leader ever directly holding more than a handful of the units below.

SystemBase unitEach tier groupsTop formation
Roman legioncontubernium, 8 men~6 to 10 of the unit belowlegion, ~5,000
Mongol hordearban, 10 menexactly 10 of the unit belowtumen, 10,000
Inca armychunka, 10 menexactly 10 of the unit belowhunu, 10,000
NATO / Ukrainesquad, ~93 to 4 of the unit belowbrigade / division

Ancient China had reached it earlier still, and Sun Tzu wrote the principle down: the control of a large force is the same as the control of a few, once you divide and number them. That kind of convergence, the way the same load finds the same kind of arch wherever it appears, is the signature of an invariant rather than a fashion.

The modern army inherited the shape rather than reinventing it, but re-proves it under fire, and that is the hardest test there is: a management fad fails a quarterly review, a broken command structure fails on a battlefield. What survives that is telling you about the problem, not the taste of the designer.

Why the number stays small

Two separate limits hold the span down, and the structure only works because it respects both.

The first is attention, which is only the raw material of decisions. One leader can properly know a handful of units: which is stuck, which is about to break, which is ready for more, caught early enough to decide on. There is a fixed amount of attention, so the good decisions it can feed are finite too.

Double the span and you do not get half the attention each, you cross a threshold past which the leader is no longer leading anyone, just reacting to whoever is loudest.

Graicunas put numbers on this in 1933: the relationships a supervisor tracks do not grow with the number of reports but explode combinatorially, every pair of subordinates its own thing to hold in mind.

The second is authority. The sergeant runs his squad because he is allowed to, without phoning the colonel. The Romans gave centurions real tactical discretion. The Mongols paired the rigid decimal structure with fast couriers and commanders expected to act on intent rather than wait for orders, the thing a later Prussian army would name Auftragstaktik and a modern one would call mission command. The point is identical across all of them: the tier that meets the situation is the tier that decides it.

An army that routed every decision to the top would lose to one whose sergeants could act. That is not a theory. It is the selection pressure that built the chart.

The newest war in Europe ran that experiment live. Ukraine, drilled since 2014 on NATO mission command, pushed decisions down to junior officers and a real corps of sergeants who adapted move by move without waiting for permission.

Russia brought the opposite: a thin NCO layer, initiative discouraged, orders from the top, so units stalled when the plan met reality and senior officers had to drive to the front to unstick them, where a striking number were killed. With comparable weapons at the start, the side whose sergeants could act outfought the side whose colonels had to.

The same maths runs your agents

Span of control does not care whether the units below are people or processes, and the binding constraint was never how many agents you can run. It is how many decisions the one principal can actually make, and make well. Every agent that hits something it cannot resolve, an ambiguity, a judgement call, a confident piece of nonsense that needs catching, throws a decision back up the line that only someone with authority can make.

A handful of agents and that flow is a trickle a human meets with care. A hundred and it is a flood; a thousand and the decisions are not made so much as waved through, and the ones the human never reaches are agents left to act on no decision at all.

The throughput does not rise because you spun up more workers, and the quality of each call drops as the queue grows, which is most of what the seven-in-eight failure rate is. It is not a tooling gap waiting on a better dashboard. It is a span-of-control wall, and what it caps is decisions, not agents.

And here is what the swarm pitch leaves out. An army of a thousand is not one commander shouting at a thousand soldiers. It is tiers, each holding a handful, each trusted to decide.

Your thousand agents have no sergeants. There is no middle that owns a hundred of them and makes the local calls, so every judgement that matters routes back to the one human at the top. That is the centralisation failure mode exactly: a platoon's worth of decisions, a squad's worth of hours, and no tier below that was allowed to act. The org chart says human in the loop; in practice the loop is theatre, because the decisions did not scale just because the agent count did.

You are not running a thousand agents. You are the one principal a thousand agents route their decisions to, and no banner reading leverage makes you decide faster, or better, than one mind can.

The armies already answered the design question, and it is two questions, not one:

  • how many decisions can the one person at the top actually make, and make well;
  • which of those a tier below the human gets to make on its own.

A fleet that answers neither is not flat and fast; it is a queue with one tired reviewer at the end of it.

The number moves, if you build the tiers

None of this makes the bet foolish, and the span is not a fixed integer. It is set by attention and by trust, and both can be raised. A commander with a radio, a map, and a trusted staff holds more than one with flags and runners, and good tooling does the same for an operator: views that surface only the exceptions, agents that summarise their own work, tests that shout when something breaks.

New technology genuinely moves the number, and the agent-fleet bet is partly a bet that it moves it a lot. That bet is worth making.

But notice how armies actually reached a thousand, the part the swarm pitch skips. They did not widen one span to a thousand. They added tiers. The way a single person sits above ten thousand is that they hold a handful of commanders, each holding a handful, down to the squad.

The reach is exponential in the depth of the tree, not the width of any one node. The honest form of "run a thousand agents" is a human over a few supervisors over their clusters, with checks at every level.

And here is the genuinely new thing, the one that could make this a revolution rather than a reorganisation: those middle tiers can be AI. A supervisor agent that owns a cluster of workers, triages their output, and escalates only what it cannot resolve is a sergeant. A deterministic gate that passes or fails every change without tiring is a sergeant that never sleeps.

Build that command tree and the human at the top really can sit over a thousand, because they are holding ten supervisors, not a thousand workers. The fleets that pay off will be the ones that built the army. The ones that stall are the flat swarms with nothing between the human and the work.

You scale the span by adding tiers, not by widening one. The new part is that the tiers can be agents. The old part is that you still have to be able to trust them.

And adding a tier is only half the move. The other half, the one the swarm pitch skips, is to make fewer decisions arrive at all. A design discipline like permaculture builds its systems so the parts meet each other's needs, and brings in a manager only for the remainder.

The agent version is the same. Deterministic gates that settle what they can, agents wired to consume each other's output instead of escalating, a fixed convention in place of a judgement call: every decision you design out is one that never reaches the bottleneck. You staff the queue with tiers, but you shrink it by design, and you shrink it first.

What you cannot delegate

The catch is the one the armies also knew: a sergeant works because authority is really delegated to him and because someone is accountable for what he does. An AI sergeant you cannot trust to decide is not a tier, it is another worker that routes everything upward, a layer added without reach gained.

Trusting it means one of two things. Either you genuinely delegate, and accept it will sometimes be wrong the way a human sergeant is, and design for that. Or you check it against something outside its own judgement, the harder and safer path, because an AI supervisor grading AI workers shares their blind spots, the reporter and the reported in a single failure domain.

The check that earns the trust is usually not another model but a deterministic oracle, the test that simply passes or does not. Skip it and you have not removed the bottleneck, you have hidden it inside a layer nobody audits, which is the same failure wearing a deeper org chart.

And under all of it there is a floor no tier ever takes, human or AI. You can delegate the work, and the swarm does the bottom layer faster than any room of juniors: more shipped, more explored, more things done. You can even delegate the middle, the triage and the local calls, to a supervisor that has earned the trust.

Delegate both of those, the juniors and the sergeants, and the multiplier is already in your hand: that is the 10x, and probably more, and notice that none of it required touching the top.

What does not delegate is the principal's own job, which was never the typing. It is choosing what should get done and what should not, the path the thing takes, the safety you have to ensure, and the answer for all of it when it goes wrong. That is not a bottleneck to engineer away; it is the irreducible part, and it was never the ceiling on throughput in the first place.

Keeping it human costs you almost nothing in speed, and it keeps the fleet yours. Hand it down to the agents and you have not delegated, you have abdicated, trading a tenfold gain you already held for a fleet that answers to no one.

The humans get the same treatment

None of this is only about agents. Companies are cutting managers: Meta, Amazon, Google, and Intel have thinned their middle layers, and the average manager now carries around twelve reports and rising. The move has a name, the Great Flattening, and the same pitch as the swarm: leaner, faster, more leverage. Gartner expects a fifth of organisations to use AI to flatten by 2026, the two stories becoming one.

Cut the middle but keep every decision at the top and you break both halves. The remaining manager inherits a platoon's worth of people with a squad's worth of hours, and every call routes upward because nobody below was given the power to make it.

The coaching disappears first; what is left is clearing a queue, for a job whose title still says develop people. The word doing the damage is lean, a lossy, flattering projection that reads well from the top, and "leverage" is just this decade's word for lean.

The Great Army of no-one

There is a failure worse than the bottleneck, and the swarm is how you reach it. A bottleneck still points somewhere: the work piles up behind a human who, slowly, decides. But an agent is called an agent for a reason: it acts on behalf of another, a principal whose intent it carries.

Strip out the tiers and you cut that line of intent, not by delegating the work, the whole point of the exercise, but by delegating the deciding: pushing the principal's own job onto things that were only ever meant to carry it out. An agent nobody is directing does not stop. It keeps acting, just not on anyone's behalf, and a thing that acts for no principal is no longer your agent. It is an agent of no-one.

Multiply that by a thousand and you have not built a fleet, you have raised a Great Army of no-one: motion without command, effort without intent, a thousand workers each confidently executing a purpose nobody set.

And it cuts both ways, which is the honest part. Working with agents we might genuinely do more than we ever could with a team of humans, which is the whole reason to try. But the army that cannot be steered where its general wants it goes instead where no one wants it, and history's name for that is a rout, or a mob. The danger of the swarm is not that it outruns your attention; it is that the moment it does, it stops being yours.

The constraint is two facts: attention is finite and does not transfer, and authority either lives where the situation is met or it bottlenecks at the top. So the questions are the old ones, and answerable: how many decisions can one person actually make well, and what has the tier below been trusted to decide.

Get them right and a thousand agents, or a flat org, can work. Get them wrong and you have raised a Great Army of no-one, given it a confident name, and the people, or the outputs, in the middle are the ones who pay.

An agent no one can direct directs itself, where no one wants it. A thousand of them is not your fleet. It is a Great Army of no-one.

The pitch this year: one engineer runs a fleet of AI agents and outproduces a team, dozens then a thousand. By early 2026 only one in eight reached durable value; the rest stalled. The reason is older than software: nobody commands a thousand of anything.

Span of control

Every army that scaled hit the same wall. Rome, the Mongols and the Inca, with no contact between them, converged on one shape: a base unit of a handful, each leader holding only a few units below, a convergence that marks an invariant, not a fashion. Two limits force it:

  • Attention: one person tracks only a few things well. Graicunas showed in 1933 that the relationships explode combinatorially, so past a point you react to whoever is loudest.
  • Authority: whoever meets the situation must decide it. Ukraine, on NATO mission command, pushed calls to its sergeants and outfought a Russia whose colonels had to.

The bottleneck is decisions, not agents

Span of control does not care if the units below are people or processes. The limit was never how many agents you can run but how many decisions you can make well, and every agent that hits an ambiguity throws one back up.

A handful is a trickle; a thousand is a flood, the calls waved through, not made. Your agents have no sergeants: no middle tier makes the local calls, so every judgement routes to one human, and the loop is theatre.

You are not running a thousand agents. You are the one principal a thousand agents route their decisions to.

The fix, and the part you keep

Armies reached a thousand by adding tiers, not widening one span: reach comes from the tree's depth, not a node's width. The new thing: the tiers can be AI, a supervisor that triages a cluster and escalates only what it cannot resolve.

But an AI supervisor grading AI workers shares their blind spots, so the check that earns trust is a deterministic oracle, not another model. What never delegates is the principal's job: what to do, and who answers when it breaks. Hand off the work and the middle and that is your 10x; hand the top down and you abdicate.

A Great Army of no-one

An agent acts on behalf of someone. Delegate the deciding, not just the work, and you cut the line of intent: it keeps acting, not on anyone's behalf. A thousand of them is a Great Army of no-one.

The same is happening to people: the Great Flattening cuts middle managers on the same pitch, leaving each survivor a platoon of reports on a squad's hours. "Leverage" is just this decade's word for "lean".


Sources