sqwyz.
22 September 2026 · Written by Scott Robertson

The Scalability Gate: Why Going to Scale Too Early Is More Dangerous Than Going Too Late

Rows of warehouse shelving stacked with boxes and pallets

When an AI programme fails after the pilot succeeded, the post-mortem almost always lands on the same conclusion: the rollout didn't scale. I think that conclusion is usually wrong, or at least incomplete. I've written elsewhere about why a pilot that worked doesn't guarantee a rollout that lands. The failure I'm describing here happens earlier still, at the moment someone decides to scale, before anyone has checked whether the organisation is ready to carry what the pilot proved.

Every steering group I've sat in treats "we moved too slowly" as the risk to avoid and "we moved fast" as evidence of good management. Caution gets treated as the failure mode. In my experience it's the other way round. Scaling too early is both the more common mistake and the more expensive one, and it's rarely named as a mistake at all, because it looks like momentum until the costs surface.

Four dimensions, and the one that gets skipped

A proper scalability assessment checks readiness across four dimensions: technical scalability, organisational readiness, change impact, and commercial viability. Each one can independently sink a scaled deployment, and a pilot that scores well on one tells you nothing about the other three. The AI Transformation Playbook at transformationplaybook.ai has a structured version of this assessment if you want a starting framework.

Technical scalability is usually the dimension everyone checks, because it's visible and it's the one the vendor will happily discuss. Can the system handle production volumes, integrate with existing infrastructure, and hold its performance under load? This gets tested because it's testable, and because a technical failure is embarrassing in a way that shows up immediately.

Organisational readiness is the dimension that gets skipped, because checking it properly means admitting the operational team inheriting the capability doesn't yet have the skills, the context, or the support structure the experiment team built up over months. The experiment team knows every workaround and every failure mode. None of that transfers by default. It has to be documented, handed over, and supported through a transition period, and none of that work shows up in a demo.

Change impact and commercial viability sit in between. Change impact gets partial attention, usually a training plan, without the harder question of whether the resistance that was manageable with ten pilot users stays manageable with two hundred. Commercial viability gets the least honest scrutiny of all, because the business case that justified the pilot rarely gets rebuilt at production volumes. It gets extrapolated, which is a different exercise with a different answer.

What scaling too early actually looks like

It looks like a rollout that goes ahead on schedule and then costs two to three times the original estimate, taking three and a half times as long, with nobody able to point to a single moment where it went wrong.

A fashion retailer's product description tool cut description time by 90% in a pilot on 100 SKUs, with the buying team happy with the output. Scaling to 15,000 SKUs exposed inconsistent attribute data between categories, a token cost that scaled non-linearly with product complexity, and an integration requirement the pilot's manual upload process had never had to meet. The estimated four-week rollout took fourteen, at 2.3 times the original cost.

A grocery operator's replenishment tool worked well in five stores run by experienced managers who knew when to override a bad recommendation. Those stores had also been picked partly because their data was clean. Scaling to eighty stores surfaced data quality problems in 30% of them, incorrect shelf capacities, outdated dimensions, missing promotional flags, none of which the pilot had been designed to catch because the pilot never touched that data.

A hospitality group's pricing tool performed well across one brand's twenty sites on one POS system. The group's other two brands ran different POS systems and different pricing structures, and building those integrations took more engineering effort than the original tool build. The bottleneck sat in what the model had to plug into, which nobody had mapped because the pilot never had to.

The pilot ran fine in every one of these cases. The assessment between pilot and scale decision, the one built to check the four dimensions before committing rather than after, didn't happen.

The asymmetry steering groups get backwards

Scaling too late has a real cost: delayed benefit, a competitor getting there first, a team that loses momentum waiting for a sign-off that keeps slipping. You can usually recover from it, because the capability still exists, the pilot evidence is still valid, and the decision to scale can be revisited next month or next quarter without having burned anything down.

The cost of scaling too early compounds. It shows up as budget overrun, as a rebuild of an integration everyone thought was finished, as a steering group that stops trusting the programme's estimates because the last three came in over, as a capability that gets quietly parked because the operational team never had what it needed to run it and nobody wants to admit that in a steering meeting. Recovering from a premature scale decision costs more than the delay would have cost, and it costs credibility on top of the money.

A steering group under pressure to show progress weighs caution as the risk. The evidence says the opposite: the premature scale decision is the one that costs more, takes longer, and does more lasting damage to how much the next business case gets trusted.

Three outcomes, not one

A scalability assessment should produce one of three outcomes: scale, optimise, or kill. Most steering groups only prepare for the first, which is the problem. A team that walks into a review meeting having only modelled "yes, and here's the rollout plan" has already decided the answer before the assessment ran.

Scale means all four dimensions clear the bar, or clear it with a defined remediation plan attached to the specific gap. Optimise means the pilot proved the concept but at least one dimension, usually organisational readiness or the underlying data, needs work before scale makes sense, with an owner and a timeline attached. Kill means the pilot answered the question it was built to answer, and the answer is that this doesn't scale in its current form. That's a legitimate result. Treating it as a failure discourages the honesty the assessment depends on.

Presenting all three honestly to a steering group, with evidence attached to each dimension rather than a single overall impression, is what stops the meeting from defaulting to optimism. A steering group presented with a single number and a recommendation to proceed will usually proceed. A steering group presented with four dimensions, evidence at each, and a named binding constraint has something to actually govern.

What I'd ask before signing off a scale decision

I'd want the assessment run by a cross-functional team rather than the team that built the pilot, because the people who'll operate and support the capability see the gaps the builders have learned to work around. I'd want the binding constraint named specifically, not described as "some work still to do." I'd want the cost model built at production volume rather than extrapolated from pilot volume, because the experiment stage rarely surfaces the assumptions that fail at scale, and cost is one of them. And I'd want to know who owns the capability day to day once the experiment team moves on to the next thing, because that answer determines whether the rollout survives its first difficult month. If the plan involves a staged rollout, I'd want each stage gated by evidence, not waved through because the previous stage went well.

A pilot that worked is evidence for the steering group to weigh, not a decision already taken. Whether the rollout lands depends on the four dimensions the pilot never had to prove: technical capacity at production volume, an operational team that can run the capability without the people who built it, change management that holds at ten times the pilot's user count, and a cost model built at scale rather than extrapolated from it. Score all four honestly, name the binding constraint rather than smoothing over it, and the steering group gets to choose between scale, optimise, and kill with the evidence in front of them. Skip that step and the choice still gets made. It just gets made later, at higher cost, by the operational team that inherits the capability rather than the steering group that approved it.