sqwyz.
21 September 2026 · Written by Scott Robertson

Why Most AI Readiness Scores Are Wrong

A laptop screen showing a web analytics dashboard with load-time and bounce-rate charts

Most organisations score themselves too high on AI readiness, not from dishonesty but because self-assessment without structured challenge produces optimistic scores almost every time. The gap between a self-assessed score and a facilitated, evidence-based one is typically one to two points on a five-point scale. At that scale, it's not a rounding error. It's the difference between a programme designed for your actual starting point and one designed for a position you haven't reached yet.

The self-assessment problem

A readiness score only has value if the process producing it is designed to surface uncomfortable truths. Most self-assessments fail because they measure intent rather than evidence, anchor on the assessor's own domain rather than the programme's domain, and skip the challenge mechanism that pushes back on scores unsupported by facts.

The result is a leadership team that believes it's at a 3.5, builds a programme designed for that position, and discovers three months in that the actual starting point was closer to 2.3. That correction matters. At 3.5, you can start five parallel experiments across different business functions. At 2.3, you start one, with a parallel workstream to fix the foundations. These aren't similar programmes. One stalls. The other has a reasonable chance of building something that lasts.

What the six dimensions are measuring

A proper readiness assessment covers six dimensions, and each has its own failure mode in self-assessment.

Data quality and availability is where the largest gap sits. Teams conflate having a data warehouse with having data that's fit for AI. Having a warehouse means data has been consolidated somewhere. Having data fit for AI means it's accessible in the format AI tools need, clean enough to produce reliable outputs, and current enough to be relevant, and the second condition is far harder to achieve than the first.

Digital maturity gets scored on what the technology team already knows about, not on what an AI programme will actually need. Infrastructure that works day-to-day gets scored as stable, which is fair for daily operations and unfair for integrating AI tools that need clean API access, reliable data pipelines, and the ability to push changes without disrupting live systems.

Skills and capability tends to be scored accurately at the technical end and too generously at the organisational end. Most leadership teams are honest about whether they have data scientists. Fewer are honest about whether the wider organisation has the AI literacy to evaluate outputs critically, use tools effectively, and resist treating AI as a black box.

Culture and appetite for change is consistently the most optimistic dimension. The absence of open objection gets mistaken for genuine willingness to change how people work. Culture should be scored on evidence of past change, not stated openness to future change.

Leadership alignment often gets scored on whether leaders are interested in AI, not on whether they share a specific view of strategy, investment, and priority areas. Interest isn't alignment. A leadership team where the COO wants to automate operations, the CFO wants to cut costs, and the CMO wants to personalise the customer experience can all score a 4 on leadership alignment while holding three different programmes in their heads.

Governance and compliance readiness is typically scored by whoever owns risk or compliance, and their domain knowledge, financial and operational governance, is often strong. AI governance is different: data classification for AI-specific risk, bias monitoring, accuracy tracking, human oversight mechanisms. The two don't map cleanly onto each other.

The fashion retailer example

A fashion retailer with 200+ stores completed a self-assessment across the leadership team and arrived at an average score of 3.5. A facilitated session two weeks later brought that average down to 2.3. The biggest corrections were in data quality, self-assessed at 4 and facilitated down to 2, and culture, self-assessed at 4 and facilitated down to 2.5.

Data quality dropped because the facilitator asked the buying team to demonstrate extracting a clean, product-level sales dataset. The process took three people and four days. A self-assessed score of 4 implies data that's accessible to analysts with some effort, with known gaps documented and prioritised. Three people and four days is a 2, not a 4.

Culture dropped because the facilitator asked for three examples of successful process change in the past two years. The team struggled to name one that had fully embedded. A score of 4 on culture requires that experimentation is encouraged, learning from failure is normalised, and middle management is broadly supportive. A team that can't name a successful change that stuck isn't operating at a 4.

The corrected scores changed the programme design entirely. Instead of launching five parallel experiments, the retailer started with one experiment in an area with minimal data dependency, running a parallel data quality remediation workstream alongside it.

The data warehouse illusion

"We have a data warehouse" is the single most common justification for a high data quality score, and it's almost always incomplete. A data warehouse tells you someone made an effort to consolidate data from multiple systems into one place. It says nothing about the quality of that data, its currency, its coverage, or whether it's structured in a way AI tools can actually consume.

The test that surfaces the truth is simple: ask for a clean, analysis-ready dataset for a realistic AI use case, something specific that would be used in a first experiment. Ask how long it takes to produce, how many people it involves, and how much manual cleaning is needed once it arrives. If the answer involves a two-week turnaround and a spreadsheet that needs significant manual correction, the data quality score isn't a 4, whatever the warehouse contains.

The domain anchoring problem

Different members of a leadership team anchor their scores on different parts of the business, and those anchors rarely match where AI will actually be applied.

A multi-category retailer running a leadership self-assessment found a significant split in data quality scores. The CFO scored it a 4: financial data was clean, well-governed, audited quarterly, and accessible with minimal friction. The commercial director scored it a 2: product and customer data was fragmented across five systems, inconsistently formatted, partially manual, with nobody holding a clear view of how to extract it cleanly for analysis.

Both scores were correct for their respective domains. The problem was that the most promising AI use cases for this business sat in commercial operations: demand forecasting, range optimisation, markdown timing, customer segmentation, all of which depend on commercial and operational data, not financial data. The CFO's 4 was irrelevant to programme planning. The commercial director's 2 was the number that mattered, and the facilitated session was what surfaced it. Left unchallenged, the self-assessment average would have obscured the whole picture.

What honest assessment enables

A score produced through facilitated, evidence-based assessment is a planning tool. A score produced through unchallenged self-assessment is, at best, a starting point for a harder conversation. The AI Transformation Playbook at transformationplaybook.ai has a facilitation guide for running this kind of session properly, with the six dimensions and the evidence questions that go with each one.

What you do with a low score on one dimension depends on the programme you're planning. A low data quality score paired with an optimisation programme means starting with use cases that don't depend on clean, integrated data while fixing the foundations in parallel. The same score paired with a transformation ambition means resolving a sequencing problem before committing significant resources.

An honest readiness position costs a few weeks up front and saves you from building a programme that stalls on contact with reality.