Every credible assessment scores the same five pillars. The sixth — whether your entities actually run the same operating model — is the one that decides whether a pilot can ever become production. The questions are published below rather than gated.
An AI readiness assessment is a structured evaluation of whether an organization can put a specific AI use case into production and keep it running — scored across strategy, data, infrastructure, governance, talent, and operating-model consistency. It is not a maturity score. Its output is a decision: proceed, proceed under stated conditions, or not yet.
The distinction matters because a score is unfalsifiable and a decision is not. A readiness report that returns “62 out of 100, moderate maturity” cannot be wrong, and cannot be acted on. A readiness report that says “this use case can go to production once data ownership is assigned and the retention policy is written, and not before” can be checked against reality in ninety days. We write the second kind.
The five-pillar model — strategy, data, infrastructure, governance, talent — is now effectively a standard. It appears, with small wording changes, in the readiness tools published by the large platform vendors, the accounting firms and the systems integrators. There is nothing wrong with it. It is necessary and it is incomplete.
It is incomplete because every pillar is scored at the level of the enterprise, and almost no mid-market company is a single enterprise in the way the scoring assumes. A company that has made six acquisitions has six definitions of a customer record, four ways of calculating a margin and three service-desk tools. Scored at the top, its data pillar looks adequate: a warehouse exists, a governance council meets. Scored where the work happens, the same data cannot support a model that has to behave identically in every entity.
This is the usual explanation for a result that gets reported as a technology failure. The widely cited Gartner forecast that at least 30 percent of generative AI projects would be abandoned after proof of concept is not mainly a story about models underperforming. Pilots succeed inside one entity, on one clean extract, with one willing team. Production requires the same definitions everywhere. That requirement is an operating-model question, and the five pillars do not ask it.
The second failure is softer and more common. Most published assessments are lead-capture instruments: ten questions behind an email gate, a score, a call. The questions are withheld because the questions are the product. We take the opposite position. The questions are below, in full, and they are worth more to you than a score is.
THE FRAMEWORK
Each dimension is assessed on evidence rather than on self-report. The last line of each is the failure mode we see most often — what it looks like in month nine, when the pilot is over and the thing is supposed to be running.
Whether a named business outcome exists, owned by a named executive, with a baseline you could defend to an auditor. Not “we want to use AI” — which decision gets faster, which cost comes out, which revenue moves.
Evidence: a written use case with a before-figure. Failure mode: a tool in search of a problem, defended after the fact.
Quality, lineage, accessibility and ownership of the specific data the use case needs — not the estate in general. Who owns the field, how it is defined, how often it is wrong, and whether anyone is accountable when it is.
Evidence: a named owner per critical field, and a measured error rate. Failure mode: a clean pilot extract that cannot be reproduced monthly.
Compute, integration surface, identity and licensing — assessed at production volume and production cost, not at pilot volume. Including the licensing question almost nobody asks before signing: what the unit cost looks like at ten times the pilot's throughput.
Evidence: a modelled cost at projected volume. Failure mode: an economic case that inverts at scale.
Policy, approval path, human-review points, model inventory, retention, and an audit trail that exists before anyone asks for it. Built against the NIST AI Risk Management Framework and ISO/IEC 42001 rather than an in-house checklist, because a buyer, an auditor or an insurer will diligence against those and not against ours.
Evidence: a model inventory and a documented approval path. Failure mode: shadow use already in production, undocumented.
Whether the people who have to use the output can, and whether the people whose work it changes were told before it arrived. Also whether the work sits with staff already carrying a full operating role, which is the most reliable predictor of a stall.
Evidence: named capacity, not named enthusiasm. Failure mode: adoption that reverts the first busy week.
Whether the entities, business units or acquired companies in scope run the same process, the same data definitions and the same controls — closely enough that one model can behave the same way in all of them.
Evidence: one process definition and one data dictionary per entity in scope, compared. Failure mode: a pilot that works in one unit and cannot be replicated in the others.
THE SIXTH DIMENSION
Here is the specific thing a five-pillar score cannot see. A company runs a document-classification pilot in its largest business unit. The data is good, the team is willing, the result is a sixty percent reduction in handling time. The board approves the rollout.
Then it meets the other five units. One classifies the same document type under a different taxonomy inherited from an acquisition. One holds the documents in a system with no API. One has a retention policy that forbids the processing the model performs. None of this was visible in an enterprise-level readiness score, because at the enterprise level there is a data warehouse, a governance council and a cloud contract, and all three scored well.
The honest assessment in that case does not say the company is unready for AI. It says the company is ready for AI in one unit and will be ready across all six once the taxonomy is reconciled and the retention policy is rewritten — and it says which of those two has to happen first, and what it costs to do them in the wrong order.
That is why we assess this dimension at the entity level rather than the enterprise level, and why we are willing to write not yet as a conclusion. An assessment that cannot return not yet is a sales document.
PUBLISHED, NOT GATED
These are the questions we ask. Take them and run the assessment yourself if that is what you need — the order matters more than the scoring, because the order is the order the work has to be sequenced in. Answer honestly and the gaps will be obvious without anyone scoring them.
Name one decision or task that would be measurably better if this worked, and the figure it is at today.
Which named executive loses something if it does not work?
For the data this use case needs, who owns each critical field — by name, not by department?
How often is that data wrong, and how do you know — measured, or assumed?
Across the entities in scope, is the process this touches defined once, or once per entity?
If the pilot succeeds in one unit, what specifically stops it being repeated in the others?
What is already in use that nobody approved, and would you find it if you looked?
Who approves a model going live, who reviews its output, and where is that written down?
At ten times the pilot's volume, what is the unit cost — and does the business case survive it?
How much of this work sits with people who also carry a full operating role?
If a customer, an acquirer or an auditor asked today how your models are governed, what would you hand them?
Twelve months out, what has to be measurably different for this to have been worth doing?
If you would rather answer them with a response at the end of it, the full diagnostic instrument — twelve questions, five to ten minutes — is on the approach page. Partial answers are still useful; the gaps tell us as much as the answers do.
Four outcomes rather than a number out of a hundred. Each one carries the condition attached to it, and the condition is what gets agreed in writing.
A precondition is missing that no amount of pilot effort substitutes for — usually data ownership, or a process that is defined differently in every entity. The report names the precondition and what it takes to meet it. Nothing is scoped until it is met.
The use case can proceed, in a stated scope, provided specific conditions hold — each with a named owner who has accepted it. This is the most common honest answer.
One use case, one scope, in production rather than in pilot, with governance in place before launch instead of after it. The point of phase one is to earn the right to a phase two, with evidence.
The structural layer holds under production volume across every entity in scope, the economics have reached run-rate rather than forecast, and the thing survives the people who built it leaving.
Most published maturity models describe the same five stages. They are useful for locating yourself, and useless for deciding what to do next — which is why the assessment above returns a condition rather than a stage.
The jump from three to four is the only one that is hard, and it is hard for operating-model reasons rather than technical ones.
A written assessment, two to four weeks for a single business unit and six to twelve across multiple entities, delivered as a document you can hand to a board rather than a dashboard login.
Governance findings are written against published standards — the NIST AI Risk Management Framework and ISO/IEC 42001 — rather than a proprietary maturity grid, because that is what a buyer, an auditor or an insurer will diligence against. Process and data findings reference APQC’s Process Classification Framework; control findings reference ISO 27001 and the NIST Cybersecurity Framework. Where a standard exists, we use it instead of restating it.
Scope and fees are set on a fit call, once the shape of the problem is clear.
WHERE THIS SITS
An AI readiness assessment is not a standalone product here. It is the first stage of Genexis ARC, our four-stage method that runs from readiness to run-rate: Readiness, then Sequence, then Spine, then Run-Rate. Each stage has a condition that has to be met in writing before the next one begins.
That structure is the reason the assessment can return not yet without ending the conversation. “Not yet” is a statement about stage one, and stage one has a defined way out of it. Most transformation programs are not undone by the wrong strategy — they are undone by the right work in the wrong order.
A structured evaluation of whether an organization can put a specific AI use case into production and keep it running, scored across strategy, data, infrastructure, governance, talent and operating-model consistency. A credible one returns a decision with conditions attached, not a maturity score.
Strategy, data, infrastructure, governance and talent. That five-pillar model is now effectively standard across published assessment tools. We score a sixth — operating-model consistency across entities — because a company that has grown by acquisition can pass all five at enterprise level and still be unable to move a pilot into production.
Unaware, experimenting, piloting, operating and scaled. The only genuinely difficult transition is from piloting to operating, and it is usually blocked by inconsistent process and data definitions rather than by technology.
It usually refers to the widely cited Gartner forecast that at least 30 percent of generative AI projects would be abandoned after proof of concept. The common causes given are poor data quality, inadequate risk controls, escalating cost and unclear business value — which in practice tend to be symptoms of the same thing: a pilot scoped where conditions were favourable and a production rollout that met the rest of the company.
Two to four weeks for a single business unit, six to twelve weeks across multiple entities. A one-hour self-assessment is worth running first; it will not tell you what to do, but it will tell you whether the expensive version is warranted.
Yes, and the twelve questions above are published so you can. An outside party is worth paying for in one specific case: when the answer is likely to be unwelcome, and the person who has to deliver it reports to someone who sponsored the pilot.
A data strategy describes an intended end state. A readiness assessment measures the current one against a specific use case. The two routinely disagree, and the disagreement is the useful part.
No. Readiness asks whether you can deploy; governance is one of the six dimensions it scores, and the one most often addressed after launch rather than before it. AI readiness and governance consulting covers both, including the governance build itself.
Answer the twelve questions and we will come back with a written read within one business day. No score, no dashboard.
Take the readiness diagnostic