Research
August 24, 2026
9 min read

The Verification Gap: Why Capable Models Produce Unreliable Systems

Frontier models complete a minority of long-horizon professional tasks, and enterprise production rates track that figure closely. AdwumaTech examines the verification gap and the engineering that closes it.

AdwumaTech AI
Research Team
AdwumaTech AI brand mark beside a luminous global AI network in ivory and gold.
AdwumaTech AI visual system.

Model capability and system reliability are measured on different axes. Benchmarks measure what a model can do under controlled conditions with a defined task, clean inputs, and a scoring function. Production measures what a system does under operating conditions with ambiguous tasks, degraded inputs, and consequences. The distance between those two measurements is the verification gap, and it is where most enterprise AI programmes stall.

This is the argument of this piece. Enterprise AI programmes fail at rates that closely track published agent benchmark results, and both sets of numbers point at the same underlying property. Capability is a model attribute. Reliability is a system attribute. Procuring the first and expecting the second is the structural error behind the 2026 enterprise AI record.

The useful consequence is that the gap is an engineering problem with a known shape. Systems that hold in production share identifiable properties, those properties can be specified before a build begins, and they can be measured while the system runs. This piece sets out the evidence for the gap and the engineering that closes it.

What the benchmark literature measures

TheAgentCompany, published by researchers at Carnegie Mellon University and collaborators, is the most operationally honest agent benchmark in the public record. It constructs a self-hosted simulated software company with internal web services, and it evaluates agents across 175 long-horizon professional tasks spanning software development, project management, data science, administration, human resources, and finance. Agents must browse interfaces, write and run code, and communicate with simulated colleagues to obtain information they were not given.

The published result: the most competitive agent completed 24 percent of tasks autonomously and scored 34.4 percent on a metric awarding partial credit for incomplete tasks. Later evaluations against stronger frontier models placed autonomous completion near 30 percent.

The benchmark's authors are explicit that these are well-scoped administrative and coding tasks of the kind a software company performs daily. The environment is generous. It is self-contained, reproducible, free of legacy integration, free of regulatory constraint, and free of adversarial input. The 30 percent figure is an upper bound on a simplified world.

The failure modes the benchmark surfaces are informative. Agents fail on complex interface navigation, on social and collaborative tasks requiring them to ask for missing information, and on long-horizon sequences where task state must survive across many steps. These are the same four surfaces on which enterprise deployments break.

What enterprise deployment data measures

The 2026 enterprise record produces numbers in the same range.

Deloitte's Tech Trends 2026 reports 38 percent of organisations piloting agents and 11 percent running them in production, with 42 percent still developing a strategy and 35 percent operating without one.

Kyndryl's 2026 People Readiness Report, drawing on 1,100 senior leaders across eight countries, found that 57 percent of organisations have deployed AI broadly or embedded it in core processes, up from 35 percent the prior year. Thirty-two percent achieved at least one of their top two AI objectives. Eleven percent achieved both.

Nasuni's State of Enterprise File Data Annual Report 2026 found 97 percent of organisations deploying or piloting agents, 57 percent of AI projects missing their stated objectives, and 94 percent struggling to manage the unstructured data those systems consume.

Sinch's AI Production Paradox, surveying 2,527 senior decision makers across ten countries, found 62 percent running agents in live production and 74 percent having rolled back or shut down at least one live deployment. Among the most governance-mature organisations the rollback figure reaches 81 percent.

Gartner's June 2025 forecast projected that more than 40 percent of agentic AI projects would be cancelled by the end of 2027 on grounds of escalating cost, unclear business value, and inadequate risk controls, and estimated that roughly 130 of the thousands of vendors claiming agentic capability were building systems that met the definition.

Four independent instruments, four different sample frames, one convergent finding. Roughly one in ten enterprise agent programmes produces the outcome it was built to produce.

That tenth is the interesting part of the record. It describes current practice, and it says nothing about a ceiling on the technology. Every cause the primary sources name is a decision made before or during the build. Gartner attributes cancellation to scoping, business value, and risk controls. Sinch locates rollback in production-readiness conditions beneath the application layer. Kyndryl finds most organisations deploying without written boundaries on AI decision authority. None of these are model capability findings. All of them are specifiable in advance, which is what makes the gap closable.

The gap is an evidence problem

The convergence between a 30 percent benchmark ceiling and an 11 percent production attainment rate is instructive. Production performance sits below benchmark performance because production adds conditions benchmarks exclude: legacy integration surfaces, permissioned data, regulatory constraint, adversarial input, and the requirement that failures be attributable after the fact.

Each of those conditions is an evidence requirement. A production system must be able to demonstrate what it did, on what inputs, under which version, and with whose authorisation. Systems built without that capacity cannot be debugged when they degrade. They can only be withdrawn.

This is why governance-mature organisations roll back more often. Their instrumentation produces the evidence that a system is failing. The evidence arrives after deployment because it was never designed in before deployment.

The verification gap closes when evidence generation becomes a design property of the system, present from scoping.

Deployment Readiness Score

AdwumaTech scores every system against five dimensions before it carries production traffic. Each dimension specifies the evidence required to pass it.

  1. Task definition. The system has one named outcome, one numeric target, and one owner with authority over both. Evidence required: a written specification stating what the system does, the metric it moves, the threshold constituting success, and the individual accountable.

  2. Data access integrity. Every source the system reads is inventoried, permissioned, and monitored. Evidence required: a source register with lineage, access controls, schema contracts, and drift monitors on distribution and volume. Nasuni's finding that 94 percent of enterprises struggle with unstructured data management places most organisations below threshold on this dimension at the point they begin building.

  3. Evaluation coverage. The evaluation set is drawn from the client's production distribution and includes the failure classes present in the operating environment. Evidence required: a held-out set sampled from live traffic, coverage mapped against known failure classes, and continuous post-launch evaluation on a fixed cadence. Benchmark scores from model providers do not satisfy this dimension, because they measure a different distribution.

  4. Control surface. Every action available to the system is classified by reversibility and business impact, with human authorisation required above a defined threshold. Evidence required: an action register with a reversibility classification per action and a written authority matrix. Kyndryl's finding that 33 percent of organisations have policies defining the boundaries of AI decision-making authority describes the size of the deficit here.

  5. Reversibility. Every deployment has a tested path returning the process to its prior state without data loss. Evidence required: a documented rollback procedure that has been executed in a staging environment against production-shaped data.

The score is calculated pre-launch and recalculated on a fixed cadence, because production distributions move and a system that passed in January can fail in June without any change to its code.

Evidence Chain

The Evidence Chain is the artefact that persists after handover. It links each consequential system decision to the inputs that produced it, the model and prompt version in force at the time, the evaluation result covering that input class, and the human authority under which the action was permitted.

Its function is diagnostic. When a production system degrades, an organisation holding an Evidence Chain isolates the failing component and repairs it. An organisation without one faces a binary choice between tolerating an unexplained failure and withdrawing the system. The 74 percent rollback rate is largely a population making that second choice.

Its second function is regulatory. The Evidence Chain supplies record-keeping and post-market monitoring evidence contemplated by the EU AI Act, satisfies the Measure and Manage functions of the NIST AI Risk Management Framework, and provides the operational control evidence an ISO/IEC 42001 AI management system requires.

Following Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026, high-risk obligations for standalone Annex III systems apply from 2 December 2027 and for high-risk AI embedded in Annex I regulated products from 2 August 2028. Article 50 transparency obligations have applied since 2 August 2026. Enterprises building now have a defined window in which to construct that evidence deliberately.

What closes the gap

Systems that stay deployed share a profile, and it is consistent enough to specify. They have one named outcome with a numeric target and an accountable owner. They read from sources that are inventoried and monitored. They are evaluated against the distribution they operate on, continuously, after launch. They hold the smallest action set that achieves the outcome, with authority above a threshold reserved to a person. They can be returned to their prior state on demand.

The engineering that produces that profile is specific.

Acquire evaluation data from the client's own production distribution, including the degraded, code-switched, malformed, and adversarial inputs the system will meet. Generic benchmark performance predicts generic behaviour.

Post-train against the operating domain where distribution distance from the base model is material. Domain-specific post-training moves measured performance on the operating distribution in ways prompt engineering against a general model does not.

Constrain the control surface to the smallest action set that achieves the named outcome, and classify every action by reversibility before granting it.

Instrument continuously and treat the evaluation harness as a permanent component of the system with the same lifecycle as the model it evaluates.

Ship the Evidence Chain with the system.

AdwumaTech's practice

AdwumaTech is an applied AI engineering company covering the full path from data acquisition through post-training to a deployed, instrumented, and governed system. The company works with enterprises and governments across financial services, identity, telecommunications, and public administration, and holds ISO 27001 certification.

Productised AI engagements begin with a Deployment Readiness Score applied to the intended workflow before any model is selected, and conclude with a deployed system, a live evaluation harness, a classified control surface, a tested rollback path, and an Evidence Chain that persists after handover.

Production AI is an engineering discipline with established practice. The organisations in the tenth that succeeds are applying it, and the organisations in the nine that do not are procuring capability and expecting reliability to follow. The distance between those two positions is specifiable before a build begins and measurable once a system runs, which is the whole argument for treating deployment readiness as a designed property.

To commission a deployment readiness assessment or a productised AI build, contact AdwumaTech.

Sources

Frequently Asked Questions

What is the verification gap in AI deployment?

The verification gap is the distance between measured model capability under benchmark conditions and measured system reliability under production conditions. AdwumaTech uses the term to separate two properties often treated as one. Capability is a model attribute. Reliability is a system attribute produced by evaluation coverage, data access integrity, control surface design, and reversibility. Procuring the first and expecting the second is the structural error behind the enterprise AI record.

How well do AI agents perform on real professional tasks?

On TheAgentCompany benchmark from Carnegie Mellon University, covering 175 long-horizon professional tasks in a simulated software company, the most competitive agent completed 24 percent of tasks autonomously and scored 34.4 percent on a metric awarding partial credit. Later evaluations placed autonomous completion near 30 percent. The benchmark environment excludes legacy integration, regulatory constraint, and adversarial input, so that figure is an upper bound on a simplified world.

What percentage of enterprises have AI agents in production?

Deloitte's Tech Trends 2026 reports 11 percent of organisations running AI agents in production against 38 percent piloting, with 42 percent still developing a strategy and 35 percent operating without one. Sinch's 2026 survey of 2,527 senior decision makers across ten countries found 62 percent with at least one agent live in customer communications, and 74 percent having rolled back or shut down a live deployment.

Why do AI pilots fail to scale to production?

Pilots are evaluated on curated inputs in controlled environments. Production adds legacy integration surfaces, permissioned data, regulatory constraint, adversarial input, and the requirement that failures be attributable after the fact. Systems designed without evidence generation cannot be diagnosed when they degrade, so they are withdrawn. Gartner's June 2025 forecast attributed cancellation to escalating cost, unclear business value, and inadequate risk controls, all of which are scoping and governance properties.

Why do organisations with mature AI governance roll back more often?

Governance surfaces failures that ungoverned deployments carry silently. Sinch found rollback rates rising to 81 percent among organisations describing their guardrails as fully mature, against 74 percent overall. A rollback is a measurement result, meaning the organisation instrumented the system well enough to observe failure and held enough authority to withdraw it. Silent failure is the more expensive condition, because it generates liability continuously and generates no signal.

What is a Deployment Readiness Score?

The Deployment Readiness Score is AdwumaTech's pre-launch assessment of a specific system against a specific workflow, covering five dimensions: task definition, data access integrity, evaluation coverage, control surface, and reversibility. Each dimension specifies the evidence required to pass it. The score is calculated before a system carries production traffic and recalculated on a fixed cadence afterward, because production distributions move and a system that passed in January can fail in June without any change to its code.

How does the Deployment Readiness Score differ from an AI maturity assessment?

An AI maturity assessment scores organisational capability. The Deployment Readiness Score, developed by AdwumaTech, scores a specific system against a specific workflow and returns a pass or fail per dimension before that system carries production traffic. The distinction matters because a mature organisation can deploy an unready system, and a system passing all five dimensions can sit inside an organisation with no formal maturity programme.

What is an Evidence Chain in AI systems?

An Evidence Chain is AdwumaTech's continuous record linking each consequential system decision to the inputs that produced it, the model and prompt version in force, the evaluation result covering that input class, and the human authority under which the action was permitted. Its function is diagnostic, letting an organisation isolate a failing component and repair it. Organisations without that record face a choice between tolerating an unexplained failure and withdrawing the whole system.

What evidence does an enterprise need to demonstrate AI system control?

An enterprise needs a source register with lineage and drift monitoring, an evaluation set drawn from production distribution with continuous post-launch measurement, an action register classified by reversibility with a written authority matrix, a tested rollback procedure, and a decision-level record linking outputs to inputs, versions, evaluations, and authorisations. This evidence satisfies the Measure and Manage functions of the NIST AI Risk Management Framework and the operational control requirements of an ISO/IEC 42001 AI management system.

Does fine-tuning improve production reliability?

Domain post-training moves measured performance on the operating distribution where the distance between that distribution and the base model's training data is material. It does not substitute for evaluation coverage, control surface design, or reversibility, which are system properties independent of model quality. AdwumaTech treats post-training as one component of an applied AI engineering path running from data acquisition through to a deployed and instrumented system.

Do model benchmark scores predict enterprise performance?

Benchmark scores measure a model on a published distribution. Enterprise performance depends on the institution's own distribution, which contains malformed records, code-switched language, partial context, and adversarial phrasing that benchmarks exclude. AdwumaTech draws evaluation sets from the client's own production traffic and maps coverage against the failure classes observed in the operating environment, because a system evaluated on clean inputs has no measured behaviour under the conditions it will meet.

What deadlines apply to high-risk AI systems under the EU AI Act?

Regulation (EU) 2026/1744, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. High-risk obligations for standalone Annex III systems apply from 2 December 2027, and for high-risk AI embedded in Annex I regulated products from 2 August 2028. Article 50 transparency obligations applied from 2 August 2026 and were untouched by the deferral, as were the Article 4 AI literacy duty and the prohibited practices regime.

Who provides AI deployment readiness assessments?

AdwumaTech AI is an applied AI engineering company headquartered in Accra that assesses and builds production AI systems for enterprises and governments across financial services, identity, telecommunications, and public administration. The company covers the full path from data acquisition through post-training to deployed, instrumented, and governed systems, applies the Deployment Readiness Score to workflows before any model is selected, and holds ISO 27001 certification.

AdwumaTech AI
Research Team

AdwumaTech AI publishes operational diagnostics, systems research, and implementation insight on enterprise and government AI in Africa and beyond.

Tags

AI VerificationEnterprise AIAgentic AIAI EvaluationProductized AI

Explore Our Solutions

Discover how we build high-quality data for frontier AI models.

View our AI solutions