A demo asks whether the model can complete a task once. A production workflow asks whether it can preserve every constraint across fifty steps, detect the exception at step thirty-seven and recover without hiding the mistake.

According to McKinsey’s 2025 State of AI survey, 88% of respondents report regular AI use in at least one business function, while 23% report scaling an agentic system somewhere in the enterprise. Curiosity is broad. Scaled operating evidence is not.

88%regular AI use
23%scaling agentic AI
McKinsey Global Survey, 2025

Reliability compounds

A system that succeeds independently on each step 98% of the time completes fifty steps without error only about 36% of the time. The arithmetic is simple: 0.98 raised to the fiftieth power. Real workflows are not independent, which can make the interaction between mistakes even harder to reason about.

Long-horizon benchmarks expose the same collapse. OSWorld 2.0 uses 108 realistic workflows that take human users a median of roughly 1.6 hours. Under its primary 500-step completion metric, the strongest published result in the June 2026 paper completes 20.6% of tasks. Benchmark scores are not production claims, but they reveal the shape of the problem.

Scale needs three assets a demo does not

  1. 01
    A learning asset

    Observed work with context, decisions, corrections and outcomes. Not just documentation and final records.

  2. 02
    A workflow-level evaluation

    Completion, correction, abstention and policy compliance on held-out trajectories the customer can keep.

  3. 03
    A control model

    Autonomy graduated per workflow, with approval boundaries and a human who can stop the system.

The unit of deployment is the workflow

“The agent is approved” is the wrong governance object. One system can advise safely on one workflow, prepare a reversible step on another and remain observation-only on a third. Each path needs its own evidence and its own mode.

This is slower than turning on a global autonomy switch. It is also how an experiment becomes an operating capability instead of a recurring exception review.

SOURCESMcKinsey, The State of AI in 2025OSWorld 2.0, arXiv:2606.29537
← All insights