Five people complete the same task in 5, 6, 7, 20 and 22 minutes. Their mean is 12. Nobody works that way.

Yet the default recipe for learning from enterprise work treats every demonstration as equally representative. Collect the traces, mix them together and optimize for the behavior in the middle. The result is statistically tidy and operationally wrong. It reproduces the average employee, including avoidable rework, slow paths and the habits senior operators learned to stop years ago.

The average is not a person

Performance is rarely one-dimensional. The fastest employee may skip a compliance step. The most accurate person may handle only easy work. This is why the answer is not “copy the fastest.” It is a two-stage system: apply a non-negotiable compliance gate, then weight the surviving demonstrations by measured efficiency and outcome quality.

THE ORDER MATTERSCompliance gate first. Expertise weighting second. Customer control throughout.

The customer holds the dial between the broadly representative workflow and the best measured workflow. What the customer cannot do is dial away the compliance floor. This separates expert weighting from simple performance surveillance. The object being evaluated is the trace, not the person.

Corrections contain more signal than finished records

A finished case record keeps the final answer and erases the path. When an expert reverses a decision, the execution trace preserves four things: what first looked plausible, what evidence changed the conclusion, how the workflow was repaired and what outcome confirmed the repair.

That transition is supervision. It teaches the model both the tempting mistake and the recovery. An archive of clean final states cannot do that. Observed work can.

Cloning still has a ceiling

Even perfect expert weighting only reaches the best behavior already present in the demonstrations. A clone can tie the best person it watched. It cannot discover an approach nobody tried.

ILLUSTRATIVE LEARNINGIMITATION CEILING
Behavior cloningVerified-outcome learning

Dashed line: best observed demonstration.

Behavior cloning approaches the best observed human. Reinforcement learning, rewarded only by verified outcomes and bounded by compliance rules, can cross that ceiling.

Learn the best compliant behavior already present in expert demonstrations.

Explore each step

That is the role of reinforcement learning after cloning. The model re-attempts the work, explores bounded variations and keeps only changes that produce a verified outcome. The reward is not “look human.” It is “pass the real check,” with compliance as the first term.

The practical implication

If an automation program has plateaued, more undifferentiated data may deepen the plateau. The better question is whether the data distinguishes compliant excellence from common behavior, preserves corrections and connects actions to outcomes.

The learning asset is not the document library. It is the history of good judgment in motion.

← All insights