A compact model was fine-tuned on one enterprise billing process, then compared with a frontier chatbot and the untrained base model. The documented process was the sole authority.
Across 52 questions and 156 blind judgments, procedural accuracy was 88% for the trained specialist, 38% for the frontier chatbot and 50% for the same untrained base model. Factual accuracy was 97% / 91% / 81%; grounded relevance was 90% / 85% / 80%, in that same order. The questions comprised 16 baseline, 21 generalization and 15 trap prompts. Metric-specific denominators were not supplied, so the percentages should not be converted into counts out of 52. The same-base comparison isolates the improvement from process training more directly.