Digital HumanRESULTS · THE DILIGENCE LEDGER

The evidence.
In full context.

Every number says what it is: benchmark, live evaluation, internal evaluation or target. AI that works starts with claims that hold up.

Examine the evidence
ANALOG CIRCUIT OPTIMIZATIONREPORTED RESULT
9/10

held-out circuits won

Specialist-led system with frontier fallback. Equal iteration budgets. Wins rank targets met first, then external tokens. Internal evaluation.

01 / Engineering02 / Billing03 / Recorded work
Methods, results and limits. Together.

Evaluations at a glance

Ten unseen circuits. Nine wins.

A specialist-led system won on 9 of 10 held-out circuits against a frontier baseline, with equal iteration budgets.

Internal evaluation · specialist with frontier fallbackInside the evaluation
ANALOG CIRCUIT OPTIMIZATIONMEASURED
9/10held-out circuits won
Equal iteration budgets136 / 146 targets met

Specialist with frontier fallback. Targets met determine wins; external tokens break ties. Internal evaluation.

First, establish the baseline.

Before specialist training, the base model achieved 50% procedural accuracy on the enterprise process evaluation.

Same model · enterprise process knowledge
ENTERPRISE PROCESS KNOWLEDGEMEASURED
50%before specialist training
The same untrained base model52 questions · 156 blind judgments

Enterprise billing process. Internally designed, LLM-judged evaluation. Not insurance execution accuracy. Internal evaluation.

The same model. A deeper understanding.

Procedural accuracy reached 88% after training: an improvement of 38 percentage points over the 50% baseline.

52 questions · 156 blind LLM judgmentsRead the method
ENTERPRISE PROCESS KNOWLEDGEMEASURED
88%after specialist training
+38 percentage points52 questions · 156 blind judgments

Enterprise billing process. Internally designed, LLM-judged evaluation. Not insurance execution accuracy. Internal evaluation.

ANALOG CIRCUIT OPTIMIZATIONMEASURED
9/10held-out circuits won
Equal iteration budgets136 / 146 targets met

Specialist with frontier fallback. Targets met determine wins; external tokens break ties. Internal evaluation.

How to read these evaluations

What task was tested? What counted as success? What was it compared against? Who designed and judged the test? What conclusion would go beyond the evidence? The methods and scope below answer those questions.

DESIGN PARTNER BENCHMARK

Run with a named partner engagement. Internal and not independently audited.

SPECIALIST ENGINEERING

9 of 10 against the frontier

Method and context

An analog circuit arrives as a specification table whose targets trade off against one another. The model learns from replayable expert reasoning traces, then re-solves the circuits under a deterministic manufacturability check and real simulation.

The August 2026 internal evaluation used ten held-out circuits: six op-amps and four LDO regulators, with equal iteration budgets. The specialist-led system called a GPT-5.5 frontier fallback when it stalled; the baseline used GPT-5.5 throughout. Wins were ranked by targets met, then external tokens as a tiebreaker. Four circuits used no external calls.

9 / 10 circuits · 136 / 146 targets · 6 vs 3 full-spec
Bounded analog-circuit optimization. No tape-out, yield or production-readiness claim.
LIVE EVALUATION

Unscripted probes against a live model.

INSURANCE · RECORDED Q&A

The reasoning, open to inspection

Method and context

The live questions examine source precedence, missing fields and dependencies between linked claim records. The product walkthrough recreates one recorded source-conflict question and separates it from rules taken directly from the workflow documentation.

The transcript includes assumptions and imperfect answers. Individual responses should be checked against the documented process; the number of probes is not a success rate.

Explore the source-backed questions
8 recorded probes · 2 rounds · May 2026
A workflow-reasoning transcript, not an accuracy score or proof of completed execution.
RECORDED DEMO

Behavior shown in a recorded application session. This is not a production deployment or an accuracy benchmark.

INSURANCE · APP EXECUTION

From incoming email to recorded outcome

Method and context

The recorded run extracts an attached claim form, checks policy dates, registers a claim, handles an added instruction to create a client note, and returns to email for confirmation and marking the request as read.

The silent highlights below keep the application and actual reasoning visible in the same frame. Selected sequences are accelerated. It is a demonstration of app behavior, not a completion-rate benchmark.

54.5 seconds · registration, context and communication
Silent highlights of the original synthetic-data app recording. The claim is registered and under review, not approved or paid.
BLIND EVALUATION

Model identities were hidden from the judge in an internally designed, LLM-judged evaluation.

CASE TO INVOICE · ANONYMIZED

A blind, 52-question evaluation

Method and context

A compact model was fine-tuned on one enterprise billing process, then compared with a frontier chatbot and the untrained base model. The documented process was the sole authority.

Across 52 questions and 156 blind judgments, procedural accuracy was 88% for the trained specialist, 38% for the frontier chatbot and 50% for the same untrained base model. Factual accuracy was 97% / 91% / 81%; grounded relevance was 90% / 85% / 80%, in that same order. The questions comprised 16 baseline, 21 generalization and 15 trap prompts. Metric-specific denominators were not supplied, so the percentages should not be converted into counts out of 52. The same-base comparison isolates the improvement from process training more directly.

88% procedural accuracy · 38% frontier · 50% base model
Internally designed and LLM-judged, with model identities hidden. No production-execution claim.
PRODUCT REQUIREMENT

A requirement the product is being built to satisfy.

MEDICAL DEVICE ORDER OPERATIONS

Cloning the order-entry specialist

Method and context

The purchase order can contain the visible fields while omitting the contextual account decision. The twin is designed to learn that decision from observed work and held-out historical orders.

It is designed to abstain with reasoning attached and remain human-approved until sustained per-customer accuracy is proven.

Bounded proof structure · target locked before validation
A proof structure or engagement status, not a completed result.
FROM THE EVALUATION / ENTERPRISE BILLING

The difference is
knowing this workflow.

Three questions from the evaluation.
Follow the rule, the decision and the difference in each answer.

01 / Work already done

Should valid work be undone?

A biller has created the debit memo before recording the billability verdict. Must the work be reversed?

THE DOCUMENTED RULE

This process allows charge entry and memo creation before the verdict is recorded.

Adapted from evaluation question Q3. The source process, rather than a general best-practice assumption, determines the answer.

THE SPECIALIST’S ANSWER

Keep the valid work.

  1. Charge entry
  2. Memo created
  3. Record verdict

Permitted sequence · no reversal needed

Keep the valid work. Record the required verdict without reversing and recreating the memo.

THE SAME QUESTION. TWO OTHER ANSWERS.
FRONTIER CHATBOT

Incorrectly treats the permitted sequence as a violation.

UNTRAINED BASE MODEL

Also calls it a violation and adds an unsupported reversal procedure.

02 / A missing note

Missing does not always mean wrong.

An invoice row has completion details but no billability note. Does that prove a required check was skipped?

THE DOCUMENTED RULE

The fixed/maintenance contract path skips that note by design. Other contract paths still require their checks.

Adapted from evaluation question Q7B. This is a specific contract exception, not a rule that missing notes are harmless.

THE SPECIALIST’S ANSWER

Know which branch applies.

IFFixed / maintenance

Note skipped by design

OTHERWISEOther contract paths

Follow the required checks

Identifies the documented contract exception and checks which branch applies.

THE SAME QUESTION. TWO OTHER ANSWERS.
FRONTIER CHATBOT

Recognizes a possible exception, but does not identify the documented contract branch.

UNTRAINED BASE MODEL

Invents a note-overwrite explanation and an unsupported lookup.

03 / Approval pending

A price discrepancy stays pending.

The verbal quote is $450, while the signed purchase order says $380. What must happen before $450 can be invoiced?

THE DOCUMENTED RULE

Verify the disputed rate against the purchase order, quote the corrected amount, wait for customer approval, and record the awaiting-approval state.

Adapted from evaluation question Q10. The amounts are test-prompt values. No recovered revenue or production transaction is claimed.

THE SPECIALIST’S ANSWER

Hold the higher amount.

SIGNED PURCHASE ORDER$380
HIGHER AMOUNT$450

Written approval required before invoicing $450

Requires written approval tied to the case before invoicing $450; holds the higher amount and records the waiting state.

THE SAME QUESTION. TWO OTHER ANSWERS.
FRONTIER CHATBOT

Gives the authorization principle, but omits the concrete hold-state record.

UNTRAINED BASE MODEL

Adds compliance and system gates that are not in the documented process.

THE SPECIALIST’S ANSWER

Keep the valid work.

  1. Charge entry
  2. Memo created
  3. Record verdict

Permitted sequence · no reversal needed

Keep the valid work. Record the required verdict without reversing and recreating the memo.

THE SAME QUESTION. TWO OTHER ANSWERS.
FRONTIER CHATBOT

Incorrectly treats the permitted sequence as a violation.

UNTRAINED BASE MODEL

Also calls it a violation and adds an unsupported reversal procedure.

Internal, blind, LLM-judged evaluation. Responses are summarized and anonymized; this tested process reasoning, not production execution.

SEE THE APP AT WORK

Real applications.
Real execution.

See the application and the model’s reasoning together: policy checks, a form filled, an added task and the final confirmation.

ACTUAL APP RECORDING00:55 / SILENT HIGHLIGHTS
01 / Verify coverage

The policy coverage is checked before the claim is registered.

A silent 54.5-second highlights cut using synthetic demo data. Selected sequences play at 2× speed; form recovery and the expanded reasoning have time to read. The full application and reasoning panel remain visible together. Pause or use the chapters to inspect a moment. The claim is registered and under review, not approved or paid.

Read the scene overview
  1. Verify coverage. The policy coverage is checked before the claim is registered.
  2. Fill the form. The application fills with claim details while the actual reasoning trail remains alongside it.
  3. Recover the amount. An edited amount is corrected back to 2,500. Watch the recovery and the remaining fields.
  4. Add a task. The registered claim appears with New status. A client-note task is added.
  5. Write the note. The client note and expanded reasoning show the work and its explanation in the same frame.
  6. Confirm the outcome. The reasoning records completion. The confirmation email keeps the claim in New status for review.
A DIFFERENT COST STRUCTURE

More work.
Less dependence on frontier APIs.

72%

Lower external token use

Company-reported reduction in external billed tokens in the specialist engineering comparison. The benchmark retained a frontier fallback for measurable stalls; this is not a zero-external-call claim.

16×

Lower cost

Neurologic’s comparison includes build costs, hardware for execution and ongoing management on our side, against frontier API token costs for the compared approach.

For every 100 external billed tokens in the compared approach, the specialist approach used 28.

Explore each step
COMPANY COMPARISON

Neurologic-reported, workload-specific comparison. Not independently audited. Supporting scope and assumptions available during evaluation.

The two figures measure different things. They are company-reported, workload-specific comparisons, not universal savings guarantees. Request the workload, accounting period, usage assumptions and supporting cost breakdown during your evaluation.

Request comparison details ↗
THE NEXT EVIDENCE

The bar keeps rising, in public.

01

Independent review

Human review with disclosed rubrics on top of the internal evaluations.

02

Customer-held validation

The validation design keeps the held-out set with the partner and computes the result inside its environment.

03

Longitudinal evidence

Correction rate, abstention quality, drift and overrides over time.

04

Verified controls

Proof along the full path from evidence to action, not just first-pass accuracy.

TARGETS, IN THE OPEN

Goals, not guarantees.

90%+

action accuracy

TARGET

A declared bar, not current performance.

85%+

error self-correction

TARGET

A declared bar, not current performance.

90%+

autonomous task completion

TARGET

A declared bar, not current performance.

TECHNICAL DILIGENCE

Ask for the packs behind the labels.

Qualified partners and investors can review the complete methods under NDA.