VERAHELM HOLDINGS LLC

SYNTHETIC DECISION RECORD / VH-DR-0013

PUBLIC SAMPLE01 / 12

INDEPENDENT AI SYSTEM EVALUATION / SYNTHETIC DEMONSTRATION

DECISION
RECORD 0013

Candidate may proceed under the quality gate in this synthetic scope. Runtime remains unresolved; no speed or savings claim is authorized.

STATUS
CONDITIONAL
REVIEW STATE
INDEPENDENTLY RECOMPUTED
WORKLOAD
30 SYNTHETIC CASES
CUSTOMER DATA
NONE
DOCUMENT CLASS
PUBLIC SYNTHETIC
APPROVAL STATE
EXAMPLE / UNSIGNED
FIELD VALIDATION
NOT PERFORMED
SYNTHETIC EVIDENCE BOUNDARY

This dossier demonstrates the depth, traceability, and restraint of a Verahelm delivery. It is not a customer result, certification, security review, legal opinion, or promised operational or financial outcome.

00

DOSSIER MAP

QUESTION → EVIDENCE → DECISION
01

Decision contract

Question, comparators, primary gate, exclusions, and stop rules.

02

Workload audit

Fixture composition, class balance, authority labels, and evidence boundary.

03

Result matrix

Field, case, risk, segment, runtime, and paired-comparison results.

04

Failure accounting

Complete candidate error inventory, regressions, and critical handling defect.

05

Verification

Controls, ablation, replay, independent checks, and uncertainty.

06

Handoff

Risks, owners, stage gates, approval state, pilot design, and next discriminating test.

RECORD IDVH-DR-0013

Stable example identifier for discussion and review.

EVIDENCE CLASSSYNTHETIC / CONTROLLED

No customer or field evidence is represented.

DECISION OWNERBUYER-DESIGNATED

A real engagement names the person authorized to decide.

SUPERSESSION RULENEW AUTHORIZED EVIDENCE

Material input, version, environment, or control changes require reissue.

All organizations, tickets, outputs, and measurements in this public dossier are fictional or synthetic.

01

OPERATING DECISION

BOUNDARY / EXPLICIT

THE QUESTION

Does the quality evidence support controlled progression?

RELEASE INTERPRETATION CONDITIONAL

Quality observations favor the candidate. Runtime is slower, so a speed claim is withheld and the limitation remains open.

DECISION

Whether the guarded candidate merits a controlled, recommendation-only customer benchmark.

COMPARATORS

Frozen straightforward baseline, guarded candidate, and differentiator ablation on identical records.

PRIMARY GATE

Correctness first; high-consequence handling controls the release interpretation.

STOP RULE

No autonomous disposition while any uncontrolled high-risk route or action error remains.

DECISION USESTATUSACCOUNTABLE OWNERRELEASE CONDITION
Recommendation-only customer benchmarkCONDITIONALLY AUTHORIZEDBuyer decision ownerSigned pilot plan, frozen gates, named reviewers, and rollback route
Autonomous dispositionWITHHELDSafety and governance ownerSeparate approval after zero uncontrolled high-consequence errors
External quality claimWITHHELDEvidence ownerBuyer-blinded replication on an authorized holdout
Speed, savings, or return claimWITHHELDPlatform and commercial ownersProduction runtime, full cost, reviewer burden, and accepted-value evidence
02

WORKLOAD AND LABEL CONTRACT

30 RECORDS / 120 SCORED FIELDS

SYNTHETIC FIXTURE

Balanced for diagnosis, not production prevalence.

Each ticket has four required outputs: route, risk, action, and review flag. A case is exact only when all four fields match the frozen gold record.

ROUTE BALANCE
Engineering
06
Maintenance
06
Procurement
06
Quality
06
Safety
06
RISK BALANCE
High
09
Medium
13
Low
08

Rare high-consequence cases are intentionally overrepresented for the demonstration.

REVIEW AUTHORITY
Review required
11
No review
19

Review correctness is measured separately from route, risk, and action.

ROUTE

Correct accountable function receives the record.

RISK

Consequence class matches the gold boundary.

ACTION

Required operational response is preserved.

REVIEW

Human authority is invoked when the contract requires it.

03

RESULT MATRIX

BASELINE / CANDIDATE
MEASUREBASELINECANDIDATEINTERPRETATION
Route accuracy83.3%86.7%+3.3 PP
Risk accuracy86.7%96.7%+10.0 PP
Action accuracy73.3%90.0%+16.7 PP
Review accuracy80.0%100.0%+20.0 PP
Overall field accuracy80.8%93.3%+12.5 PP
Exact cases20 / 3025 / 30DESCRIPTIVE SUPPORT
High-risk recall06 / 0909 / 09DETECTION SUPPORT
High-risk fully correct06 / 0908 / 091 HANDLING DEFECT
Total field errors2308−15
RuntimeREFERENCE≈1.32× TIMESPEED CLAIM WITHHELD

Averages do not control the release state. The remaining high-risk route/action defect keeps the recommendation conditional.

04

SEGMENT AND REGRESSION AUDIT

AGGREGATES DO NOT HIDE FAILURES

BY GOLD RISK

RISKNOVERALLACTION
High0994.4%88.9%
Medium1396.2%92.3%
Low0887.5%87.5%

Low-risk errors can still create unnecessary review load and alert fatigue.

BY GOLD ROUTE

ROUTENOVERALLACTION
Engineering0687.5%83.3%
Maintenance06100.0%100.0%
Procurement06100.0%100.0%
Quality0691.7%83.3%
Safety0687.5%83.3%

Six records per route are diagnostic, not a stable production estimate.

IMPROVED 09REGRESSED 04BOTH CORRECT 16BOTH WRONG 01
05

COMPLETE CANDIDATE ERROR INVENTORY

5 CASES / 8 FIELD ERRORS
CASEOBSERVED ERRORSEVERITYOPERATING EFFECT
T006Quality/document routed to engineering/drawing review.MODERATE + MAJOR

Unnecessary engineering work and incorrect action.

T007Engineering/drawing review routed to quality/document.MODERATE + MAJOR

Drawing conflict may bypass engineering ownership.

T017Low risk classified as medium.MAJOR

Excess review and alert burden.

T020Safety observation routed to quality.MODERATE

Ownership error despite a low-risk label.

T025Safety/stop-line routed to quality/quarantine.CRITICAL

High risk detected; required authority and action were wrong.

Perfect high-risk detection is not perfect high-risk handling. Case T025 controls the autonomy boundary.

06

EVIDENCE CHAIN

CHECKS / CONTROLS
01

Workload frozen

Comparison rules and the synthetic case set remained fixed for the evaluated result.

PASS
02

Controls executed

Positive, negative, coverage, determinism, and ablation controls passed within the demonstration.

PASS
03

Independent verification

A separately implemented score recomputation and a manual record-level review agreed with the stated synthetic result.

PASS
04

Runs interleaved

Twenty-one comparison runs were interleaved across three synthetic workload sizes.

PASS
05

Ablation retained

The ablation scored between baseline and candidate; contribution is supported only inside this fixture.

89.2%
06

Contrary evidence kept

Four regressions, one critical action error, and slower runtime remain in the decision record.

OPEN
PROPOSED CLAIMCONTROLLING EVIDENCESTATUSPERMITTED LANGUAGE
Candidate improves exact-case results25/30 candidate versus 20/30 baseline on matched synthetic casesDESCRIPTIVE SUPPORTQuality signal inside this fixture
Candidate is safe for autonomous dispositionCritical route/action defect T025NOT SUPPORTEDRecommendation-only progression
Candidate is fasterCandidate required approximately 1.32× baseline timeCONTRADICTEDNo speed or throughput claim
Candidate creates financial valueProduction spend, review burden, and accepted-result value not measuredNOT EVALUATEDNo savings or return claim
CONTROLLED ARTIFACTSTATEEXAMPLE IDINTEGRITY CONTROLINVALIDATION RULE
Evaluation and claim contractFROZENVH-EC-0013Content digest recorded in a real deliveryAny material question, gate, exclusion, or claim change requires reissue
Authorized fixture and label inventoryFROZENVH-FX-0030Record count, schema, label, and source digestsCount, schema, authority, or content mismatch stops scoring
Baseline and candidate manifestsRECORDEDVH-SM-B / VH-SM-CVersion, configuration, dependency, and execution digestsAn unpinned or changed component is a new system under test
Interleaved execution recordRECORDEDVH-RUN-021Order, environment, seed, timestamp, and outcome logMissing or malformed runs are excluded and disclosed
Independent score recomputationPASSEDVH-IV-0013Separate implementation plus result digest comparisonAny disagreement reopens the result and blocks authorization

Independent agreement inside a synthetic test is not customer-blinded production validation.

07

UNCERTAINTY AND PAIRED COMPARISON

DESCRIPTIVE ≠ POPULATION PROOF
EXACT-CASE ACCURACY83.3%

Candidate: 25/30 exact cases. Simple 95% Wilson interval: 66.4%–92.7%.

PAIRED DIRECTION9 : 4

Nine candidate improvements versus four regressions on matched cases.

MCNEMARp = 0.267

The paired difference is not significant at α=0.05 on this 30-case fixture. Values are reported to three decimal places.

SYSTEMEXACTRATE95% WILSON INTERVAL
Baseline20 / 3066.7%48.8%–80.8%
Ablation22 / 3073.3%55.6%–85.8%
Candidate25 / 3083.3%66.4%–92.7%

These intervals do not account for distribution shift, label error, dependence between records, or customer-specific workflow conditions.

08

RUNTIME AND VALUE BOUNDARY

NO SPEEDUP / NO SAVINGS CLAIM
RECORDSBASELINE MEDIANCANDIDATE MEDIANRELATIVE TIME
300.192 ms0.254 ms1.32×
3001.828 ms2.424 ms1.33×
3,00017.862 ms23.621 ms1.32×

The candidate is slower in this local proxy. Quality evidence does not authorize a runtime, throughput, cost, or return-on-investment claim.

09

LIMITATION REGISTER

OPEN / RETAINED
  1. L-01

    Synthetic scope

    The workload contains no customer data and does not reproduce a live operating environment.

    OPEN
  2. L-02

    Runtime regression

    The higher-quality candidate is slower than the simple baseline; no speed claim is made.

    OPEN
  3. L-03

    No external validation

    No buyer-controlled blinded holdout or field deployment has tested transfer beyond this demonstration.

    OPEN
  4. L-05

    No promised value

    Any financial illustration would be a planning assumption, not a savings or return guarantee.

    OPEN
RISKTRIGGER / EVIDENCESEVERITYCONTROL / NEXT ACTIONRESIDUAL STATEACCOUNTABLE OWNER
R-01Critical route/action defect T025CRITICALRecommendation-only use, human authority, repair, and blinded replayOPEN / BLOCKINGSafety and governance
R-02No buyer-held external validationHIGHAuthorized holdout, segment audit, shift analysis, and predefined stop rulesOPENBuyer evidence owner
R-03Candidate runtime is approximately 1.32× baselineMODERATEProfile production service time, queueing, retries, and full reviewer burdenOPENPlatform owner
R-04No buyer-authoritative labels or adjudication historyHIGHDual review, documented authority, disagreement register, and sealed resolutionNOT ASSESSEDBuyer subject-matter owner
R-05Customer data, privacy, and security controls not exercisedHIGHSeparate data plan, threat review, access path, retention rule, and incident dutiesOUT OF SCOPECustomer security owner
10

CUSTOMER-BLINDED PILOT DESIGN

PROPOSED / NOT YET EXECUTED

NEXT EVIDENCE LEVEL

Use 100–300 authorized records and a buyer-held holdout.

The buyer retains the de-identified blind set until decision rules, candidate configuration, metrics, exclusions, and stop conditions are frozen.

INPUT DESIGN
  • Stratify by route, risk, source, age, and exception type.
  • Report rare high-consequence oversampling separately.
  • Use dual review for disputed or high-risk gold labels.
REQUIRED METRICS
  • Field and exact-case accuracy.
  • High-risk route, action, risk, and review recall separately.
  • Override, disagreement, runtime, spend, retry, and reviewer burden.
STOP CONDITIONS
  • Any uncontrolled high-risk action error.
  • Unresolved gold-label disagreement or material segment regression.
  • Failed replay, missed value threshold, or changed evidence contract.
STAGEEXIT REQUIREMENTSTATUS
Synthetic demonstrationFrozen fixture, controls, independent verificationPASSED
Customer benchmarkSegment and exact-case evidence on authorized recordsNOT STARTED
Recommendation-only pilotSealed holdout, owners, rollback, measured burdenNOT STARTED
Narrow low-risk automationMonitored, reversible production sliceNOT AUTHORIZED
High-consequence autonomySeparate qualified safety and governance approvalPROHIBITED INITIALLY
11

DECISION HANDOFF

NEXT TEST / MINIMUM

NEXT DISCRIMINATING TEST

Repeat the frozen gates on a buyer-blinded holdout.

Use authorized, de-identified buyer records. Retain the baseline, record distribution shifts and exceptions, and authorize only the claims that survive the new evidence boundary.

EXAMPLE DELIVERY PACKAGE
  • 01 / FROZEN EVALUATION + CLAIM CONTRACT
  • 02 / AUTHORIZED INPUT + LABEL INVENTORY
  • 03 / VERSION + EXECUTION MANIFEST
  • 04 / RESULT + SEGMENT MATRIX
  • 05 / COMPLETE ERROR + EXCEPTION REGISTER
  • 06 / CONTROL + INDEPENDENT VERIFICATION RECORD
  • 07 / CLAIM–EVIDENCE + LIMITATION REGISTER
  • 08 / RISK + CONTROL OWNERSHIP REGISTER
  • 09 / APPROVAL + ACTION RECORD
  • 10 / DECISION MEMO + NEXT TEST
REQUEST A MINIMUM-FIT CHECK
PREPARED BYINDEPENDENT EVALUATOR

Example role only; a real delivery names the responsible preparer and issue date.

VERIFIED BYSEPARATE RECOMPUTATION

Synthetic score agreement is recorded; customer-blinded validation remains outstanding.

DECISION AUTHORITYBUYER-DESIGNATED / UNASSIGNED

No person or organization is authorized by this public sample.

RELEASE RECORDCONDITIONAL / UNSIGNED

Recommendation-only progression; autonomy and external claims remain withheld.

VERAHELM HOLDINGS LLC / INDEPENDENT TECHNICAL EVALUATION

A defensible no-go is still a useful result.

RETURN TO VERAHELMSTART SCOPE REQUEST
Use your browser's print command to save this sample as a PDF. The print layout removes navigation and converts the document to a white archival record.