The honesty ledger

Last updated 2026-08-07. Every claim on the main page traces to a row here. Measured on recordings the system never saw during setup; protocols named per row.

Measured capabilities

ClaimNumberDataProtocol
Rank agreement with competition judges, 10 m platform diving+0.71 Spearman ρPublic judged-competition diving corpus, 70 held-out divesTrained on a disjoint split; scored dives never seen in setup
Surgical phase recognition, laparoscopic cholecystectomy81% frame accuracyHeiChole — public research benchmark, 24 videosSix-fold cross-validation: every video scored by a model that never trained on it
Confidence honesty (phase predictions)~2.5% expected calibration errorSame surgical corpusTemperature scaling fit on validation logits; reliability diagram ships in-product
Breadth13 corpora · 6 fieldsSurgery, simulation (cataract, peg transfer), judged sport, industrial assembly, fitness, musicSame product, per-field heads trained on a few dozen scored recordings each

Measured limits — published on purpose

BoundaryWhat we found
Music performance (piano)Piece difficulty is readable from video with meaningful accuracy; a player's finesse lives largely in sound and fine finger technique that silent video can't fully carry. We publish the boundary instead of shipping a confidently wrong score.
Small corporaWith few validation videos, confidence intervals are wide — the product shows them and runs significance tests so noise is labeled as noise, not sold as improvement.
Stage of validationAll surgical results are on a public research benchmark (HeiChole, 24 videos). ActionQuality has not yet run a hospital pilot; that is the next step, not a claim we make.

How this page works

New capability claims appear here only with a number, the dataset, and the protocol. If a claim on the main page ever lacks a row here, that's a bug — write to [email protected].