Checks before claims.
Every episode moves through signal checks, structured review and acceptance against an agreed task rubric.
A dataset is useful only when its structure, signals and outcomes can be trusted.
| Metric | Definition | Target / reporting method | Commitment |
|---|---|---|---|
| Episode QA coverage | Delivered episodes reviewed against the agreed task rubric. | Episode-level review | 100% Protocol |
| Schema validation | Required fields and data types conform to the agreed schema. | Automated validation | 100% Protocol |
| Timestamp integrity | Alignment between recorded streams. | Synchronization check | Reported in ms Measured after collection |
| Dropped frames | Frames absent from expected capture cadence. | Stream integrity check | Reported as % Measured after collection |
| Trajectory completeness | Required action and state traces present for the episode. | Continuity validation | Reported as % Measured after collection |
| Takeover boundary accuracy | Start and end of human intervention correctly marked. | Boundary review | Reported as % Measured after collection |
| Inter-reviewer agreement | Agreement between independent QA reviews. | Double-review sample | Reported as % Measured after collection |
| Final rejection rate | Episodes removed before delivery. | Final acceptance review | Reported as % Measured after collection |
Protocol values are commitments. Performance values are reported from the delivered collection.
Every episode moves through signal checks, structured review and acceptance against an agreed task rubric.
Task training, calibration and review criteria are agreed for each capture program.
Initial pilot QA protocol and scorecard.