TAKELOOP / RESEARCH / 09
The field is scaling its data.
Published industry evidence is shown separately from Takeloop operating targets and proposed experiments.
PUBLISHED BY THIRD PARTIES / NOT TAKELOOP RESULTS
76K trajectories350 hours86 tasks564 scenes50 collectors
DROID dataset / RSS 2024
Published DROID dataset scale. Not Takeloop data.
160,266 tasks527 skills22 robots
Open X-Embodiment / arXiv:2310.08864
Published Open X-Embodiment dataset scale. Not Takeloop data.
10 → 60 training hours6.3s → 4.3s per package88% → ~95% barcode orientation
Figure, published logistics result
Published Figure result. Not a Takeloop benchmark.
380 long-horizon demonstrations / 50–100 teleoperation hours2,000 demonstrations / 86% success100K hours / 66% ICL vs 9% language prompting
Skild AI, Introducing S1
Vendor-reported Skild result. Not a Takeloop result or independently replicated benchmark.
PROPOSED TAKELOOP BENCHMARK
Does recovery data change the outcome?
Recovery Lift measures the change in task success after adding human takeover and recovery data.
Protocol defined. Experiment not yet conducted.
RECOVERY LIFTPending measurementChange in task success rate, reported in percentage points
Control
500 demonstrations without takeover or recovery segments
Takeloop
500 demonstrations with takeover and recovery segments
- One robot and one defined task family
- 500 training demonstrations per cohort
- 100 held-out evaluation episodes
- Controlled training conditions
- Held-out evaluation with randomized episode order
- Secondary metrics: recovery success, intervention rate, completion time and failure rate
QUALITY SYSTEMS / V0.1
QA Technical Note
Episode integrity, review methodology and pilot acceptance reporting.
PDF