TAKELOOP / RESEARCH / 09

The field is scaling its data.

Published industry evidence is shown separately from Takeloop operating targets and proposed experiments.

PUBLISHED BY THIRD PARTIES / NOT TAKELOOP RESULTS

DROID

Primary source
76K trajectories350 hours86 tasks564 scenes50 collectors

DROID dataset / RSS 2024

Published DROID dataset scale. Not Takeloop data.

Open X-Embodiment

Primary source
160,266 tasks527 skills22 robots

Open X-Embodiment / arXiv:2310.08864

Published Open X-Embodiment dataset scale. Not Takeloop data.

Figure Helix

Primary source
10 → 60 training hours6.3s → 4.3s per package88% → ~95% barcode orientation

Figure, published logistics result

Published Figure result. Not a Takeloop benchmark.

Skild S1

Primary source
380 long-horizon demonstrations / 50–100 teleoperation hours2,000 demonstrations / 86% success100K hours / 66% ICL vs 9% language prompting

Skild AI, Introducing S1

Vendor-reported Skild result. Not a Takeloop result or independently replicated benchmark.

PROPOSED TAKELOOP BENCHMARK

Does recovery data change the outcome?

Recovery Lift measures the change in task success after adding human takeover and recovery data.

Protocol defined. Experiment not yet conducted.

RECOVERY LIFTPending measurementChange in task success rate, reported in percentage points

Control

500 demonstrations without takeover or recovery segments

Takeloop

500 demonstrations with takeover and recovery segments

  • One robot and one defined task family
  • 500 training demonstrations per cohort
  • 100 held-out evaluation episodes
  • Controlled training conditions
  • Held-out evaluation with randomized episode order
  • Secondary metrics: recovery success, intervention rate, completion time and failure rate

QUALITY SYSTEMS / V0.1

QA Technical Note

Episode integrity, review methodology and pilot acceptance reporting.

PDF