Audit of 13 Public Robot Datasets Reveals Clean Data Is Not Always Good Training Data
A automated audit tool called RDA v0.9.7 was used to evaluate 13 publicly available LeRobot-format datasets from HuggingFace Hub, covering 4,940 episodes across simulation and real-world robot manipulation tasks. All episodes passed standard integrity checks, showing no missing frames, NaN values, or timestamp errors. However, deeper behavioral and efficiency analysis uncovered dramatic differences between datasets, including one case where two datasets collected on the same robot in the same lab showed a fourfold gap in effective motion ratio. For example, one dataset had robots actively moving 79% of the time, while another had them idle for over 83% of episodes — a distribution that can structurally bias a model toward predicting inaction. The findings highlight that conventional data quality checks are insufficient, and that understanding a dataset's behavioral profile is critical to predicting and diagnosing training outcomes.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in