Behavioral Dataset: Multimodal Video, Depth & Actions

Human-teleoperated demonstrations of everyday household tasks that synchronize RGB video with quantized depth and instance segmentation from head and wrist cameras, shipped in LeRobot format so they drop into existing robot-learning pipelines.

Embodied AITeleoperationSegmentation

Overview

  • 1.1

    Human activity video with task-level context

  • 1.2

    Metadata suited for reviewing action labels and capture structure

  • 1.3

    Signals for evaluating scene diversity and multimodal alignment

Specifications

License
Non-exclusive license
File types
mp4
Video format
H.264
Frame rate
30fps
Resolution
720x720
Depth
8-bit (quantized)
Segmentation
Instance ID maps
Multi view
Head & wrist cameras
Format
LeRobot

Metadata coverage

sample identifierssample titlesmedia previewsdurationmedia typePII statustagsfeature flagsmedia metadatamodality

Samples

The sample includes 3 clips. Broader coverage can be scoped by action category, environment, modality, and annotation depth.

Related