Datasets

Real-world training data across voice, video, 3D, and paired text. Collected with consent, licensed for commercial use.

Filter datasets by type

Video

8

Real-world video: interviews, egocentric tasks, and robots at work.

Robotic Arm: Pick-and-Place & Manipulation

Robotic-arm pick-and-place episodes pairing 1280x720 video at 30 fps with synchronized robot telemetry, camera information, and manifests, so policies can be evaluated against ground-truth sensor state rather than video alone.

RoboticsComputer Vision

Cultural Interviews: Public Conversations & Perspectives

Public interview videos with unscripted cultural conversations, shot in natural settings at 1080p. Metadata covers interview setting, conversational style, and framing, so coverage can be reviewed before scoping more.

Cultural ContextUnscripted

Egocentric Video: Cleaning, Laundry & Car Wash

First-person task videos covering appliance cleaning, car wash and repair, interior work, and pet sitting, filmed from the worker's own viewpoint at 1080p and 30 fps with hands, tools, and task flow in frame.

Activity RecognitionTask Automation

Humanoid Robot Training: Egocentric Manipulation

Egocentric video of real-world manipulation for robotics and embodied AI, captured at 1080p and 60 fps through a super-wide lens, close to what a humanoid platform's head camera sees during household work, assembly, and collaboration tasks.

Humanoid RoboticsEmbodied AI

First-Person Task Video: Household & Maintenance Activities

Curated first-person videos spanning kitchen work, gardening, light maintenance, cleaning, and fine-motor manipulation, with clear hand-object interaction in every clip and per-clip metadata for category, task type, and capture properties.

EgocentricHousehold Tasks

Warehouse: Kitting, Pick-and-Place & Packaging

Human workers performing kitting, pick-and-place, and packaging tasks in a simulated warehouse, captured from third-person and top-down views at 1280x720 with synchronized audio for logistics and robotics evaluation.

Warehouse AutomationRobotics Training

Behavioral Dataset: Multimodal Video, Depth & Actions

Human-teleoperated demonstrations of everyday household tasks that synchronize RGB video with quantized depth and instance segmentation from head and wrist cameras, shipped in LeRobot format so they drop into existing robot-learning pipelines.

Embodied AITeleoperation

Computer Use: Agent Operating a Desktop Interface

Agent-view recordings of desktop interaction captured at 60 fps, with cursor movement and GUI actions visible frame by frame and subtitle tracks marking on-screen task execution for computer-use agents.

Digital AgentComputer Interaction

Don’t see what you need?

Tell us the modality, language, or scenario, and we’ll scope it with you and collect it.

Send us your spec

Audio

12

Speech and dialect audio recorded by native speakers, each clip paired with a transcript.

Bengali Speech: Single & Multi-Speaker

Bengali speech spanning single-speaker narration and two-speaker conversations, so models see natural turn-taking as well as clean prompts. Audio ships as mono 44.1 kHz WAV, and every clip carries a ground-truth transcript with speaker structure and format metadata.

TranscriptsSpeaker metadata

Hindi Speech: Contributor Recordings

Hindi speech recorded through a moderated contributor app, keeping capture quality and formatting consistent across speakers. Audio ships as mono 48 kHz WAV, and each clip pairs the recording with a validated transcript plus duration, codec, and speaker metadata.

Single-speakerTranscripts

Spanish Speech: Contributor Recordings

Spanish speech recorded through a moderated contributor app, keeping capture quality and formatting consistent across speakers. Each clip pairs the recording with a validated transcript plus duration, codec, and speaker metadata.

Single-speakerTranscripts

Korean Speech: Contributor Recordings

Korean speech recorded through a moderated contributor app, keeping capture quality and formatting consistent across speakers. Audio ships as mono 48 kHz WAV, and each clip pairs the recording with a validated transcript plus duration, codec, and speaker metadata.

Single-speakerTranscripts

Urdu Speech: Contributor Recordings

Urdu speech recorded through a moderated contributor app, keeping capture quality and formatting consistent across speakers. Audio is mono at 48 kHz, and each clip pairs the recording with a validated transcript plus duration, codec, and speaker metadata.

Single-speakerTranscripts

French Speech: Contributor Recordings

French speech recorded through a moderated contributor app, keeping capture quality and formatting consistent across speakers. Audio ships as 48 kHz WAV in mono and stereo, and each clip pairs the recording with a validated transcript plus duration, codec, and speaker metadata.

Single-speakerTranscripts

Bengali Speech: Read & Conversational

Bengali speech captured across read prompts, narration, and free conversation, so the same voices appear in both controlled and natural speech. Audio ships as WAV at 44.1 and 48 kHz with a ground-truth transcript on every clip, ready for ASR, voice interfaces, and alignment work.

TranscriptsASR

Tamil Speech: Read & Conversational

Tamil speech covering read and conversational sessions side by side, which matters for a language where the spoken form drifts far from written text. Audio ships as 48 kHz WAV with ground-truth transcripts on every clip, built for ASR and low-resource language research.

TranscriptsASR

Telugu Speech: Read & Conversational

Telugu speech spanning scripted prompts, narration, and conversational samples, so models can be checked against both careful and casual speech. Audio ships as 48 kHz WAV, and each recording is paired with a transcript for ASR evaluation and voice assistant workflows.

TranscriptsASR

Chinese (Mandarin) Speech Samples

Mandarin speech from native speakers, with each clip pairing mono 48 kHz audio with a ground-truth transcript, language labels, and speaker context for speech recognition and voice interface evaluation.

Native speakersTranscripts

Chinese (Cantonese) Speech Samples

Cantonese speech from native speakers, with each clip pairing mono 16 kHz audio with a ground-truth transcript, language labels, and speaker context, a match for common telephony and ASR pipelines.

Native speakersTranscripts

Chinese (Henan) Speech Samples

Henan Chinese speech from native speakers, a regional variety rarely available as labeled data. Each clip pairs mono 16 kHz audio with a ground-truth transcript, language labels, and speaker context for regional speech analysis.

Native speakersTranscripts

Don’t see what you need?

Tell us the modality, language, or scenario, and we’ll scope it with you and collect it.

Send us your spec