Robotics
What is robot teleoperation data?
Robot teleoperation data is a recording of a person controlling a real robot through a task, captured so a policy can learn from it. Each recording, called an episode, holds what the robot saw, the state of its joints, the commands it received, and whether the task succeeded.
One episode, from its metadata
Imitation learning trains a policy to copy a skilled operator, and teleoperation records that behavior on the robot's own body. An episode typically syncs color video and depth, proprioception (joint positions, velocities and efforts), the actions sent, a plain-language instruction and a success label, with timestamps on every stream. This abridged excerpt is from a published sample on Gerra's Robot Teleoperation Demonstrations page, where a Booster T1 humanoid moves a light box. Gerra has more than 400,000 accepted episodes from Booster T1, Unitree and LimX Tron robots.
{
"episode": {
"duration_seconds": 29.51,
"task": { "instruction": "Pick up the light box, turn around, and place it on the chair behind you." }
},
"sensors": {
"visual": { "rgb_frames": 1182, "depth_frames": 925, "resolution": { "width": 1280, "height": 720 } },
"proprioception": { "joint_samples": 2338, "joint_sample_rate_hz": 79.23 }
}
}
Align by timestamp, never by frame index
Streams don't start, stop or tick together. In the public sample episode_010, the camera feed starts 62 milliseconds after the episode clock, about two frames at 30 fps, and stops 1.1 seconds before the audio does, while joint states arrive about 77 times a second. Pairing a frame with a joint sample by index is wrong from the first frame; joining on timestamps is right throughout.
So give every sample its own timestamp, and join streams on time. Gerra stamps each stream with an episode-relative and an absolute Unix timestamp at its native rate, keeps native rates instead of resampling, and records a sync-tolerance budget in the format.
Keep the failures
The public sample episode_010 is a box pick-and-place run with "success": false and a failure_reason: the robot dropped the box after placing it over the target table. Behavior cloning usually trains on successes alone, but a success detector, a learned reward or offline RL needs the failures as well, and needs the label to tell them apart.
Formats for pooling robots
Open X-Embodiment pooled data from 22 robots, and every robot brings its own joints, action space and camera layout. LeRobot and RLDS are the common open formats. Gerra packages episodes in the Open Robot Training Format (ORTF), which states the action space and coordinate frames explicitly, keeps native stream rates on a shared clock, and converts to both. The ORTF file format documentation shows an episode archive's full layout.
What real episodes measure about simulation
Real recordings also measure the gap to simulation. Gerra's sim2real study replayed 123 real Booster T1 episodes, 284,794 joint-state samples, in MuJoCo with position control and compared them joint by joint. Mean position error was 5.56 degrees. The knee pitch joints were worst at 12.0 and 12.2 degrees, leg joints had about 1.2 times the error of arm joints, and the study points to contact dynamics as the main modeling problem.