AkaiEgoStack: AkaiEgo V2: Egocentric Capture Rig for Scale
Bringing Scaling laws for humanoids
– Research @Akai Space Labs
In Episode 01, we introduced the thesis behind AkaiEgoStack: as robots become more capable, access to high-quality physical-world data becomes equally important.
In Episode 02, we built the first open-spec capture rig to test that thesis in hardware: synchronized stereo cameras and an IMU.
Our further R&D around this led to AkaiEgo V1, our first capture product for operations. Across 500+ hours of recordings, V1 helped us move from validating the capture stack to putting it into practical collection workflows.
Beyond the lab, factors like deployment, wearer comfort, mounting stability, device consistency, and downstream data quality start to matter. Those learnings shaped what came next: AkaiEgo V2.
Introducing AkaiEgo V2
AkaiEgo V2 is our next-generation egocentric capture rig, redesigned around the requirements of repeated, real-world data collection. It combines synchronized stereo global-shutter cameras, a lidar sensor, inertial sensing, and on-device recording in a compact, stable, cap-worn form factor designed for natural movement and consistent capture.

Figure
AkaiEgo V2
| Specification | AkaiEgo V2 |
|---|---|
| Cameras | 2× Global-Shutter RGB |
| LIDAR | 64 zone 30 Hz TOF sensor |
| Resolution | 1920 × 1080 @ 30 FPS per cam |
| FOV | 100°V / 160°H / 200°D |
| Capture | Synchronized Stereo |
| Codec | MJPEG / H.265 |
| Synchronization | Hardware-synchronized Cameras + IMU |
| IMU | 9-DoF Inertial Sensing |
| Recording | On-device |
| Audio | On-demand |
| Storage | Local removable storage / SD |
| Connectivity | Wi-Fi / USB |
| Power | USB-C External Battery/Power Bank |
| Mount | Integrated cap-worn mount |
| Weight | < 50 grams |
Table
AkaiEgo V2 technical specifications.
The system is built around five practical outcomes:
- High-fidelity capture — Global-shutter stereo cameras, 64 zone LIDAR & synchronized inertial sensing preserve the visual, depth and motion information required for accurate reconstruction.
- Stable sensing geometry — The integrated mount maintains a consistent camera position and orientation throughout a session.
- Non-intrusive form factor — The compact cap-worn design lets the wearer move and work naturally without the device getting in the way.
- Repeatable collection — An integrated architecture simplifies setup and operation, enabling more consistent capture across workers and sessions.
- High-quality physical-world data — Synchronized multimodal recordings retain the visual, spatial, and temporal context needed for reliable downstream processing in the AkaiEgoStack.
The result is a high-fidelity capture system that converts more real-world activity into reliable, high-quality physical-world data that can flow through the AkaiEgoStack into downstream robotics training applications.
What the V1 Gave Us
AkaiEgo V1 was built on off-the-shelf hardware on purpose. We needed to answer the sensing questions before optimizing the device.
That paid off. Over 500+ hours of recordings, we got a working path through the full stack: synchronized visual-inertial capture, a camera-following trajectory, and articulated hand landmarks from the same sequence — more details in the next section.
Figure
V1 in the field — 500+ hours of egocentric recordings across real workplaces
We plan to open-source these V1 recordings soon, so others can explore and build on the data. Early access is now open for teams interested in getting data access ahead of the public release — Request Data
V1 also showed where the bottleneck moved next. Once sensing worked — comfort, setup time, and device consistency started deciding how much useful data we could collect. A rig that needs constant adjustment, or becomes uncomfortable, shortens sessions and adds variation between them.
V2 is the response: a form factor we can actually scale.
Recording to Representation
The value of the capture system becomes clearer once we look at what happens after a session is recorded. A typical recording session contains stereo RGB, IMU measurements, timestamps, calibration, and session metadata. When these streams are synchronized, they provide the foundation for structured information from the same physical episode.
The first layer is visual-inertial reconstruction, which gives us an estimate of how the camera moved through the workspace. From that, we can represent the camera/head trajectory and establish the motion of the observer alongside the visual scene.
The next layer is an articulated hand pose. A single frame tells us where the hand is; a sequence tells us how it moved. Our current pipeline recovers hand trajectory across the recording, which can then be interpreted alongside the camera trajectory and original visual stream.
Figure
Head / device movement and hand pose estimation from the same synchronized recording
Together, these outputs add geometry, motion, and temporal structure to the raw recording, extracting the physical structure within the observation.
Collection Efficiency: Recorded to Usable Hours
Scale is not simply a question of how many devices we can deploy or how many hours they can record. The more useful measure is how much high-fidelity, training-useful data each device and worker can produce over the course of a day.
The capture system should stay out of the way of the person wearing it. The device should not force the worker to change how they move, look, reach, or interact with objects.
A few minutes of friction in one session may seem insignificant; repeated across hundreds or thousands of sessions, it becomes a meaningful constraint on collection capacity.
The same is true for fidelity. A recording that is long but unstable, poorly synchronized, or inconsistent across sessions is far less valuable.
The progression we therefore care about is:
Our experience from operations shows that scaling collection requires more than capture accuracy alone; the device also needs to behave consistently. Small variations in mounting position or device handling can become systematic sources of variation as the number of recordings grows.
That is the transition to AkaiEgo V2: an efficient, reliable, high-fidelity capture rig designed for repeated collection at scale, while remaining stable, comfortable, and unobtrusive for the wearer.
In the next episode, we'll move from the capture layer to the data enrichment layer of AkaiEgoStack, and look at how raw egocentric recordings are transformed into high-quality training data.

