All posts
Episode 03

AkaiEgoStack: AkaiEgo V2: Egocentric Capture Rig for Scale

Bringing Scaling laws for humanoids

– Research @Akai Space Labs

In Episode 01, we introduced the thesis behind AkaiEgoStack: as robots become more capable, access to high-quality physical-world data becomes equally important.

In Episode 02, we built the first open-spec capture rig to test that thesis in hardware: synchronized stereo cameras and an IMU.

Our further R&D around this led to AkaiEgo V1, our first capture product for operations. Across 500+ hours of recordings, V1 helped us move from validating the capture stack to putting it into practical collection workflows.

Beyond the lab, factors like deployment, wearer comfort, mounting stability, device consistency, and downstream data quality start to matter. Those learnings shaped what came next: AkaiEgo V2.

01

Introducing AkaiEgo V2

AkaiEgo V2 is our next-generation egocentric capture rig, redesigned around the requirements of repeated, real-world data collection. It combines synchronized stereo global-shutter cameras, a lidar sensor, inertial sensing, and on-device recording in a compact, stable, cap-worn form factor designed for natural movement and consistent capture.

Mannequin wearing the AkaiEgo V2 cap-mounted stereo capture rig

Figure

AkaiEgo V2

SpecificationAkaiEgo V2
Cameras2× Global-Shutter RGB
LIDAR64 zone 30 Hz TOF sensor
Resolution1920 × 1080 @ 30 FPS per cam
FOV100°V / 160°H / 200°D
CaptureSynchronized Stereo
CodecMJPEG / H.265
SynchronizationHardware-synchronized Cameras + IMU
IMU9-DoF Inertial Sensing
RecordingOn-device
AudioOn-demand
StorageLocal removable storage / SD
ConnectivityWi-Fi / USB
PowerUSB-C External Battery/Power Bank
MountIntegrated cap-worn mount
Weight< 50 grams

Table

AkaiEgo V2 technical specifications.

The system is built around five practical outcomes:

  1. High-fidelity capture — Global-shutter stereo cameras, 64 zone LIDAR & synchronized inertial sensing preserve the visual, depth and motion information required for accurate reconstruction.
  2. Stable sensing geometry — The integrated mount maintains a consistent camera position and orientation throughout a session.
  3. Non-intrusive form factor — The compact cap-worn design lets the wearer move and work naturally without the device getting in the way.
  4. Repeatable collection — An integrated architecture simplifies setup and operation, enabling more consistent capture across workers and sessions.
  5. High-quality physical-world data — Synchronized multimodal recordings retain the visual, spatial, and temporal context needed for reliable downstream processing in the AkaiEgoStack.

The result is a high-fidelity capture system that converts more real-world activity into reliable, high-quality physical-world data that can flow through the AkaiEgoStack into downstream robotics training applications.

02

What the V1 Gave Us

AkaiEgo V1 was built on off-the-shelf hardware on purpose. We needed to answer the sensing questions before optimizing the device.

That paid off. Over 500+ hours of recordings, we got a working path through the full stack: synchronized visual-inertial capture, a camera-following trajectory, and articulated hand landmarks from the same sequence — more details in the next section.

Figure

V1 in the field — 500+ hours of egocentric recordings across real workplaces

We plan to open-source these V1 recordings soon, so others can explore and build on the data. Early access is now open for teams interested in getting data access ahead of the public release — Request Data

V1 also showed where the bottleneck moved next. Once sensing worked — comfort, setup time, and device consistency started deciding how much useful data we could collect. A rig that needs constant adjustment, or becomes uncomfortable, shortens sessions and adds variation between them.

V2 is the response: a form factor we can actually scale.

03

Recording to Representation

The value of the capture system becomes clearer once we look at what happens after a session is recorded. A typical recording session contains stereo RGB, IMU measurements, timestamps, calibration, and session metadata. When these streams are synchronized, they provide the foundation for structured information from the same physical episode.

The first layer is visual-inertial reconstruction, which gives us an estimate of how the camera moved through the workspace. From that, we can represent the camera/head trajectory and establish the motion of the observer alongside the visual scene.

The next layer is an articulated hand pose. A single frame tells us where the hand is; a sequence tells us how it moved. Our current pipeline recovers hand trajectory across the recording, which can then be interpreted alongside the camera trajectory and original visual stream.

Head / Device Movement
Hand Pose Estimation

Figure

Head / device movement and hand pose estimation from the same synchronized recording

Together, these outputs add geometry, motion, and temporal structure to the raw recording, extracting the physical structure within the observation.

04

Collection Efficiency: Recorded to Usable Hours

Scale is not simply a question of how many devices we can deploy or how many hours they can record. The more useful measure is how much high-fidelity, training-useful data each device and worker can produce over the course of a day.

The capture system should stay out of the way of the person wearing it. The device should not force the worker to change how they move, look, reach, or interact with objects.

A few minutes of friction in one session may seem insignificant; repeated across hundreds or thousands of sessions, it becomes a meaningful constraint on collection capacity.

The same is true for fidelity. A recording that is long but unstable, poorly synchronized, or inconsistent across sessions is far less valuable.

The progression we therefore care about is:

recorded hours → usable hours → high-quality training data

Our experience from operations shows that scaling collection requires more than capture accuracy alone; the device also needs to behave consistently. Small variations in mounting position or device handling can become systematic sources of variation as the number of recordings grows.

That is the transition to AkaiEgo V2: an efficient, reliable, high-fidelity capture rig designed for repeated collection at scale, while remaining stable, comfortable, and unobtrusive for the wearer.

In the next episode, we'll move from the capture layer to the data enrichment layer of AkaiEgoStack, and look at how raw egocentric recordings are transformed into high-quality training data.

Partner with us

Akai Space

Copyright © 2026 Akai Space Labs. All rights reserved.