Embodied intelligence models must jointly learn environmental understanding, task
intent, spatial movement, and object manipulation from real human behavior. Existing
egocentric datasets, however, are commonly constrained by a fixed field of view.
Pointing a camera toward a local workspace may remove the surrounding environment and
future targets; pointing it toward the broader scene may move the body, hands, and
contact events outside the frame.
Ego360 uses a lightweight neck-worn panoramic camera to turn this irreversible capture
decision into a configurable observation after capture. The complete omnidirectional
source supports arbitrary viewing directions, fields of view, and task-conditioned
perspective renderings while maintaining a persistent spatial reference for gaze,
motion, and interaction.
Property
Conventional RGB
Ego360 source
Horizontal field of view≈109°360°
Vertical field of view≈60°180°
Instantaneous azimuth coverage≈30.2%100%
Lateral and opposite contextOutside viewRetained
Post-hoc view selectionUnsupportedSupported
What one demonstration preserves
Intent
Operator speech, episode-level instructions, and procedurally ordered task steps.
Interaction
Hands, objects, contact states, object motion, and state changes on a shared timeline.
Motion
Whole-body and articulated hand motion transformed into world-referenced coordinates.
Environment
Panoramic scene context, episode-level reconstruction, camera pose, and global trajectory.
Research position
The initial release is deliberately scoped as a data and infrastructure contribution.
It establishes observation, annotation, geometry, privacy, and quality-control
foundations without tying the dataset to a particular VLA architecture, training recipe,
or robot platform.