A joint research project by

Research dataset · Physical intelligence

Ego360: A Panoramic View of Human Experience for Physical Intelligence

DeepCybo & Insta360

A large-scale panoramic human-demonstration dataset that preserves environmental context, visible body regions, hand–object interaction, operator speech, human motion, and scene geometry within one temporally aligned record.

Overview of Ego360 collection domains and panoramic sensing compared with fixed field-of-view capture
Figure 1. Ego360 overview. Demonstrations span factories, homes, laboratories, and retail settings. Unlike fixed-FOV capture, neck-worn panoramic capture retains surrounding context, body pose, hand–object interaction, and gaze for post-hoc view selection.
10,000 hreal-world demonstrations
1,000+contributors
4 domainsfactory, home, laboratory, retail
22,582speech-aligned captions characterized

Overview

Retain first, select later

Embodied intelligence models must jointly learn environmental understanding, task intent, spatial movement, and object manipulation from real human behavior. Existing egocentric datasets, however, are commonly constrained by a fixed field of view. Pointing a camera toward a local workspace may remove the surrounding environment and future targets; pointing it toward the broader scene may move the body, hands, and contact events outside the frame.

Ego360 uses a lightweight neck-worn panoramic camera to turn this irreversible capture decision into a configurable observation after capture. The complete omnidirectional source supports arbitrary viewing directions, fields of view, and task-conditioned perspective renderings while maintaining a persistent spatial reference for gaze, motion, and interaction.

Property Conventional RGB Ego360 source
Horizontal field of view≈109°360°
Vertical field of view≈60°180°
Instantaneous azimuth coverage≈30.2%100%
Lateral and opposite contextOutside viewRetained
Post-hoc view selectionUnsupportedSupported

What one demonstration preserves

Intent

Operator speech, episode-level instructions, and procedurally ordered task steps.

Interaction

Hands, objects, contact states, object motion, and state changes on a shared timeline.

Motion

Whole-body and articulated hand motion transformed into world-referenced coordinates.

Environment

Panoramic scene context, episode-level reconstruction, camera pose, and global trajectory.

Research position

The initial release is deliberately scoped as a data and infrastructure contribution. It establishes observation, annotation, geometry, privacy, and quality-control foundations without tying the dataset to a particular VLA architecture, training recipe, or robot platform.

Dataset construction

From continuous panoramic recordings to structured supervision

Ego360 converts raw demonstrations into privacy-safe, task-structured, and spatially grounded segment packages through a two-stage pipeline. Temporal provenance and processing status are retained throughout the process for traceable quality control.

Two-stage privacy-aware Ego360 data pipeline
Panoramic data pipeline. Stage I extracts privacy-safe action segments from synchronized panoramic video and speech. Stage II combines temporal frames with spatial evidence to produce graphs, structured captions, and grounded QA.
Stage IPrivacy-aware temporal segmentation

Collection and capture

Demonstrations are recorded across factories, homes, laboratories, and retail environments. A dual-fisheye 360° camera with synchronized audio preserves the surrounding workspace, visible body, interaction, and natural task-oriented narration on one timeline. Sessions are organized by participant, location, and task with calibration, synchronization, and capture-quality checks.

Visual screening and speech alignment

A fixed-view proxy enables efficient visual screening while the panoramic stream remains canonical. Severe occlusion, capture failures, and degraded windows are filtered. Synchronized audio is processed with voice activity detection, automatic speech recognition, and word-level temporal alignment; neighboring utterances are consolidated into narration units.

Narration-guided interval construction

Narration provides a proposal rather than a final boundary because an operator may speak before acting, combine steps, or mention unrelated content. Each unit records narration onset, instruction completion, and retained action endpoint. The interval is reconciled with visual-quality windows to preserve the context needed to interpret the complete physical action.

Semantic validation and privacy

Transcript and visual evidence are jointly checked for action relevance. Original ASR, normalized descriptions, match labels, scores, and reasons remain auditable. Privacy protection is applied to the full panoramic sequence: faces are blurred and screens, identity documents, addresses, and other identifying content are reviewed and masked or removed before release.

Stage IISpatially grounded multimodal annotation

Temporal and spatial evidence

Caption and QA generation sample the full temporal extent at one frame per second, uniformly limiting long segments to 100 frames. A separate spatial branch samples approximately once every three seconds with at most 30 frames, emphasizing persistent objects, interaction regions, scene text, and geometric context.

Ego-Element Graph

The graph centers on an explicit ego node and a compact set of task-relevant object and region nodes. Nodes store semantic identity, image region, ego-relative position, visible text, confidence, supporting frames, temporal state, and normalized relative depth. Directed edges encode left/right, above/below, near/far, in-front-of/behind, support, containment, and contact.

Structured captions and QA

Captions separate scene elements, spatial dynamics, and action execution. Procedural descriptions organize each activity into temporally ordered steps. QA pairs cover action ordering, visible text, counting, attributes, spatial relations, state changes, and higher-level visual reasoning using the same evidence reference.

Validation and packaging

Deterministic checks validate graph parsing and normalization, caption structure, QA structure, requested counts, and status. Failures trigger a frames-only fallback rather than terminating the pipeline. Each package aligns privacy-safe video, temporal metadata, sampled frames, graph, caption, QA, raw responses, and processing status under one stable segment identifier.

Factory and home panoramic annotations
Factory and home tasks. Synchronized directions, temporal frames, cross-view evidence, action annotations, and grounded QA.
Retail and laboratory panoramic annotations
Retail and laboratory tasks. The same evidence structure spans commercial and biochemical workspaces.

Geometric grounding

Human motion and scene geometry in compatible world frames

Semantic labels describe what occurs; geometric grounding describes how the operator moves through and acts within the physical environment. Ego360 links local hand–object interaction to global locomotion and persistent scene structure.

World-referenced human motion

The human-motion branch begins with a task-relevant monocular RGB crop from the dewarped panoramic recording. SAM 3 masks or YOLOv7 person detections localize the wearer, and optical-flow tracking maintains temporal association. Motion-adaptive interpolation adds frames only during high-speed hand intervals that risk blur or temporal under-sampling.

SAM-Body4D recovers temporally consistent 3D motion. The representation retains a 16-joint body skeleton and 21 articulated joints for each hand. Timestamped camera poses transform the result from camera coordinates into the calibrated world frame, followed by temporal smoothing that suppresses jitter while preserving fast hand motion and contact transitions.

Temporally aligned RGB views, 3D body and hand skeletons, and relative-depth maps in factory and retail tasks
Human-motion alignment with depth. As in the paper, factory and retail sequences pair projected body and hand joints with the corresponding 3D skeletons and temporally aligned relative-depth maps.

Panoramic scene reconstruction and global trajectory

Scene reconstruction uses the complete panoramic observation rather than the task crop. Each episode is sampled at a configurable one-second interval. A registered timestamp yields one central fisheye-camera pose and four co-located virtual fisheye views facing different directions, increasing coverage without pretending that the views are independently captured.

Aerial triangulation and structure from motion establish an episode-level coordinate frame. Multi-view depth estimates are filtered for cross-view consistency and fused into a persistent textured surface while transient foreground observations are excluded. Timestamped camera centers and viewing directions then form the wearer’s global trajectory through the episode map.

Panoramic scene reconstruction with selected views, camera poses, and global trajectory
Scene reconstruction. Selected observations are placed at their recovered positions and viewing directions along the episode-level camera trajectory.

Long-horizon navigation–manipulation continuity

Navigation and manipulation retain the source-video timeline. Manipulation intervals connect to ordered steps, contact states, and aligned body and hand motion; navigation intervals retain panoramic observations, camera poses, viewing directions, and scene positions. This preserves target search, approach, local action, object transport, and transition to the next interaction site as one continuous demonstration rather than unrelated clips.

Long-horizon cooking episode alternating between navigation and manipulation
Navigation–manipulation continuity. A cooking episode connects stove operation, refrigerator access, object retrieval, return movement, and final placement on one shared timeline.

Dataset characterization

Diversity across participants, environments, semantics, and task duration

Ego360 participant, environment, and semantic diversity
Dataset composition. Participant metadata, collection domains, and concepts extracted from speech-aligned captions.

Participants and embodiment

Ego360 includes more than 1,000 contributors. A detailed 50-person sample records demographic and anthropometric metadata including height, arm span, leg length, body mass, shoulder width, upper-arm length, forearm length, and hand length. These variables affect reachable workspace, manipulation posture, hand placement, and spatial relationships to objects.

Scenes and tasks

A characterized 500-hour subset contains 209.0 hours of home activity, 206.5 hours of commercial activity, 59.5 hours of factory activity, and 25.0 hours of laboratory activity. The environments differ substantially in layout, object inventory, procedure, safety constraints, and interaction style.

Semantic coverage

Analysis of 22,582 speech-aligned captions organizes vocabulary into actions, objects and entities, scenes and spatial anchors, and attributes and spatial states. The resulting supervision captures what action occurs, which objects participate, where it occurs, and how physical state changes.

Long-horizon structure

Episodes span local manipulation, room-scale navigation, target search and approach, device use, multi-stage procedures, and transitions between interaction sites. All intervals remain connected by source timestamps and the hierarchical task representation.

Data-centric evaluation

Observational coverage and annotation integrity

The current evaluation measures Ego360 as a data resource rather than reporting gains from a particular downstream model. Metrics are computed at the event level and macro-averaged over eight independent episodes and 357 events; 95% confidence intervals use 10,000 episode-cluster bootstrap samples.

Panoramic evidence retention across eight episodes and 357 events
MetricOverallCommercialHomeFactoryLaboratory
Task context visible ↑100.0%100.0%100.0%100.0%100.0%
Any hand visible ↑88.0%77.1%91.5%88.7%94.6%
Both hands visible ↑78.8%64.1%87.6%71.8%91.5%
Target object visible ↑90.6%69.3%99.9%99.0%94.2%
Interaction visible ↑83.4%64.7%91.1%87.7%90.3%
Mean event coverage ↑85.4%68.3%91.5%89.2%92.7%
Complete-event recall ↑82.7%63.2%90.0%86.7%91.0%
Full torso visible ↑4.4%0.0%12.6%1.9%2.9%
Mean maximum invisible gap ↓1.66 s3.70 s0.79 s1.07 s1.07 s

Complete-event recall is the proportion of events with at least 90% action-evidence coverage. Commercial settings remain the most challenging because dense shelving, frequent reorientation, and local occlusion interrupt interaction evidence.

Annotation coverage and automatic structural consistency
ComponentEvaluated dataCheckResult
Temporal segmentation357 eventsExact boundary transfer357/357
ASR357 eventsOutput provided357/357
Caption and VLM parsing357 eventsCaption and parse success100% / 100%
Question answering357 eventsOutput availability714 pairs; 100% coverage
Visibility annotations357 eventsOutput availability357/357; 35 uncertain as NA
Ego-Element Graph357 eventsOutput availability357/357
Graph structural audit357 graphsSchema, boxes, frames, relations100% for all checks
Audio–visual audit325 valid eventsMatch rate / relevance score92.5% / 73.7%
Screen-segment QC578 segmentsValid / review / rejected497 / 39 / 42

Supported conclusion

Panoramic capture substantially expands instantaneous field of view and jointly retains environmental, hand, object, and interaction evidence for most evaluated events. The pipeline consistently produces complete and structurally valid multimodal outputs while identifying lower-quality candidates.

Important qualification

The neck-worn configuration does not guarantee continuous full-body visibility. Output availability and structural validity also do not establish semantic correctness. Human-verified caption factuality, hallucination rate, QA accuracy, and graph semantic accuracy are not claimed until the corresponding audit is complete.

Ethics, privacy, and limitations

Privacy is part of dataset construction, not a post-hoc release step

Collection and release safeguards

Data is collected under informed consent and only in environments approved by owners or authorized managers. Every sequence is reviewed for personally identifiable information. Visible faces are blurred; other identifying content is removed or masked where necessary; only sequences that pass the de-identification audit are retained for release.

Current limitations

Panoramic imagery may contain stitching artifacts, severe distortion, motion blur, and self-occlusion. Pose and reconstruction can degrade under weak texture, rapid motion, or persistent occlusion. Participant, geographic, and task distributions may not represent the full diversity of real-world activity. Downstream VLA pretraining and human-to-robot transfer remain future evaluations.

Project resources

Paper available. Dataset release in preparation.

The public dataset URL will replace the placeholder below when release review, documentation, licensing, and access safeguards are complete.

The dataset link will be published here when the release is ready.