awesome-egocentric-vision
github.com/sid2697/awesome-egocentric-vision ↗A curated list of egocentric (first-person) vision and related area resources
Use this list with your AI agent
Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:
"Show me eccv resources from awesome-egocentric-vision"
Installation instructions →What's inside
Papers
- 3D Hand Pose Estimation in Everyday Egocentric ImagesECCV
Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. In ECCV 2024.
- 3D Human Pose Perception from Egocentric Stereo VideosCVPR
Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. In CVPR 2024.
- 4Diff: 3D-Aware Diffusion Model for Third-to-First Viewpoint TranslationECCV
Feng Cheng, Mi Luo, Huiyu Wang, Alex Dimakis, Lorenzo Torresani, Gedas Bertasius, and Kristen Grauman. In ECCV 2024.
- A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task PerspectivesCVPR
Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, and Giuseppe Averta. In CVPR 2024.
- Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric VideosECCV
Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, and Kristen Grauman. In ECCV 2024.
- Action Scene Graphs for Long-Form Understanding of Egocentric VideosCVPR
Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi, and Giovanni Maria Farinella. In CVPR 2024.
Datasets
- ADLAll Datasets
20 subjects performing daily activities in their native environments.
- Aria Digital TwinAll Datasets
A comprehensive egocentric dataset containing 200 sequences of real-world activities conducted by Aria wearers in two real indoor scenes with 398 object instances (324 stationary and 74 dynamic).
- Aria Everyday Activities (AEA)All Datasets
143 daily-activity sequences recorded by multiple wearers in five indoor locations with Project Aria, with globally-aligned 3D trajectories, scene point clouds, per-frame eye gaze, and time-aligned speech.
- Aria Gen 2 Pilot Dataset (A2PD)All Datasets
Egocentric multimodal dataset captured with Aria Gen 2 glasses across everyday scenarios (cleaning, cooking, eating, playing, walking) between a primary user and three friends, with rich sensor streams (RGB, eye tracking, IMU, spatial audio, PPG, GPS) and machine-perception outputs.
- Aria Navigation Dataset (AND)All Datasets
Roughly 4 hours of Project Aria egocentric recordings for humanoid/egocentric navigation, introduced with LookOut.
- Assembly101All Datasets
4,321 videos of assembling/disassembling 101 take-apart toy vehicles, with 8 static and 4 egocentric views, 100K+ coarse and 1M fine-grained action segments, and 18M 3D hand poses, for procedural activity recognition, anticipation, segmentation, and mistake detection.
Showing a sample of 445 resources. View the full list on GitHub →