awesome-audio-visual
github.com/krantiparida/awesome-audio-visual ↗A curated list of different papers and datasets in various areas of audio-visual processing
Use this list with your AI agent
Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:
"Show me resources resources from awesome-audio-visual"
Installation instructions →What's inside
Resources
- 2.5D Visual Sound
Gao, R., & Grauman, K. (CVPR 2019)
- ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning
Lee, S., Chung, J., Yu, Y., Kim, G., Breuel, T., Chechik, G., & Song, Y. (ICCV 2021)
- A Closer Look at Weakly-Supervised Audio-Visual Source Localization
Mo, S., & Morgado, P. (NeurIPS 2022)
- Active Audio-Visual Separation of Dynamic Sound Sources
Majumder, S. & Grauman, K. (ECCV 2022)
- Active Contrastive Learning of Ausio-Visual Video Representations
Ma, S., Zeng, Z., McDuff, D., & Song, Y. (ICLR 2021)
- AlignNet: A Unifying Approach to Audio-Visual Alignment
Wang, J., Fang, Z., & Zhao, H. (WACV 2020)
Datasets
- ACAV100M
140 million full-length videos (total duration 1,030 years) and produce a dataset of 100 million 10-second clips (31 years) with high audio-visual correspondence.
- AIST++
A large-scale 3D human dance motion dataset, which contains a wide variety of 3D motion paired with music It is built upon the AIST Dance Database, which is an uncalibrated multi-view collection of dance videos.
- AudioSet
Audio-Visual Classification
- AudioSet Single Source
Subset of AudioSet videos containing only a single souding object
- AudioSetZSL
Audio-Visual Zero-shot Learning
- AuDio Visual Aerial sceNe reCognition datasEt (ADVANCE)
Geotagged aerial images and sounds, classified into 13 scene classes
Showing a sample of 262 resources. View the full list on GitHub →