awesome-large-audio-models
github.com/emulationai/awesome-large-audio-models ↗Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
Use this list with your AI agent
Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:
"Show me ai model catalogs resources from awesome-large-audio-models"
Installation instructions →What's inside
AI Model Catalogs
- AI Models Catalog
Structured catalog of 4,587+ AI models across 95 providers. Includes 118 audio input models and 34 audio output models with pricing and capabilities. Interactive catalog:
Other Speech Applications
- AnveVoice
Voice AI for real-time website interaction using audio models. Sub-700ms STT-to-action, 50+ languages, real DOM actions. MIT-0 license.
- StoryRoute
GPS-triggered LLM story generation synthesized via a large TTS model into real-time audio narration as users walk through any city
Audio Datasets
- Audiocaps
Audiocaps: Generating captions for audios in the wild
- Audio set
Audio set: An ontology and human-labeled dataset for audio events
- Clotho
Clotho: An audio captioning dataset
- CommonVoice 11
CommonVoice: A Massively Multilingual Speech Corpus
- CoVoST
CoVoST: A Large-Scale Multilingual Speech-To-Text Translation Corpus
- CVSS
CVSS: A Massively Multilingual Speech-to-Speech Translation Corpus
Large Audio Models in Music
- voicetoinstrument.com
Convert voice to instrumental tracks using AI
Showing a sample of 25 resources. View the full list on GitHub →