Skip to main content

Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.

738
GitHub Stars
25
Curated Resources
4
Categories
2 hours ago
Last Refreshed
Other Speech ApplicationsLarge Audio Models in MusicAudio DatasetsAI Model Catalogs

Use this list with your AI agent

Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:

"Show me ai model catalogs resources from awesome-large-audio-models"

Installation instructions →

What's inside

AI Model Catalogs

  • AI Models Catalog

    Structured catalog of 4,587+ AI models across 95 providers. Includes 118 audio input models and 34 audio output models with pricing and capabilities. Interactive catalog:

Other Speech Applications

  • AnveVoice

    Voice AI for real-time website interaction using audio models. Sub-700ms STT-to-action, 50+ languages, real DOM actions. MIT-0 license.

  • StoryRoute

    GPS-triggered LLM story generation synthesized via a large TTS model into real-time audio narration as users walk through any city

Audio Datasets

  • Audiocaps

    Audiocaps: Generating captions for audios in the wild

  • Audio set

    Audio set: An ontology and human-labeled dataset for audio events

  • Clotho

    Clotho: An audio captioning dataset

  • CommonVoice 11

    CommonVoice: A Massively Multilingual Speech Corpus

  • CoVoST

    CoVoST: A Large-Scale Multilingual Speech-To-Text Translation Corpus

  • CVSS

    CVSS: A Massively Multilingual Speech-to-Speech Translation Corpus

Large Audio Models in Music

Showing a sample of 25 resources. View the full list on GitHub →