Skip to main content

Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.

1.3k
GitHub Stars
4
Curated Resources
2
Categories
1 month ago
Last Refreshed
ToolsImportant References

Use this list with your AI agent

Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:

"Show me tools resources from awesome-vlm-architectures"

Installation instructions →

What's inside

Tools

  • DualView

    Free side-by-side comparison tool for VLM outputs, images, videos, and AI prompts

Showing a sample of 4 resources. View the full list on GitHub →