awesome-document-understanding
github.com/harrytea/awesome-document-understanding ↗Document Artifical Intelligence
201
GitHub Stars
201
Curated Resources
5
Categories
2 hours ago
Last Refreshed
🏆 Milestone📑 Document Understanding🔮 MLLM🎯 Grounded MLLM🎬 Video LLM
Use this list with your AI agent
Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:
"Show me 📑 document understanding resources from awesome-document-understanding"
Installation instructions →What's inside
📑 Document Understanding
- A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
- A Simple yet Effective Layout Token in Large Language Models for Document Understanding
- BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions
- DeepSeek-OCR 2: Visual Causal Flow
- DeepSeek-OCR: Contexts Optical Compression
- DiT: Self-supervised Pre-training for Document Image Transformer
🎬 Video LLM
- Artemis: Towards Referential Understanding in Complex Videos
- Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
- Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
🔮 MLLM
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
- FastVLM: Efficient Vision Encoding for Vision Language Models
- Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
- Flamingo: a Visual Language Model for Few-Shot Learning
🎯 Grounded MLLM
- BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
- Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
- Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
- GroundingGPT: Language Enhanced Multi-modal Grounding Model
🏆 Milestone
- CogVLM2: Visual Language Models for Image and Video Understanding
- CogVLM: Visual Expert for Pretrained Language Models
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Showing a sample of 201 resources. View the full list on GitHub →