awesome-document-understanding
github.com/harrytea/awesome-document-understanding ↗Document Artifical Intelligence
201
GitHub Stars
133
Curated Resources
5
Categories
14 hours ago
Last Refreshed
🏆 Milestone📑 Document Understanding🔮 MLLM🎯 Grounded MLLM🎬 Video LLM
Use this list with your AI agent
Add the Context Awesome MCP server to Claude, Cursor, or any MCP client, then ask:
"Show me 📑 document understanding resources from awesome-document-understanding"
Installation instructions →What's inside
📑 Document Understanding
- A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
- A Simple yet Effective Layout Token in Large Language Models for Document Understanding
- BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions
- DiT: Self-supervised Pre-training for Document Image Transformer
- DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
- DocLLM: A layout-aware generative language model for multimodal document understanding
🎬 Video LLM
- Artemis: Towards Referential Understanding in Complex Videos
- Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
- Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
🔮 MLLM
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
- FastVLM: Efficient Vision Encoding for Vision Language Models
- Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
- Flamingo: a Visual Language Model for Few-Shot Learning
🎯 Grounded MLLM
- BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
- Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
- Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
- GroundingGPT: Language Enhanced Multi-modal Grounding Model
🏆 Milestone
- InternLM2 Technical Report
- InternLM: A Multilingual Language Model with Progressively Enhanced Capabilities
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
Showing a sample of 133 resources. View the full list on GitHub →