VideoChat3 logo

VideoChat3

A Fully Open Video MLLM for Efficient, Generalist Video Understanding

arXiv Paper     Homepage Homepage     Hugging Face Models & Data

VideoChat3 overview across temporal perception, long-video reasoning, temporal grounding, and online proactive response

One model for fine-grained motion, long-video reasoning, temporal grounding, and online proactive response.

## ✨ Overview **VideoChat3** is a 4B generalist video MLLM built to understand video across timeβ€”from subtle motion and hour-long stories to precise temporal evidence and live streams. It combines **I3D-ViT** for 16Γ— spatiotemporal compression with **Adaptive Frame Resolution** for evidence-aware streaming, trained on **Academic2M**, **LV116K**, and **OL617K**. ## πŸš€ Highlights - 🎬 **Generalist video understanding:** one model for motion, long video, temporal grounding, and live streaming. - ⚑ **Token-efficient architecture:** I3D-ViT compresses redundant visual tokens while preserving spatiotemporal evidence. - πŸ” **Adaptive streaming perception:** frame resolution is increased only when closer visual inspection is needed. - πŸ”“ **Open resources:** model weights and the complete training datasets are publicly available. ## πŸ“‹ TODO - [x] πŸ€— Release model weights and data - [ ] πŸ› οΈ Release training code ## Citation ``` @misc{videochat3, title={VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding}, author={Xinhao Li and Yuhan Zhu and Xiangyu Zeng and Yuhao Dong and Haoning Wu and Zhiqiu Zhang and Yuandong Yang and Changlian Ma and Qingyu Zhang and Yansong Shi and Xinyu Chen and Haoran Chen and Zizheng Huang and Jun Zhang and Kun Ouyang and Lin Sui and Ziang Yan and Yicheng Xu and Chenting Wang and Yinan He and Hongjie Zhang and Yi Wang and Yu Qiao and Yali Wang and Ziwei Liu and Kai Chen and Limin Wang}, year={2026}, eprint={2607.14935}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2607.14935}, } ```