LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs




[[Project Page](https://kd-tao.github.io/LVOmniBench/)] [[Paper](https://arxiv.org/abs/2603.19217)] [[Dataset](https://huggingface.co/datasets/KD-TAO/LVOmniBench)]
LVOmniBench is a new audio-visual understanding evaluation benchmark in long-form audio-video inputs. 🌟
## 🔥 News
* **`2026.03.19`** 🌟 We are very proud to launch LVOmniBench, the pioneering comprehensive evaluation benchmark of OmniLLMs in Long Audio-Video Understanding Evaluation!
## ✨ LVOmniBench Introduction
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds to 5 minutes, failing to reflect the demands of real-world applications, where videos typically run for tens of minutes. To address this critical gap, we introduce LVOmniBench, a new benchmark designed specifically for the cross-modal comprehension of long-form audio and video.
* We curated a diverse collection of long videos, with durations ranging from
**10 to 90 minutes** and an average duration of **2,069s**. This duration represents
a greater than sixfold increase in temporal scale compared to that of existing
benchmarks for audio-visual understanding.
* We **manually constructed 1,014
high-quality multiple-choice questions**, which are explicitly designed to require
joint reasoning across the audio and visual modalities, thereby facilitating a more
comprehensive evaluation of OmniLLMs.
* Each QA is ranked by difficulty level, and long audio-video understanding poses significant challenges for both current proprietary and open source models!