Research · · 6 min read · Seldon Research

VGI-Bench: A Holistic Benchmark for Video General Intelligence

Dithered flock of birds mid-flight, ascending in sequence across a deep indigo field, trailing streams of dotted data.

We introduce VGI-Bench (Video General Intelligence Bench), a holistic multimodal benchmark probing performance across twelve distinct visual and audio-visual skills. VGI-Bench comprises 550 human-curated questions designed to mitigate the common mistakes of today’s video benchmarks and to expose pragmatic challenges for modern state-of-the-art models. We evaluate a range of multimodal vision-language models and find that the best-performing model reaches 64.73% accuracy under our default protocol, against a 27% random-chance baseline, while humans reach 85.4%.

Vision-language models (VLMs) are today’s de-facto standard for visual reasoning tasks. They are widely applied in embodied systems such as humanoid robots[1] and self-driving cars,[2] and they serve as the visual backbones of computer-use[3] and design agents.[4] This renders the need for robust, holistic and, in particular, diagnostic benchmarks that let model developers see exactly where a model can be trusted and what remains difficult.

Where video benchmarks fall short

While existing benchmarks are available, most of them suffer from at least one of the following shortcomings:

  • Language priors Questions do not require an actual understanding of the video or audio and can be solved with strong language priors alone.[5]
  • Single-frame shortcuts Questions do not require understanding many frames in order; a single frame, or an unordered bag of frames, may suffice.[6]
  • Training-data contamination The questions may have been included in the training data, inflating the results.[7]
  • Setup artifacts The model’s design or input settings make it impossible for the model to answer, deflating the results in a non-informative way.[8]
  • Unclear real-world meaning The interpretation of the results in terms of real-world applications remains unclear.[9]

The performance

We present results for a range of widely used vision-language models. We call the models with their default settings where possible, and where a long video does not fit the context window we pass a downsampled version of it that retains as many of the relevant frames as possible. All models receive the same neutral system prompt and answer in a multiple-choice format in which the number of distractors may vary. The baseline random-chance accuracy is 27%. As a human baseline, we report 85.4%, where each question is answered by at least two distinct human reviewers and we take the average of their answers.

Horizontal bar chart ranking the evaluated models by accuracy on VGI-Bench. Gemini 3.1 Pro Preview leads at 64.73%, ahead of a 27% random-chance baseline; humans reach 85.4%. The weakest model scores 29.82%.

The live leaderboard is on Hugging Face, and the evaluation harness we used to produce these numbers is on GitHub.

Small context windows create a measurement challenge, because passing the whole video at a high frame rate is impossible, while naively downsampling the video to a uniform FPS is likely to miss the relevant information needed to answer the question. We do two things to mitigate this:

  • Short videos for dense skills Long videos do not require dense frame understanding (such as motion direction or rapidly changing details). Videos that test those skills are all shorter than 60 seconds, and every long-video question can be answered from discrete frames in sequence.
  • A 256-frame sampling floor We ensure that at a floor of 256 frames sampled uniformly across the video, each section relevant to the question is seen in at least one sampled frame.

Examples

hard-negative · transcript/visual0 models correct

When Tom says a ‘webhook could send the alert to Slack’, what color is the Slack Post button beside the URL field?

Answer This does not happen in the video.

01 / 07

The dataset

Our dataset consists of 398 public videos and programmatically generated videos, the latter to avoid models having been exposed to the test data at train time, spanning just under 100 hours.

Histogram of public YouTube video durations across the VGI-Bench dataset, excluding the hidden split and the sub-minute programmatically generated clips: 302 videos with a median of 19.0 minutes, P90 of 33.1 minutes, and a long tail out to 44.7 minutes.

Where the camera sits matters just as much as the length, because it decides which kind of system the footage resembles. We therefore classified every video by viewpoint.

Bar chart breaking the VGI-Bench dataset down by camera viewpoint. Share of the 550 questions: third-person observer 197, synthetic 105, screen computer-use 85, first-person embodied 80, mixed viewpoint 72, slides 11, across 398 videos and 99.1 hours of footage.

All videos have been annotated by human experts and we apply strict filtering to mitigate the issues discussed above. Every question in our final dataset must pass these two gates, in order:

  • Gate 1 pass^3 < 100% for a blind, text-only model.
  • Gate 2 pass^3 < 100% for an older baseline VLM. We chose Gemini 2.5 Flash-Lite.

We chose to pose questions in a hybrid format. We use a combination of public multiple-choice and hidden open-ended questions that can be submitted to prevent training data contamination.

Questions in our benchmark fall into one or more of the following skill categories:

Skills tested

Each of the 550 questions probes one or more of twelve visual and audio-visual skills. Click a skill to see per-model results on that subset.

SkillRelevance
Interacting with a persistent environment to perform long-running tasks.
Identifying relevant moments over long contexts. Recalling important events.
Understanding multi-step processes or visual instructions.
Preventing hallucination in high-risk scenarios.
Understanding multiple moments in the context of one final goal or task.
Distinguishing between multiple objects or events, even when they look very similar.
Understanding trajectories of objects in the physical world.
Keeping track of occurrences across moments and counting inside a single moment.
Navigation in the physical world.
Ignoring distractors and adversarial attacks in goal-driven scenarios.
Understanding when audio and video overlap causally and temporally.
Distinguishing left vs. right in both object orientation (hands) and directions.

Peeks into the future

We also ran models inside coding agents, which can write code and call tools, instead of prompting them directly. For Gemini 3.6 Flash this helped: driven by the Antigravity CLI it answers 77.36% of the benchmark correctly, against 63.45% as a bare call. In some smaller experiments we did see Coding Agents use OpenCV and other standard image libraries in sophisticated ways to answer questions which bare Vision Language Models could not solve.[11] The question of whether the harness might be the bottleneck for visual reasoning tasks, and the natural next question about which failure modes are harness dependent, are subject of future work.

Agents

Accuracy for the same model called directly versus driven by a coding-agent harness with tool access. GPT-5.6-sol was given the video transcript.

ModelHarnessBareIn harness
GPT-5.6-solCodexn/a65.76%
Gemini 3.6 FlashAntigravity CLI63.45%77.36%

Infinibench-v0

Some of the questions in VGI-Bench have been created by a continuously exploring agent that is rewarded for identifying new failure modes in VLMs. This agent creates the videos itself, using standard Python libraries like PyGame. The outcome is quite interesting, and our team will continue researching how to automate evaluation in this way. We find that a continuously exploring agent can create very interesting, non-trivial problems which serve as challenging out-of-distribution eval data and diagnose failure modes that were previously unknown.

Six columns of small dot clusters on a light background, changing subtly between frames.
infinibench · lossy compression3 models correct

Six groups of moving dots are shown side by side. In some, the dots mark a person’s joints and move together like a real human walking; in the rest they swing out of sync, like a jumbled puppet. How many groups show a real person walking?

Answer 4

Dark field with drifting shapes (a striped circle, a hexagon, a triangle) passing between paddle bars, with a counter in the corner.
infinibench · counting cross frame2 models correct

Which shape moves from the left side through the opening in the divider? A. The striped circle B. The triangle C. The hexagon

Answer B, the triangle

Four numbered tokens on parallel horizontal tracks, moving behind a vertical bar that hides them as they cross.
infinibench · persistent entity movement tracking2 models correct

Each ball crossing makes the sensor flash gold exactly 0.8 seconds later. Which numbered ball caused the gold flash that occurs at the exact moment ball 4 crosses the sensor?

Answer Ball 2

Field of small emoji icons (fruit, animals, instruments, vehicles) scattered across a light background.
infinibench · lossy compression2 models correct

How many of the animal icons are moving?

Answer 3

Conclusion

We constructed VGI-Bench to serve as a North Star guiding a direction for the improvement of vision models. It shows that, despite impressive progress over the last years, VLMs still lack fundamental skills needed to understand and reason about everyday situations. Consequently, pushing the frontier of embodied AI requires a map of failure modes for the strongest VLMs that is continuously expanded and updated. VGI-Bench serves as the beginning of this effort. If this research is interesting to you, or you want to use our dataset for evaluations, please reach out here.

Footnotes & Sources
  1. A. Brohan et al., RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, 2023. arxiv.org/abs/2307.15818
  2. J.-J. Hwang et al., EMMA: End-to-End Multimodal Model for Autonomous Driving, 2024. arxiv.org/abs/2410.23262
  3. W. Hong et al., CogAgent: A Visual Language Model for GUI Agents, CVPR 2024. arxiv.org/abs/2312.08914
  4. M. Xu et al., WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation, 2025. arxiv.org/abs/2511.06251
  5. L. Chen et al., Are We on the Right Way for Evaluating Large Vision-Language Models?, 2024. arxiv.org/abs/2403.20330
  6. S. Buch, C. Eyzaguirre, A. Gaidon, J. Wu, L. Fei-Fei, J. C. Niebles, Revisiting the “Video” in Video-Language Understanding, CVPR 2022. arxiv.org/abs/2206.01720
  7. B. C. Xu, L. Wu, A. Ryu, A Controlled Audit of Pretraining Contamination in Public Medical Vision-Language Benchmarks, 2026. arxiv.org/abs/2606.10066
  8. X. Fang et al., MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding, NeurIPS 2024. arxiv.org/abs/2406.14515
  9. D. Gurari et al., VizWiz Grand Challenge: Answering Visual Questions from Blind People, CVPR 2018. arxiv.org/abs/1802.08218
  10. X. Yue et al., MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark, CVPR 2024. arxiv.org/abs/2311.16502
  11. D. Chen et al., Sandboxed Coding Agents are Competitive Omni-modal Task Solvers, 2026. arxiv.org/abs/2606.00579