Publications — research from Seldon

Research, reflections and results on building the multimodal data foundry for video-native AI.

VGI-Bench: A Holistic Benchmark for Video General Intelligence

Research · · 6 min read · Seldon Research

Introducing VGI-Bench, a holistic multimodal benchmark of 550 human-curated questions probing twelve distinct visual and audio-visual skills, built so models cannot win with language priors, single-frame shortcuts, or memorized training data.

Hard Times: A Step Toward Closing the Multimodal Data Gap

Research · · 6 min read · Seldon Research

Introducing HARD-TIME, a benchmark for long-video understanding that tests whether models ground answers in visual and audio evidence across time — targeting hallucination, sycophancy, and temporal retrieval over long videos. Frontier models drop as low as 38.9% on hard negatives.

From Minutes to Days: Agents Beat Frontier VLMs in Video QA

Research · · 12 min read · Seldon Research

Introducing MULE, which establishes a new state of the art on Video-MME-v2 (55.34, +5.94 over Gemini-3-Pro under the benchmark's grouped non-linear metric) and leads EgoLifeQA overall at 66.0 (+8.5 over the prior SOTA), winning four of five categories. No fine-tuning of the underlying VLMs.