activity
20242026
most citedXModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

1 citations · 1 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun +5

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…

cs.CV2026

VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking

Jingyang Lin, Jialian Wu, Jiang Liu +6

Video agentic models have advanced challenging video-language tasks. However, most agentic approaches still heavily rely on greedy parsing over densely sampled video frames, result…

cs.CV2025

Learning from Online Videos at Inference Time for Computer-Use Agents

Yujian Liu, Ze Wang, Hao Chen +7

Computer-use agents can operate computers and automate laborious tasks, but despite recent rapid progress, they still lag behind human users, especially when tasks require domain-s…

cs.CV2025

CaptionQA: Is Your Caption as Useful as the Image Itself?

Shijia Yang, Yunong Liu, Bohan Zhai +5

Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic inference pipelines. Yet current eva…

cs.CV2025

Latent Visual Reasoning

Bangzheng Li, Ximeng Sun, Jiang Liu +7

Multimodal Large Language Models (MLLMs) have achieved notable gains in various tasks by incorporating Chain-of-Thought (CoT) reasoning in language spaces. Recent work extends this…

cs.CV20251 cited

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

Xingrui Wang, Jiang Liu, Chao Huang +7

Omni-modal large language models (OLLMs) aim to unify audio, vision, and text understanding within a single framework. While existing benchmarks primarily evaluate general cross-mo…