8 papers
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
Jiajun Wu, Haoyu Kang, Yining Sun +13
Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. Howeve…
The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs
Zihan Chen, Yiming Zhang, Wenxiang Geng +2
Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benc…
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
Zihan Chen, Yiming Zhang, Hengguang Zhou +3
Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).Despite recent benchmark gains reported for RL, we find th…
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
Yiming Zhang, Zhuokai Zhao, Chengzhang Yu +8
Autoregressive large vision--language models (LVLMs) interface video and language by projecting video features into the LLM's embedding space as continuous visual token embeddings.…
FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights
Chengzhang Yu, Yiming Zhang, Zhixin Liu +3
The automation of scientific research through large language models (LLMs) presents significant opportunities but faces critical challenges in knowledge synthesis and quality assur…
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
Yiming Zhang, Zhuokai Zhao, Zhaorun Chen +3
Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains c…