4 papers
Concentrate After Imagination: Text-Conditioned Evidence Grounding for Partially Relevant Video Retrieval
Shuaiqi Cheng, Siyu You, Yanbi Wu +3
Partially Relevant Video Retrieval (PRVR) retrieves untrimmed videos when queries describe only short moments. Although recent methods improve local representations, uncertainty mo…
PhysDox: Benchmarking LLMs on Physical Feasibility Auditing of Physiological Sensing Protocols
He Liu, Boyuan Gu, Shuaiqi Cheng +3
Large language models (LLMs) increasingly assist in experimental design, yet fluent protocols often remain physically infeasible. We introduce PhysDox, a physical feasibility audit…
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
Fanpu Cao, Xin Zou, Xuming Hu +1
Multimodal large language models (MLLMs) have become a key interface for visual reasoning and grounded question answering, yet they remain vulnerable to visual hallucinations, wher…
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
Yuhang Han, Yuyang Wu, Zhengbo Jiao +6
Reinforcement Learning from Verifiable Rewards (RLVR) has substantially enhanced the reasoning capabilities of large language models in abstract reasoning tasks. However, its appli…