4 papers
Find, Fix, Reason: Context Repair for Video Reasoning
Haojian Huang, Chuanyu Qin, Yinchuan Li +1
Reinforcement learning has advanced video reasoning in large multi-modal models, yet dominant pipelines either rely on on-policy self-exploration, which plateaus at the model's kno…
STAR: Mitigating Cascading Errors in Spatial Reasoning via Turn-point Alignment and Segment-level DPO
Pukun Zhao, Longxiang Wang, Chen Chen +4
Structured spatial navigation is a core benchmark for Large Language Models (LLMs) spatial reasoning. Existing paradigms like Visualization-of-Thought (VoT) are prone to cascading…
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
Chen Chen, Runze Li, Zejun Zhang +4
We address multimodal deepfake detection requiring both robustness and interpretability by proposing FakeHunter, a unified framework that combines memory guided retrieval, a struct…
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
Pukun Zhao, Longxiang Wang, Miaowei Wang +3
Most existing spatial reasoning benchmarks focus on static or globally observable environments, failing to capture the challenges of long-horizon reasoning and memory utilization u…