11 papers
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
Yixin Ji, Fanghua Ye, Juntao Li +5
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing…
AffectVerse: Emotional World Models for Multimodal Affective Computing
Bo Zhao, Fanghua Ye, Yixin Ji +3
Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, o…
When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling
Zhimin Lin, Yixin Ji, Jinpeng Li +5
Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-time scaling methods, such as…
When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning
Ruotao Xu, Yixin Ji, Yu Luo +5
Large reasoning models (LRMs) have achieved strong performance enhancement through scaling test time computation, but due to the inherent limitations of the underlying language mod…
When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning
Yang Xiang, Yixin Ji, Ruotao Xu +4
Large reasoning models (LRMs) have achieved remarkable performance in complex reasoning tasks, driven by their powerful inference-time scaling capability. However, LRMs often suffe…
GAST: Gradient-aligned Sparse Tuning of Large Language Models with Data-layer Selection
Kai Yao, Zhenghan Song, Kaixin Wu +5
Parameter-Efficient Fine-Tuning (PEFT) has become a key strategy for adapting large language models, with recent advances in sparse tuning reducing overhead by selectively updating…