6 papers
Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
Zirui Song, Huaxing Liu, Xiang Wang +8
Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in their outpu…
CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding
Wei Jia, Zhicong Lu, Yu Chen +5
Large vision-language models (LVLMs) have achieved substantial performance gains in Video Temporal Grounding (VTG) through reinforcement learning (RL). However, existing methods pr…
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Zixiang Xu, Sixian Li, Huaxing Liu +4
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argu…
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
Jiawei Guo, Donglei Yu, Yu Chen +6
Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interpretation can be premature: a t…
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
Jiawei Guo, Yu Chen, Xiang Wang +6
Recent latent visual reasoning methods achieve substantial gains by inserting continuous latent tokens into multimodal language models. These gains are commonly attributed to the t…
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
Changyuan Tian, Zhicong Lu, Huaxing Liu +7
Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and recent work extends RLVR to…