16 papers
Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories
Hengyu Jin, Shu Yang, Di Wang
Latent reasoning methods perform multi-step inference entirely in the model's continuous hidden states, promising more compact and efficient reasoning. However, these opaque hidden…
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
Xinyan Jiang, Ninghao Liu, Di Wang +1
Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning. We introduce TRACED, a framework that assesses reasoning quality th…
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
Juangui Xu, Zikun Guo, Jingwei Lv +5
Visual language models (VLMs) have the potential to transform medical workflows. However, the deployment is limited by sycophancy. Despite this serious threat to patient safety, a…
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
Wenrui Zhou, Mohamed Hendy, Shu Yang +5
As video large language models (Video-LLMs) become increasingly integrated into real-world applications that demand grounded multimodal reasoning, ensuring their factual consistenc…
Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
Qishun Yang, Shu Yang, Lijie Hu +1
Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods require explicit safety labels or c…
Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency
Xinyan Jiang, Wenjing Yu, Di Wang +1
Activation engineering enables precise control over Large Language Models (LLMs) without the computational cost of fine-tuning. However, existing methods deriving vectors from stat…