4 papers · 1 filter
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
Xinyan Jiang, Ninghao Liu, Di Wang +1
Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning. We introduce TRACED, a framework that assesses reasoning quality th…
PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
Manjiang Yu, Hongji Li, Priyanka Singh +3
Reliable behavior control is central to deploying large language models (LLMs) on the web. Activation steering offers a tuning-free route to align attributes (e.g., truthfulness) t…
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
Zhuoran Zhang, Tengyue Wang, Xilin Gong +4
Multimodal large language models (MLLMs) must resolve conflicts when different modalities provide contradictory information, a process we term modality following. Prior work measur…
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
Liangliang You, Junchi Yao, Shu Yang +3
While multimodal large language models excel at various tasks, they still suffer from hallucinations, which limit their reliability and scalability for broader domain applications.…