37 papers
Algorithmic Recourse of In-Context Learning for Tabular Data
Wenshuo Dong, Jiaming Zhang, Shaopeng Fu +3
The paper introduces a theoretical and practical framework for providing algorithmic recourse on tabular data using in-context learning with large language models, proposing a zero…
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
Xinyan Jiang, Ninghao Liu, Di Wang +1
Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning. We introduce TRACED, a framework that assesses reasoning quality th…
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
Juangui Xu, Zikun Guo, Jingwei Lv +5
Visual language models (VLMs) have the potential to transform medical workflows. However, the deployment is limited by sycophancy. Despite this serious threat to patient safety, a…
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
Wenrui Zhou, Mohamed Hendy, Shu Yang +5
As video large language models (Video-LLMs) become increasingly integrated into real-world applications that demand grounded multimodal reasoning, ensuring their factual consistenc…
Predicting LLM Output Length via Entropy-Guided Representations
Huanyi Xie, Yubin Chen, Liangyu Wang +2
The long-tailed distribution of sequence lengths in LLM serving and reinforcement learning (RL) sampling causes significant computational waste due to excessive padding in batched…
Concept-Based Dictionary Learning for Inference-Time Safety in Vision Language Action Models
Siqi Wen, Shu Yang, Shaopeng Fu +3
Vision Language Action (VLA) models close the perception action loop by translating multimodal instructions into executable behaviors, but this very capability magnifies safety ris…