24 papers
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
Haoqing Wang, Xingrun Xing, Wei Xia +2
FaithEyes proposes a multi‑agent framework where a vision‑language model judges its own tool calls to ensure they are useful, improving both accuracy and tool faithfulness on visua…
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4
Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued pretraining on long-context data can help but is…
PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback
Wei Xia, Jin Wu, Haoran Shi +2
Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback. However, existing approaches often treat scoring and…
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns
Vlad Sobal, Shuo Yang, Yuting Zhang +2
We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing. Unlike t…
Reinforcement-aware Knowledge Distillation for LLM Reasoning
Zhaoyang Zhang, Shuli Jiang, Yantao Shen +6
Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such model…
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
Wei Xia, Haoqing Wang, Zhi-Hong Deng +1
Chain-of-thought (CoT) reasoning has become the default strategy for enhancing LLM capabilities, yet its application raises a fundamental question: when is explicit reasoning actua…