activity
20242026
collaborators

16 papers

cs.CL2026

Does Accuracy Equal Evidence? Reasoning Faithfulness under KV Cache Compression

Mengting Ai, Jingrui He, Yue Guo

KV cache compression is commonly evaluated by final-answer accuracy, implicitly assuming that preserving the answer also preserves the reasoning that supports it. We test this assu…

cs.LG2026

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

Xinrui He, Mengting Ai, Junting Wang +2

Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness. While many im…

cs.CL2026

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

Jiaru Zou, Ling Yang, Yunzhe Qi +5

Agentic reinforcement learning has advanced large language models (LLMs) to reason through long chain-of-thought trajectories while interleaving external tool use. Existing approac…

cs.AI2026

Harnessing Generalist Agents for Contextualized Time Series

Zihao Li, Kaifeng Jin, Yuanchen Bei +8

Time series are often embedded in rich contexts that are essential for holistic modeling. Moreover, real-world practitioners often require end-to-end workflows for analyzing tempor…

cs.CL2026

Code as Agent Harness

Xuying Ning, Katherine Tieu, Dongqi Fu +39

Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineerin…

cs.CL2026

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Tianxin Wei, Noveen Sachdeva, Benjamin Coleman +12

Statefulness is essential for large language model (LLM) agents to perform long-term planning and problem-solving. This makes memory a critical component, yet its management and ev…