collaborators

7 papers

cs.CL2026

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

Jiaming Wei, Zekun Wu, Adriano Koshiyama +1

Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combination…

cs.LG2026

Training Large Language Models for Self-Explanation Faithfulness

Yeoktatt Cheah, María Pérez-Ortiz, Noah Y. Siegel +1

We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects i…

cs.CL2026

Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

Noah Y. Siegel, Nicolas Heess, Maria Perez-Ortiz +1

When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans. But are these explanations faithful, i.e. do they convey the factors actual…

cs.CL2026

Tool Calling is Linearly Readable and Steerable in Language Models

Zekun Wu, Ze Wang, Seonglae Cho +4

When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. As agents take on consequential actions, one…

cs.CL2026

Sell Me This Stock: Unsafe Recommendation Drift in LLM Agents

Zekun Wu, Adriano Koshiyama, Sahan Bulathwela +1

People increasingly use LLM agents for multi-turn financial recommendations, where the agent pulls market data through tools and tracks user preferences across turns. When tool out…

cs.CL2024

The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models

Noah Y. Siegel, Oana-Maria Camburu, Nicolas Heess +1

In order to oversee advanced AI systems, it is important to understand their underlying decision-making process. When prompted, large language models (LLMs) can provide natural lan…