5 papers
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
Yuanli Wang, Yaoyao Qian, Yue Zhang +8
LLM agents have made rapid progress on software engineering and ML research tasks, but these advances often assume access to a working runnable environment. For research artifacts…
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
Hanhan Zhou, Shamik Roy, Rashmi Gangadharaiah
Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive models. Controlled generation…
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
Nishanth Sridhar Nakshatri, Shamik Roy, Manoj Ghuhan Arivazhagan +3
LLMs often fail to handle temporal knowledge conflicts--contradictions arising when facts evolve over time within their training data. Existing studies evaluate this phenomenon thr…
WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
Yaoyao Qian, Yuanli Wang, Jinda Zhang +8
Current evaluation of web agents largely reduces to binary success metrics or conformity to a single reference trajectory, ignoring the structural diversity present in benchmark da…
WHEN TO ACT, WHEN TO WAIT: Modeling the Intent-Action Alignment Problem in Dialogue
Yaoyao Qian, Jindan Huang, Yuanli Wang +5
Dialogue systems often fail when user utterances are semantically complete yet lack the clarity and completeness required for appropriate system action. This mismatch arises becaus…