collaborators

6 papers

cs.AI2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

Yu Liu, Zhilin Liu, Zhiwei Yang +7

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception,…

cs.AI2026

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Xingming Long, Yu Liu, Zhiwei Yang +7

Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or extern…

cs.CL2026

ClueWeaver: Reward-Guided Dual-Agent Evidence Reasoning for Compact LLMs on Literary Long Narratives

Jihao Zhu, Zhiwei Yang, Wenxiao Zhang +7

Humanities and social science research requires close reading of long narrative materials such as novels, scripts, archives, and case reports, yet many users have limited access to…

cs.CR2026

CARE: Pre-Execution Command Verification for Shell-Executing LLM Agents

Yu Liu, Wenxiao Zhang, Zhiwei Yang +7

Large Language Model (LLM) agents are increasingly used for coding and terminal automation, making shell-command dispatch a high-stakes runtime control point. We study command-leve…

cs.AI2026

MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing

Yu Liu, Zhiwei Yang, Diandian Guo +7

Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural…

cs.SD2026

When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting

Yu Liu, Zhiwei Yang, Wenxiao Zhang +8

A model can learn that the piano piece Für Elise is calm and reflective by listening to the audio or by reading a text description, but does it matter which route that knowledge to…