collaborators

5 papers

cs.CR2026

CARE: Pre-Execution Command Verification for Shell-Executing LLM Agents

Wenxiao Zhang, Yu Liu, Zhiwei Yang +7

Large Language Model (LLM) agents are increasingly used for coding and terminal automation, making shell-command dispatch a high-stakes runtime control point. We study command-leve…

cs.SD2026

When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting

Yu Liu, Zhiwei Yang, Wenxiao Zhang +8

A model can learn that the piano piece Für Elise is calm and reflective by listening to the audio or by reading a text description, but does it matter which route that knowledge t…

cs.CL2026

Does Faithfulness-Guided Alignment Hurt Accuracy? Unlocking Accurate and Faithful Post-Retrieval Reasoning

Yu Liu, Wenxiao Zhang, Diandian Guo +6

Retrieval-augmented generation (RAG) can achieve strong answer accuracy on multi-hop questions, but outcome-level rewards often leave reasoning traces weakly grounded and difficult…

cs.RO2025

Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety

Wenxiao Zhang, Xiangrui Kong, Conan Dewitt +2

Integrating large language models (LLMs) into robotic systems has revolutionised embodied artificial intelligence, enabling advanced decision-making and adaptability. However, ensu…

cs.CL2025

KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding

Bokwang Hwang, Seonkyu Lim, Taewoong Kim +24

We introduce KFinEval-Pilot, a benchmark suite specifically designed to evaluate large language models (LLMs) in the Korean financial domain. Addressing the limitations of existing…