2 citations · 3 across the 17 of their papers we have counts for
10 papers · 1 filter
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
Yu Liu, Zhilin Liu, Zhiwei Yang +7
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception,…
CARE: Pre-Execution Command Verification for Shell-Executing LLM Agents
Yu Liu, Wenxiao Zhang, Zhiwei Yang +7
Large Language Model (LLM) agents are increasingly used for coding and terminal automation, making shell-command dispatch a high-stakes runtime control point. We study command-leve…
MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
Yu Liu, Zhiwei Yang, Diandian Guo +7
Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural…
When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting
Yu Liu, Zhiwei Yang, Wenxiao Zhang +8
A model can learn that the piano piece Für Elise is calm and reflective by listening to the audio or by reading a text description, but does it matter which route that knowledge to…
Trans-RAG: Query-Centric Vector Transformation for Secure Cross-Organizational Retrieval
Yu Liu, Kun Peng, Wenxiao Zhang +4
Retrieval Augmented Generation (RAG) systems deployed across organizational boundaries face fundamental tensions between security, accuracy, and efficiency. Current encryption meth…
STIndex: A Context-Aware Multi-Dimensional Spatiotemporal Information Extraction System
Wenxiao Zhang, Yu Liu, Qiang sun +5
Extracting structured knowledge from unstructured data still faces practical limitations: entity and event extraction pipelines remain brittle, knowledge graph construction require…