4 papers
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
Minghui Pan, Jiayuxuan Yang, Yuanyuan Yuan +2
AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLM…
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
Chengyu Shen, Yanheng Hou, Minghui Pan +8
Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify approp…
Reallocating Attention Across Layers to Reduce Multimodal Hallucination
Haolang Lu, Bolun Chu, WeiYe Fu +7
Multimodal large reasoning models (MLRMs) often suffer from hallucinations that stem not only from insufficient visual grounding but also from imbalanced allocation between percept…
Streaming Hallucination Detection in Long Chain-of-Thought Reasoning
Haolang Lu, Minghui Pan, Ripeng Li +6
Long chain-of-thought (CoT) reasoning improves the performance of large language models, yet hallucinations in such settings often emerge subtly and propagate across reasoning step…