8 papers
When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models
Yufei Zhang, Chenlu Zhan, Hongwei Wang
Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechanistically poorly understood.…
SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance
Yufei Zhang, Chenlu Zhan, Donghui Sun +2
Affordance grounding aims to localize the functional region for interaction, such as the handle to grasp or the button to press, rather than the whole object. This makes it more ch…
SAD-GS: Learning Reliable 3D Semantic Gaussian Fields via Dynamic Geo-Semantic Anchoring
Yufei Zhang, Chenlu Zhan, Gaoang Wang +1
Open-vocabulary 3D semantic Gaussian field learning relies on multi-view 2D supervision, whose semantic targets and spatial assignments are often unreliable. Across varying viewpoi…
FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding
Chenlu Zhan, Yufei Zhang, Gaoang Wang +1
Semantic querying in complex 3D scenes through free-form language presents a significant challenge. Existing 3D scene understanding methods use large-scale training data and CLIP t…
Hi-LSplat: Hierarchical 3D Language Gaussian Splatting
Chenlu Zhan, Yufei Zhang, Gaoang Wang +1
Modeling 3D language fields with Gaussian Splatting for open-ended language queries has recently garnered increasing attention. However, recent 3DGS-based models leverage view-depe…
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
Hanrong Zhang, Jingyuan Huang, Kai Mei +5
Although LLM-based agents, powered by Large Language Models (LLMs), can use external tools and memory mechanisms to solve complex real-world tasks, they may also introduce critical…