activity
20242026
most citedMedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

Yufei Zhang, Chenlu Zhan, Hongwei Wang

Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechanistically poorly understood.…

cs.CV2026

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

Yufei Zhang, Chenlu Zhan, Donghui Sun +2

Affordance grounding aims to localize the functional region for interaction, such as the handle to grasp or the button to press, rather than the whole object. This makes it more ch…

cs.CV2026

SAD-GS: Learning Reliable 3D Semantic Gaussian Fields via Dynamic Geo-Semantic Anchoring

Yufei Zhang, Chenlu Zhan, Gaoang Wang +1

Open-vocabulary 3D semantic Gaussian field learning relies on multi-view 2D supervision, whose semantic targets and spatial assignments are often unreliable. Across varying viewpoi…

cs.CV2026

See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent

Tianci Tang, Tielong Cai, Hongwei Wang +1

Pre-trained perception models excel in generic image domains but degrade significantly in novel environments like indoor scenes. The conventional remedy is fine-tuning on downstrea…

cs.CV2026

DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target

BoCheng Hu, Zhonghan Zhao, Kaiyue Zhou +2

Most existing hand motion generation benchmarks for hand-object interaction (HOI) focus on static objects, leaving dynamic scenarios with moving targets and time-critical coordinat…

cs.CV2025

Understanding Dynamic Scenes in Ego Centric 4D Point Clouds

Junsheng Huang, Shengyu Hao, Bocheng Hu +2

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, a…