111 citations · 333 across the 33 of their papers we have counts for
44 papers
Rethinking Multi-Label Image Classification With Deep Learning: Taxonomy, Challenge, and Outlook
Xuelin Zhu, Xiu-Shen Wei, Jiawei Ge +2
Multi-label image classification (MLIC), a fundamental task in computer vision, focuses on identifying multiple objects or concepts within an image, underpinning numerous read-worl…
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: From Evaluation to Diagnosis
Hong-Tao Yu, Chen-Wei Xie, Yuxin Peng +2
Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception and reasoning capabilities. While numerous benchmarks have evaluated…
Where Should Action Generation Begin? A Learnable Source Prior for Generative Robot Policies
Meipo Dai, Qiyuan Zhuang, He-Yang Xu +4
Generative robot policies typically begin action generation from an observation-independent standard Gaussian distribution, leaving the choice of source distribution underexplored.…
Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation
He-Yang Xu, Pengyuan Zhang, Zongyuan Ge +5
Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high-fidelity spatial…
Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models
Jia-Wei Hai, Yijun Wang, Xiu-Shen Wei
Vision-Language Models (VLMs), such as CLIP, have achieved significant zero-shot performance on downstream tasks with various fine-tuning adaptation methods. However, recent studie…
Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance
Lingfeng Zhang, Xiaoshuai Hao, Xizhou Bu +11
Assisting humans in open-world outdoor environments requires robots to translate high-level natural-language intentions into safe, long-horizon, and socially compliant navigation b…