3 citations · 6 across the 5 of their papers we have counts for
5 papers · 1 filter
HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models
Feng He, Zhenting Wang, Qifan Wang +4
Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence. Prior work mainly focuses o…
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
Xinyan Cai, Shiguang Wu, Dafeng Chi +4
In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imag…
Radiance Field Learners As UAV First-Person Viewers
Liqi Yan, Qifan Wang, Junhan Zhao +4
First-Person-View (FPV) holds immense potential for revolutionizing the trajectory of Unmanned Aerial Vehicles (UAVs), offering an exhilarating avenue for navigating complex buildi…
AMD: Automatic Multi-step Distillation of Large-scale Vision Models
Cheng Han, Qifan Wang, Sohail A. Dianat +6
Transformer-based architectures have become the de-facto standard models for diverse vision tasks owing to their superior performance. As the size of the models continues to scale…
Prototypical Transformer as Unified Motion Learners
Cheng Han, Yawen Lu, Guohao Sun +9
In this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoForme…