11 papers
ScaleHP: Scale-Mediated Optimization of Coupled Errors for Metric-Space Hand Pose Estimation
Ruitao Jing, Xingyu Chen, Hongyang Li +3
In this paper, we present ScaleHP, a unified framework that explicitly represents per-instance metric scale to resolve the coupled errors in calibrated camera-space hand pose estim…
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
Pengxin Xu, Xincheng Lin, Luping Xiao +5
General scene perception has progressed from object recognition toward open-vocabulary grounding, part localization, and affordance prediction. Yet these capabilities are often rea…
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
Yiran Ling, Qing Lian, Jinghang Li +6
In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to…
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
Jiazhou Zhou, Yucheng Chen, Hongyang Li +4
Multimodal Large Language Models (MLLMs) have achieved remarkable success, yet they remain prone to perception-related hallucinations in fine-grained tasks. This vulnerability aris…
SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
Guanting Ye, Qiyan Zhao, Wenhao Yu +7
3D Large Vision-Language Models (3D LVLMs) built upon Large Language Models (LLMs) have achieved remarkable progress across various multimodal tasks. However, their inherited posit…
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
Jiazhou Zhou, Qing Jiang, Kanghao Chen +4
Object detection methods have evolved from closed-set to open-set paradigms over the years. Current open-set object detectors, however, remain constrained by their exclusive relian…