8 papers
GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection
Hetian Liu, Jin Cui, Mengcheng Shi +4
On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets. Coreset selection offers a practical solution…
Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models
Jin Cui, Yanbin Hu, Xinyue Long +3
Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previous…
HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models
Jin Cui, Chuanchang Su, Jiayi Lu +3
Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despit…
Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models
Yanbin Hu, Jin Cui, Jun Ye +4
3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs co…
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework
Jin Cui, Jiaqi Guo, Ruixuan Yang +6
Chain-of-Thought (CoT) reasoning empowers Large Language Models (LLMs) with remarkable capabilities but typically requires prohibitive parameter scales. CoT distillation has emerge…
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
Jin Cui, Xinyue Long, Xunyong Zhang +5
Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence into discrete textual thoughts,…