collaborators

8 papers

cs.LG2026

GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection

Hetian Liu, Jin Cui, Mengcheng Shi +4

On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets. Coreset selection offers a practical solution…

cs.RO2026

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models

Jin Cui, Yanbin Hu, Xinyue Long +3

Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previous…

cs.CV2026

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models

Jin Cui, Chuanchang Su, Jiayi Lu +3

Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despit…

cs.CV2026

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models

Yanbin Hu, Jin Cui, Jun Ye +4

3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs co…

cs.CL2026

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

Jin Cui, Xinyue Long, Xunyong Zhang +5

Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence into discrete textual thoughts,…

cs.CL2026

MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT Distillation

Jin Cui, Jiaqi Guo, Jiepeng Zhou +6

While Large Language Models (LLMs) have emerged with remarkable capabilities in complex tasks through Chain-of-Thought reasoning, practical resource constraints have sparked intere…