5 papers
GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection
Hetian Liu, Jin Cui, Mengcheng Shi +4
On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets. Coreset selection offers a practical solution…
Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models
Jin Cui, Yanbin Hu, Xinyue Long +3
Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previous…
HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models
Jin Cui, Chuanchang Su, Jiayi Lu +3
Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despit…
RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines
Xindian Ma, Xinyu Long, Yefei Zhang +11
Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entanglement, and measurement-induced nonlinear…
Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning
Jin Cui, Xinyue Long, Xunyong Zhang +5
Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence into discrete textual thoughts,…