collaborators

8 papers

cs.CV2026

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

Haoxiang Sun, Zhihang Yi, Langxuan Deng +6

Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in local image regions. Existin…

cs.CV2026

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

Jindi Lv, Aoyu Li, Yuhao Zhou +6

Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a sev…

cs.RO2026

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

Jindi Lv, Hao Li, Jie Li +11

Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and…

cs.CV2026

NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics

Jian Lan, Zhicheng Liu, Xinpeng Wang +5

The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related spatial intelligence tasks suc…

cs.AI2025

Multi-Path Collaborative Reasoning via Reinforcement Learning

Jindi Lv, Yuhao Zhou, Zheng Zhu +3

Chain-of-Thought (CoT) reasoning has significantly advanced the problem-solving capabilities of Large Language Models (LLMs), yet conventional CoT often exhibits internal determini…

cs.CV2025

DD-Ranking: Rethinking the Evaluation of Dataset Distillation

Zekai Li, Xinhao Zhong, Samir Khaki +49

In recent years, dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance co…