collaborators

6 papers

cs.CV2025

Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models

Shengchao Zhou, Yuxin Chen, Yuying Ge +4

Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relation…

cs.CV2025

EmbRACE-3K: Embodied Reasoning and Action in Complex Environments

Mingxian Lin, Wei Huang, Yitang Li +6

Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied…

cs.CV2025

2D Instance Editing in 3D Space

Yuhuan Xie, Aoxuan Pan, Ming-Xian Lin +3

Generative models have achieved significant progress in advancing 2D image editing, demonstrating exceptional precision and realism. However, they often struggle with consistency a…

cs.CV2025

Scaling RL to Long Videos

Yukang Chen, Wei Huang, Baifeng Shi +11

We introduce a full-stack framework that scales up reasoning in vision-language models (VLMs) to long videos, leveraging reinforcement learning. We address the unique challenges of…

cs.LG2025

DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization

Zijian Ye, Wei Huang, Yifei Yu +3

Large language models (LLMs) demonstrate remarkable performance but face substantial computational and memory challenges that limit their practical deployment. Quantization has eme…

cs.CV2024

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

Songhao Han, Wei Huang, Hairong Shi +7

The advancement of Large Vision Language Models (LVLMs) has significantly improved multimodal understanding, yet challenges remain in video reasoning tasks due to the scarcity of h…