activity
20242026
collaborators

10 papers

cs.RO2026

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Kehan Li, Bohan Hou, Minghao Zhu +28

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, Ry…

cs.CV2026

Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian Splatting

Yian Zhao, Rushi Ye, Ruochong Zheng +6

3D style transfer refers to the artistic stylization of 3D assets based on reference style images. Recently, 3DGS-based stylization methods have drawn considerable attention, prima…

cs.CV2025

Comp-Attn: Present-and-Align Attention for Compositional Video Generation

Hongyu Zhang, Yufan Deng, Shenghai Yuan +5

In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main c…

cs.CV2025

Qwen3-VL Technical Report

Shuai Bai, Yuxuan Cai, Ruizhe Chen +61

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…

cs.CV2025

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Boqiang Zhang, Kehan Li, Zesen Cheng +12

In this paper, we propose VideoLLaMA3, a more advanced multimodal foundation model for image and video understanding. The core design philosophy of VideoLLaMA3 is vision-centric. T…

cs.CV2025

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Yuqian Yuan, Hang Zhang, Wentong Li +9

Video Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and st…