most citedDeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution

2 citations · 8 across the 9 of their papers we have counts for

collaborators

14 papers

cs.CV2025

SuperCLIP: CLIP with Simple Classification Supervision

Weiheng Zhao, Zilong Huang, Jiashi Feng +1

Contrastive Language-Image Pretraining (CLIP) achieves strong generalization in vision-language tasks by aligning images and texts in a shared embedding space. However, recent find…

cs.CV20252 cited

Depth Anything 3: Recovering the Visual Space from Any Views

Haotong Lin, Sili Chen, Junhao Liew +5

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of…

cs.CV2025

Trace Anything: Representing Any Video in 4D via Trajectory Fields

Xinhang Liu, Yuxi Xiao, Donny Y. Chen +4

Effective spatio-temporal representation is fundamental to modeling, understanding, and predicting dynamics in videos. The atomic unit of a video, the pixel, traces a continuous 3D…

cs.CV2025

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Xiangtai Li, Tao Zhang, Yanwei Li +13

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack t…

cs.CV20251 cited

Seed1.5-VL Technical Report

Dong Guo, Faming Wu, Feida Zhu +194

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…

cs.CV2025

Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Tao Zhang, Xiangtai Li, Zilong Huang +6

Multimodal Large Language Models (MLLMs) achieve remarkable performance for fine-grained pixel-level understanding tasks. However, all the works rely heavily on extra components, s…