activity
20242026
most citedRADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.RO2026

Long-Horizon Manipulation via Trace-Conditioned VLA Planning

Isabella Liu, An-Chieh Cheng, Rui Yan +7

Long-horizon manipulation remains challenging for vision-language-action (VLA) policies: real tasks are multi-step, progress-dependent, and brittle to compounding execution errors.…

cs.CV2026

Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

Baifeng Shi, Stephanie Fu, Long Lian +10

Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in th…

cs.LG2025

GSPN-2: Efficient Parallel Sequence Modeling

Hongjun Wang, Yitong Jiang, Collin McCarthy +12

Efficient vision transformer remains a bottleneck for high-resolution images and long-video related real-world applications. Generalized Spatial Propagation Network (GSPN) addresse…

cs.CV2025

3D Aware Region Prompted Vision Language Model

An-Chieh Cheng, Yang Fu, Yukang Chen +10

We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flex…

cs.IR2025

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang +6

The rapid evolution of e-commerce has exposed the limitations of traditional product retrieval systems in managing complex, multi-turn user interactions. Recent advances in multimo…

cs.CV20241 cited

RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models

Greg Heinrich, Mike Ranzinger, Hongxu +6

Agglomerative models have recently emerged as a powerful approach to training vision foundation models, leveraging multi-teacher distillation from existing models such as CLIP, DIN…