activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

Junyi Zhang, Charles Herrmann, Junhwa Hur +5

Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or lim…

cs.CV2025

FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos

Yulu Gan, Ligeng Zhu, Dandan Shan +8

Motion understanding is fundamental to physical reasoning, enabling models to infer dynamics and predict future states. However, state-of-the-art models still struggle on recent mo…

cs.CV2025

Scaling Vision Pre-Training to 4K Resolution

Baifeng Shi, Boyi Li, Han Cai +8

High-resolution perception of visual details is crucial for daily tasks. Current vision pre-training, however, is still limited to low resolutions (e.g., 378 x 378 pixels) due to t…

cs.CV2025

Rethinking Patch Dependence for Masked Autoencoders

Letian Fu, Long Lian, Renhao Wang +6

In this work, we examine the impact of inter-patch dependencies in the decoder of masked autoencoders (MAE) on representation learning. We decompose the decoding mechanism for mask…

cs.CV2024

xT: Nested Tokenization for Larger Context in Large Images

Ritwik Gupta, Shufan Li, Tyler Zhu +3

Modern computer vision pipelines handle large images in one of two sub-optimal ways: down-sampling or cropping. These two methods incur significant losses in the amount of informat…

cs.CV2024

When Do We Not Need Larger Vision Models?

Baifeng Shi, Ziyang Wu, Maolin Mao +2

Scaling up the size of vision models has been the de facto standard to obtain more powerful visual representations. In this work, we discuss the point beyond which larger vision mo…