4 papers
TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning
Binnan Liu, Yechi Ma, Tian Xie +1
The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid. Looped visual reaso…
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
Tao Xie, Peishan Yang, Yudong Jin +8
This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly r…
Roadside Monocular 3D Detection Prompted by 2D Detection
Yechi Ma, Yanan Li, Wei Hua +1
Roadside monocular 3D detection requires detecting objects of predefined classes in an RGB frame and predicting their 3D attributes, such as bird's-eye-view (BEV) locations. It has…
Long-Tailed 3D Detection via Multi-Modal Fusion
Yechi Ma, Neehar Peri, Achal Dave +3
Contemporary autonomous vehicle (AV) benchmarks have advanced techniques for training 3D detectors. While class labels naturally follow a long-tailed distribution in the real world…