7 papers
Optimality of Gradient-MUSIC for spectral estimation
Albert Fannjiang, Weilin Li, Wenjing Liao
We introduce the Gradient-MUSIC algorithm for estimating the unknown frequencies and amplitudes of a nonharmonic signal from noisy time samples. While the classical MUSIC algorithm…
GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth
Yuecheng Liu, Junda Cheng, Longliang Liu +4
Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail region…
Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods
Zhaiming Shen, Alexander Hsu, Rongjie Lai +1
While in-context learning (ICL) has achieved remarkable success in natural language and vision domains, its theoretical understanding-particularly in the context of structured geom…
Transformers for Learning on Noisy and Task-Level Manifolds: Approximation and Generalization Insights
Zhaiming Shen, Alex Havrilla, Rongjie Lai +2
Transformers serve as the foundational architecture for large language and video generation models, such as GPT, BERT, SORA and their successors. Empirical studies have demonstrate…
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
Zhongjie Shi, Wenjing Liao
This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain and -dimensional compact Riemannian manifolds.…
MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
Junda Cheng, Wenjing Liao, Zhipeng Cai +10
We introduce MonSter++, a geometric foundation model for multi-view depth estimation, unifying rectified stereo matching and unrectified multi-view stereo. Both tasks fundamentally…