3 papers
cs.CV2026
DSA: Dynamic Step Allocation for Fast Autoregressive Video Generation
Thanh-Tung Le, Yunhan Zhao, Menglei Chai +5
Video diffusion transformers have achieved state-of-the-art visual quality, but their high inference cost remains a major bottleneck for real-time applications. Recent distillation…
cs.CV2025
UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis
Thanh-Tung Le, Tuan Pham, Tung Nguyen +3
Novel view synthesis (NVS) seeks to render photorealistic, 3D-consistent images of a scene from unseen camera poses given only a sparse set of posed views. Existing deterministic n…
cs.CV2025
GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation
Tuan Pham, Thanh-Tung Le, Xiaohui Xie +1
We introduce a novel framework for metric depth estimation that enhances pretrained diffusion-based monocular depth estimation (DB-MDE) models with stereo vision guidance. While ex…