2 papers
cs.RO2026
PointAction: 3D Points as Universal Action Representations for Robot Control
Mutian Tong, Han Jiang, Qiao Feng +2
Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. How…
cs.GR2025
Spatiotemporally Consistent Indoor Lighting Estimation with Diffusion Priors
Mutian Tong, Rundi Wu, Changxi Zheng
Indoor lighting estimation from a single image or video remains a challenge due to its highly ill-posed nature, especially when the lighting condition of the scene varies spatially…