collaborators

6 papers

cs.RO2025

L1 Sample Flow for Efficient Visuomotor Learning

Weixi Song, Zhetao Chen, Tao Xu +6

Denoising-based models, such as diffusion and flow matching, have been a critical component of robotic manipulation for their strong distribution-fitting and scaling capacity. Conc…

cs.CV2025

Robust Source-Free Domain Adaptation for Medical Image Segmentation based on Curriculum Learning

Ziqi Zhang, Yuexiang Li, Yawen Huang +6

Recent studies have uncovered a new research line, namely source-free domain adaptation, which adapts a model to target domains without using the source data. Such a setting can ad…

cs.LG2025

Quantum-Boosted High-Fidelity Deep Learning

Feng-ao Wang, Shaobo Chen, Yao Xuan +12

A fundamental limitation of probabilistic deep learning is its predominant reliance on Gaussian priors. This simplistic assumption prevents models from accurately capturing the com…

cs.RO2025

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning

Sheng Chen, Peiyu He, Jiaxin Hu +67

Modern robot navigation systems encounter difficulties in diverse and complex indoor environments. Traditional approaches rely on multiple modules with small models or rule-based s…

cs.CV2025

Movie Gen: A Cast of Media Foundation Models

Adam Polyak, Amit Zohar, Andrew Brown +85

We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabili…

cs.SD2024

Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio

Xinrui Yan, Jiangyan Yi, Jianhua Tao +6

Open environment oriented open set model attribution of deepfake audio is an emerging research topic, aiming to identify the generation models of deepfake audio. Most previous work…