5 papers
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
Juncheng Ma, Jianxin Bi, Yufan Deng +19
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remai…
Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving
Kewei Zhang, Jin Wang, Sensen Gao +9
End-to-end autonomous driving via Vision-Language-Action (VLA) models demands a precarious balance between high-fidelity trajectory planning and efficient inference. Existing parad…
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
Kewei Zhang, Ye Huang, Yufan Deng +5
While the Transformer architecture dominates many fields, its quadratic self-attention complexity hinders its use in large-scale applications. Linear attention offers an efficient…
DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution
Zheng Chen, Zichen Zou, Kewei Zhang +4
Diffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely sl…
QuantDemoire: Quantization with Outlier Aware for Image Demoiréing
Zheng Chen, Kewei Zhang, Xiaoyang Liu +4
Demoiréing aims to remove moiré artifacts that often occur in images. While recent deep learning-based methods have achieved promising results, they typically require substantial…