3 papers
cs.CV2025
ConvRot: Rotation-Based Plug-and-Play 4-bit Quantization for Diffusion Transformers
Feice Huang, Zuliang Han, Xing Zhou +3
Diffusion transformers have demonstrated strong capabilities in generating high-quality images. However, as model size increases, the growing memory footprint and inference latency…
cs.RO2025
RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning
Yuhong Zhang, Zihan Gao, Shengpeng Li +12
We introduce Robowheel, a data engine that converts human hand object interaction (HOI) videos into training-ready supervision for cross morphology robotic learning. From monocular…
cs.CV2025
RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency
Siqi Li, Zhengkai Jiang, Jiawei Zhou +3
Virtual try-on has emerged as a pivotal task at the intersection of computer vision and fashion, aimed at digitally simulating how clothing items fit on the human body. Despite not…