3 papers
cs.RO2026
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
Xiaofei Wu, Yi Zhang, Yumeng Liu +3
Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and…
cs.CV2025
Pack and Force Your Memory: Long-form and Consistent Video Generation
Xiaofei Wu, Guozhen Zhang, Zhiyong Xu +3
Long-form video generation presents a dual challenge: models must capture long-range dependencies while preventing the error accumulation inherent in autoregressive decoding. To ad…
cs.CV2025
GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization
Shaowen Tong, Zimin Xia, Alexandre Alahi +2
Cross-view localization, the task of estimating a camera's 3-degrees-of-freedom (3-DoF) pose by aligning ground-level images with satellite images, is crucial for large-scale outdo…