4 papers
No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
Dengyang Jiang, Mengmeng Wang, Liuzhuozheng Li +6
Recent studies have demonstrated that learning a meaningful internal representation can accelerate generative training. However, existing approaches necessitate to either introduce…
Prompt-Free Conditional Diffusion for Multi-object Image Augmentation
Haoyu Wang, Lei Zhang, Wei Wei +2
Diffusion models has underpinned much recent advances of dataset augmentation in various computer vision tasks. However, when involving generating multi-object images as real scena…
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
Mingfei Han, Liang Ma, Kamila Zhumakhanova +5
Vision-and-Language Navigation (VLN) suffers from the limited diversity and scale of training data, primarily constrained by the manual curation of existing simulators. To address…
Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning
Fei Zhou, Peng Wang, Lei Zhang +5
Meta-learning offers a promising avenue for few-shot learning (FSL), enabling models to glean a generalizable feature embedding through episodic training on synthetic FSL tasks in…