46 papers · 1 filter
UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy
Yicheng Xu, Jiangning Zhang, Zhucun Xue +5
In-context learning (ICL) enables fast task adaptation from demonstrations without per-task parameter updates but remains highly sensitive to example selection and formatting. In u…
FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
Bo Yin, Xiaobin Hu, Xingyu Zhou +7
Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapt large pretrained models to new tasks remains challenging. We revisit the reco…
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
Hangyu Lin, Chao Wen, Chengming Xu +4
Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-based video editing. The prevai…
JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation
Yinan Chen, Chuming Lin, Zhennan Chen +12
While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge t…
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
Jiayi Luo, Qiyan Liu, Tengyang Wang +8
Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens.…
PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
Haojun Chen, Haoyang He, Chengming Xu +11
Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imagin…