5 papers
LLM-Oriented Token-Adaptive Knowledge Distillation
Xurong Xie, Zhucun Xue, Jiafu Wu +5
Knowledge distillation (KD) is a key technique for compressing large-scale language models (LLMs), yet prevailing logit-based methods typically employ static strategies that are mi…
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
Zhucun Xue, Jiangning Zhang, Teng Hu +8
The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for vide…
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
Zhucun Xue, Jiangning Zhang, Xurong Xie +4
Multimodal Large Language Models (MLLMs) perform well in video understanding but degrade on long videos due to fixed-length context and weak long-term dependency modeling. Retrieva…
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
Qingdong He, Jinlong Peng, Pengcheng Xu +8
To enhance the controllability of text-to-image diffusion models, current ControlNet-like models have explored various control signals to dictate image attributes. However, existin…
MIMAFace: Face Animation via Motion-Identity Modulated Appearance Feature Learning
Yue Han, Junwei Zhu, Yuxiang Feng +5
Current diffusion-based face animation methods generally adopt a ReferenceNet (a copy of U-Net) and a large amount of curated self-acquired data to learn appearance features, as ro…