7 papers
TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation
Qi Liu, Gang Yue, Mingyu Yin +7
Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separate, task-specific models. Bu…
AB: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning
Yiyun Zhou, Zhonghua Jiang, Wenkang Han +4
Efficient transfer learning methods for large-scale vision-language models (, CLIP) enable strong few-shot transfer, yet existing adaptation methods follow a fixed fine-tunin…
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
Wenkang Han, Zhixiong Zeng, Jing Huang +7
Autonomous agents for Graphical User Interfaces (GUIs) are revolutionizing human-computer interaction, yet their reliance on text-based instructions imposes limitations on accessib…
Cognitive-Level Adaptive Generation via Capability-Aware Retrieval and Style Adaptation
Qingsong Wang, Tao Wu, Wang Lin +4
Large Language Models (LLMs) have demonstrated strong performance in open-ended generation tasks. However, they often struggle to adapt content to users with differing cognitive ca…
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
Wenkang Han, Wang Lin, Yiyun Zhou +4
Face Video Restoration (FVR) aims to recover high-quality face videos from degraded versions. Traditional methods struggle to preserve fine-grained, identity-specific features when…
CoLA: Collaborative Low-Rank Adaptation
Yiyun Zhou, Chang Yao, Jingyuan Chen
The scaling law of Large Language Models (LLMs) reveals a power-law relationship, showing diminishing return on performance as model scale increases. While training LLMs from scrat…