20 papers
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation
Yilei Hua, Beibei Jing, Ce Zheng +3
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods…
PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization
Ying Tang, Dong Li, Youjia Zhang +3
Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent i…
Large Language Model as Token Compressor and Decompressor
Wenbing Li, Yiran Wang, Zikai Song +4
In this paper, we study whether an off-the-shelf LLM can be adapted into a discrete, variable-length token compressor and decompressor for long-context processing. To this end, we…
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
Wenbing Li, Zikai Song, Hang Zhou +3
Recent attempts to combine low-rank adaptation (LoRA) with mixture-of-experts (MoE) for multi-task adaptation of Large Language Models (LLMs) often replace whole attention/FFN laye…
CurEvo: Curriculum-Guided Self-Evolution for Video Understanding
Guiyi Zeng, Junqing Yu, Yi-Ping Phoebe Chen +3
Recent advances in self-evolution video understanding frameworks have demonstrated the potential of autonomous learning without human annotations. However, existing methods often s…
GateMOT: Q-Gated Attention for Dense Object Tracking
Mingjin Lv, Zelin Liu, Feifei Shao +4
While large models demonstrate the strong representational power of vanilla attention, this core mechanism cannot be directly applied to Dense Object Tracking: its quadratic all-to…