4 papers
AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning
Sarthak Jain, Qiran Hu, Zhen Zhu +1
Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignmen…
AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation
Eric Ji, Qiran Hu, Wufei Ma +4
Synthetic data generation has emerged as a powerful tool for improving data scalability in computer vision. Recent diffusion-based pipelines have demonstrated strong photorealism.…
How to Teach Large Multimodal Models New Skills
Zhen Zhu, Yiming Gong, Yao Xiao +2
How can we teach large multimodal models (LMMs) new skills without erasing prior abilities? We study sequential fine-tuning on five target skills while monitoring general ability o…
Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference
Xiyu Guo, Shan Wang, Chunfang Ji +6
The rapid advancement of large language models (LLMs) and domain-specific AI agents has greatly expanded the ecosystem of AI-powered services. User queries, however, are highly div…