13 papers
BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training
Zili Zhang, Chengxu Yang, Shenglong Zhang +8
Training multimodal large language models (MLLMs) is challenged by both model and data heterogeneity. Existing systems redesign the training pipeline to address these challenges, b…
Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs
Wenjie Zhu, Yabin Zhang, Liang Xu +3
While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in…
An Efficient Streaming Video Understanding Framework with Agentic Control
Jinming Liu, Jianguo Huang, Zhaoyang Jia +7
Streaming video requires handling dynamic information density under strict latency budgets. Yet, existing methods typically employ static strategies, such as fixed memory compressi…
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
Siyong Jian, Siyuan Li, Luyuan Zhang +5
Discrete autoregressive (AR) text-to-image (T2I) models pair a VQ tokenizer with an AR policy, and current post-training pipelines optimize only the policy while keeping the VQ dec…
Generation Navigator: A State-Aware Agentic Framework for Image Generation
Jinming Liu, Ruoyu Feng, Yuqi Wang +2
Despite rapid advances in text-to-image generation, faithfully realizing user intent remains challenging, often requiring manual multi-turn trial and error. To automate this proces…
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
Yu Shang, Yinzhou Tang, Yiding Ma +22
World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…