5 papers
Stage-adaptive Token Selection for Efficient Omni-modal LLMs
Zijie Xin, Jie Yang, Ruixiang Zhao +4
Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window…
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
Ruixiang Zhao, Jie Yang, Zijie Xin +4
Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-moda…
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
Shuhao Han, Haotian Fan, Fangyuan Kong +112
This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoratio…
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
Xinli Yue, JianHui Sun, Junda Lu +6
With the rapid advancement of text-to-image (T2I) generation models, assessing the semantic alignment between generated images and text descriptions has become a significant resear…
Self-Bootstrapping for Versatile Test-Time Adaptation
Shuaicheng Niu, Guohao Chen, Peilin Zhao +3
In this paper, we seek to develop a versatile test-time adaptation (TTA) objective for a variety of tasks - classification and regression across image-, object-, and pixel-level pr…