collaborators

5 papers

cs.CV2026

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

Zijie Xin, Jie Yang, Ruixiang Zhao +4

Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window…

cs.CV2026

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Ruixiang Zhao, Jie Yang, Zijie Xin +4

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-moda…

cs.CV2025

NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment

Shuhao Han, Haotian Fan, Fangyuan Kong +112

This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoratio…

cs.CV2025

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching

Xinli Yue, JianHui Sun, Junda Lu +6

With the rapid advancement of text-to-image (T2I) generation models, assessing the semantic alignment between generated images and text descriptions has become a significant resear…

cs.CV2025

Self-Bootstrapping for Versatile Test-Time Adaptation

Shuaicheng Niu, Guohao Chen, Peilin Zhao +3

In this paper, we seek to develop a versatile test-time adaptation (TTA) objective for a variety of tasks - classification and regression across image-, object-, and pixel-level pr…