collaborators

6 papers

cs.CV2026

Can Text-to-Image Models Draw from the Right Frame of Reference?

Zheyuan Gu, Ruihang Li, Yong Huang +5

Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expressions are interpreted under diff…

cs.CV2026

Test-Time Curriculum for Open-Set AIGC Detection

Yiqian Zhang, Zheyuan Gu, Xiangzhao Hao +8

AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to emerge. Although existing met…

cs.CV2026

CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models

Xiangzhao Hao, Zefeng Zhang, Zhenyu Zhang +6

Image degradation from blur, noise, compression, and poor illumination severely undermines multimodal understanding in real-world settings. Unified multimodal models that combine u…

cs.CV2025

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang +8

Visual Spatial Reasoning is crucial for enabling Multimodal Large Language Models (MLLMs) to understand object properties and spatial relationships, yet current models still strugg…

cs.CV2025

Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization

Zefeng Zhang, Hengzhu Tang, Jiawei Sheng +6

Multimodal Large Language Models excel in various tasks, yet often struggle with modality bias, where the model tends to rely heavily on a single modality and overlook critical inf…

cs.CV2025

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search

Hengzhu Tang, Zefeng Zhang, Zhiping Li +5

Video Quality Assessment (VQA) is vital for large-scale video retrieval systems, aimed at identifying quality issues to prioritize high-quality videos. In industrial systems, low-q…