collaborators

7 papers

cs.CV2026

CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering

Mouxiao Huang, Qiangyu Yan, Borui Jiang +1

Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (e.g., CIDEr and SPICE) and L…

cs.CL2026

Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation

Xuan Du, Qiangyu Yan, Wenshuo Li +4

The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, ta…

cs.CV2026

EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution

Haizhen Xie, Kunpeng Du, Qiangyu Yan +5

Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. While T2I models have traditionally…

cs.CV2026

Multimodal Latent Reasoning via Hierarchical Visual Cues Injection

Yiming Zhang, Qiangyu Yan, Borui Jiang +1

The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often remains a "fast thinking" paradig…

cs.CV2026

GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection

Zhenliang Ni, Qiangyu Yan, Mouxiao Huang +5

The rapid advancement of video generation models has made it increasingly challenging to distinguish AI-generated videos from real ones. This issue underscores the urgent need for…

cs.CV2025

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Yiman Zhang, Ziheng Luo, Qiangyu Yan +4

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with exis…