activity
20242026
collaborators

7 papers

cs.CV2026

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

Zirui Zhang, Yinbo Yu, Donghai Guan +3

The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years. Compared with early generati…

cs.CV2026

SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation

Shilong Xiang, Zirui Zhang, Lijun Yu +1

Autoregressive models excel in visual generation by treating images as 1D sequences of discrete tokens, mirroring language modeling. However, this flattening discards the intrinsic…

cs.AI2026

R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning

Zirui Zhang, Haoyu Dong, Kexin Pei +1

Robust perception and reasoning require consistency across sensory modalities. Yet current multimodal models often violate this principle, yielding contradictory predictions for vi…

cs.LG2025

Reactivation: Empirical NTK Dynamics Under Task Shifts

Yuzhi Liu, Zixuan Chen, Zirui Zhang +2

The Neural Tangent Kernel (NTK) offers a powerful tool to study the functional dynamics of neural networks. In the so-called lazy, or kernel regime, the NTK remains static during t…

cs.CL2025

Yi-Lightning Technical Report

Alan Wake, Bei Chen, C. X. Lv +41

This technical report presents Yi-Lightning, our latest flagship large language model (LLM). It achieves exceptional performance, ranking 6th overall on Chatbot Arena, with particu…

cs.SD2024

I Can Hear You: Selective Robust Training for Deepfake Audio Detection

Zirui Zhang, Wei Hao, Aroon Sankoh +4

Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we…