most citedJudge Anything: MLLM as a Judge Across Any Modality

1 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.AI2026

CityPlanner: A Sandbox Agent for Executable Urban Planning

Wentao Zhang, Jingyuan Wang, Zetong Zhou +2

Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service…

cs.CV2026

LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model

Jiachun Jin, Zetong Zhou, Xiao Yang +4

Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs…

cs.CV2026

Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders

Siqi Kou, Jiachun Jin, Zetong Zhou +8

Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equippe…

cs.IR20251 cited

Generative Reasoning Recommendation via LLMs

Minjie Hong, Zetong Zhou, Zirun Guo +5

Despite their remarkable reasoning capabilities across diverse domains, large language models (LLMs) face fundamental challenges in natively functioning as generative reasoning rec…

cs.CV2025

Reinforced Visual Perception with Tools

Zetong Zhou, Dongping Chen, Zixian Ma +6

Visual reasoning, a cornerstone of human intelligence, encompasses complex perceptual and logical processes essential for solving diverse visual problems. While advances in compute…

cs.CL20251 cited

Judge Anything: MLLM as a Judge Across Any Modality

Shu Pu, Yaochen Wang, Dongping Chen +10

Evaluating generative foundation models on open-ended multimodal understanding (MMU) and generation (MMG) tasks across diverse modalities (e.g., images, audio, video) poses signifi…