most citedGenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Xinming Wang, Weinong Wang, Hongming Yang +13

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although thes…

cs.CV2026

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

Mingqi Gao, Hongyuan Dong, Yifei Chen +6

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense vi…

cs.CV20261 cited

GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?

Ruihang Li, Leigang Qu, Jingxu Zhang +6

The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this…

cs.CV2026

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

Wenjin Hou, Wei Liu, Han Hu +3

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. Howeve…

cs.CV2026

Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal Regions

Kaiwen Zheng, Junchen Fu, Songpei Xu +4

In this paper, we introduce an underexplored problem in facial analysis: generating and recognizing multi-attribute natural language descriptions, containing facial action units (A…

cs.CV2025

GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization

Yikun Wang, Zuyan Liu, Ziyi Wang +3

Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agen…