collaborators

5 papers

cs.CV2026

RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

Jian Zou, Xiaoyu Xu, Zhihua Wang +3

High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the resulting long visual token se…

cs.CV2026

Learning Stable Canonical Worlds for Novel View Synthesis and Beyond

Xiaoyu Xu, Jian Zou, Sheyang Tang +3

Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions. As more input views are adde…

cs.CV2026

MDS-VQA: Model-Informed Data Selection for Video Quality Assessment

Jian Zou, Xiaoyu Xu, Zhihua Wang +3

Learning-based video quality assessment (VQA) has advanced rapidly, yet progress is increasingly constrained by a disconnect between model design and dataset curation. Model-centri…

cs.CL2026

Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions

Shunyang Luo, Peibei Cao, Zhihui Zhu +3

Reward models (RMs) are central to aligning large language models, yet their practical effectiveness hinges on generalization to unseen prompts and shifting distributions. Most exi…

cs.LG2025

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

Kehua Feng, Keyan Ding, Hongzhi Tan +8

Reliable evaluation of large language models (LLMs) is impeded by two key challenges: objective metrics often fail to reflect human perception of natural language, and exhaustive h…