works on

From the 1 of 9 linked papers with an AI index.

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

Xin Jin, Huanqia Cai, Zhen Li +9

The paper introduces Z-Reward, a teacher‑student framework that learns to predict full rubric‑aligned score distributions for text‑to‑image generation instead of single scalar rewa…

cs.CV2026

DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation

Xifeng Xue, Xiaokang Wang, Zirui Li +2

Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as…

cs.CV2026

CrystaL: Spontaneous Emergence of Visual Latents in MLLMs

Yang Zhang, Danyang Li, Yuxuan Li +4

Multimodal Large Language Models (MLLMs) have achieved remarkable performance by integrating powerful language backbones with large-scale visual encoders. Among these, latent Chain…

cs.CV2026

Test-Time Computing for Referring Multimodal Large Language Models

Mingrui Wu, Hao Chen, Jiayi Ji +5

We propose ControlMLLM++, a novel test-time adaptation framework that injects learnable visual prompts into frozen multimodal large language models (MLLMs) to enable fine-grained r…

cs.CV2026

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

Zhong-Yu Li, Ruoyi Du, Juncheng Yan +6

Recent progress in diffusion models significantly advances various image generation tasks. However, the current mainstream approach remains focused on building task-specific models…

cs.CV2024

Multi-Token Enhancing for Vision Representation Learning

Zhong-Yu Li, Yu-Song Hu, Bo-Wen Yin +1

Vision representation learning, especially self-supervised learning, is pivotal for various vision applications. Ensemble learning has also succeeded in enhancing the performance a…