activity
20222026
most citedJudge Anything: MLLM as a Judge Across Any Modality

1 citations · 1 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV2026

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

Long Cui, Xiaoqian Liu, Qi Qin +4

Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inh…

cs.CV2026

Accelerating Masked Image Generation by Learning Controlled Latent Dynamics

Kaiwen Zhu, Quansheng Zeng, Yuandong Pu +9

Masked Image Generation Models (MIGMs) have achieved great success, yet their efficiency is hampered by the multiple steps of bi-directional attention. In fact, there exists notabl…

cs.CV2026

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

Wenhui Liao, Hongliang Li, Pengyu Xie +15

Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…

cs.CV2025

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

Shuo Cao, Jiayang Li, Xiaohui Li +12

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their abil…

cs.CV2025

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

Yi Xin, Siqi Luo, Tianxiang Xu +13

Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and ef…

cs.CV2025

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

Yi Xin, Qi Qin, Siqi Luo +29

We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by util…