activity
20232026
most citedTeG-DG: Textually Guided Domain Generalization for Face Anti-Spoofing

1 citations · 1 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2026

Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

Jiedong Zhuang, Lu Lu, Ming Dai +4

Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual to…

cs.CV2025

Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting

Jiangnan Ye, Jiedong Zhuang, Lianrui Mu +5

We introduce GS-Light, an efficient, textual position-aware pipeline for text-guided relighting of 3D scenes represented via Gaussian Splatting (3DGS). GS-Light implements a traini…

cs.CV2025

No Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image Detection

Lianrui Mu, Zou Xingze, Jianhong Bai +7

The rapid growth of high-resolution, meticulously crafted AI-generated images poses a significant challenge to existing detection methods, which are often trained and evaluated on…

cs.CV2024

ST: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

Jiedong Zhuang, Lu Lu, Ming Dai +4

Multimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual token…

cs.CV2024

FashionR2R: Texture-preserving Rendered-to-Real Image Translation with Diffusion Models

Rui Hu, Qian He, Gaofeng He +4

Modeling and producing lifelike clothed human images has attracted researchers' attention from different areas for decades, with the complexity from highly articulated and structur…

cs.CV2024

Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding

Xiaoyu Liang, Jiayuan Yu, Lianrui Mu +7

Although Visual-Language Models (VLMs) have shown impressive capabilities in tasks like visual question answering and image captioning, they still struggle with hallucinations. Ana…