collaborators

22 papers

cs.CV2026

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Zizhong Ding, Junxian Li, Kai Liu +4

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive e…

cs.CV2026

TurboClear: One-Step Object-Effect Removal via Region-Calibrated Distribution Matching and Fusion

Jiawei Guo, Junxian Li, Yixin Tang +4

Recently, diffusion-based removal methods have achieved promising visual quality in removing both target objects and their associated effects. However, they typically rely on multi…

cs.CV2026

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing

Weiwei Tan, Junxian Li, Rui Wang +3

Unified multimodal models (UMMs) have recently demonstrated powerful instruction-based image editing capabilities, while also raising serious concerns about the unauthorized manipu…

cs.CR2026

SlowBA: An efficiency backdoor attack towards VLM-based GUI agents

Junxian Li, Tu Lan, Haozhen Tan +2

Modern vision-language-model (VLM) based graphical user interface (GUI) agents are expected not only to execute actions accurately but also to respond to user instructions with low…

cs.CV2026

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

Yongsen Cheng, Kai Liu, Kaiwen Tao +5

Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challenging in resource-constrained sc…

cs.CL2026

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation

Jiatong Li, Junxian Li, Weida Wang +6

Recently, Large Language Models (LLMs) have demonstrated great potential in natural language-driven molecule discovery. However, existing datasets and benchmarks for molecule-text…