activity
20242026
collaborators

10 papers

cs.LG2026

Post-training Quantization for Hybrid Iterative Generative Models

Jing Gao, Junyi Wu, Wei Wang +2

Iterative Generative Models (IGMs) span autoregressive and diffusion paradigms, and hybrid variants that couple them can achieve remarkable image-generation fidelity. However, thei…

cs.CV2026

Inline Critic Steers Image Editing

Weitai Kang, Xiaohang Zhan, Yizhou Wang +4

Instruction-based image editing exhibits heterogeneous difficulty not only across cases but also across regions of an image, motivating refinement approaches that allocate correcti…

cs.AI2026

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

Bin Lei, Weitai Kang, Zijian Zhang +8

This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlik…

cs.AI2026

Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge

Bo Pan, Xuan Kan, Kaitai Zhang +6

Large language models (LLMs) have become widely adopted as automated judges for evaluating AI-generated content. Despite their success, aligning LLM-based evaluations with human ju…

cs.CV2025

VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction

Weitai Kang, Jason Kuen, Mengwei Ren +3

Current visual grounding models are either based on a Multimodal Large Language Model (MLLM) that performs auto-regressive decoding, which is slow and risks hallucinations, or on r…

cs.CV2025

ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model

Weitai Kang, Weiming Zhuang, Zhizhong Li +2

Fine-grained multimodal capability in Multimodal Large Language Models (MLLMs) has emerged as a critical research direction, particularly for tackling the visual grounding (VG) pro…