10 papers
Post-training Quantization for Hybrid Iterative Generative Models
Jing Gao, Junyi Wu, Wei Wang +2
Iterative Generative Models (IGMs) span autoregressive and diffusion paradigms, and hybrid variants that couple them can achieve remarkable image-generation fidelity. However, thei…
Inline Critic Steers Image Editing
Weitai Kang, Xiaohang Zhan, Yizhou Wang +4
Instruction-based image editing exhibits heterogeneous difficulty not only across cases but also across regions of an image, motivating refinement approaches that allocate correcti…
InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction
Bin Lei, Weitai Kang, Zijian Zhang +8
This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlik…
Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge
Bo Pan, Xuan Kan, Kaitai Zhang +6
Large language models (LLMs) have become widely adopted as automated judges for evaluating AI-generated content. Despite their success, aligning LLM-based evaluations with human ju…
VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
Weitai Kang, Jason Kuen, Mengwei Ren +3
Current visual grounding models are either based on a Multimodal Large Language Model (MLLM) that performs auto-regressive decoding, which is slow and risks hallucinations, or on r…
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
Weitai Kang, Weiming Zhuang, Zhizhong Li +2
Fine-grained multimodal capability in Multimodal Large Language Models (MLLMs) has emerged as a critical research direction, particularly for tackling the visual grounding (VG) pro…