3 papers
cs.CV2026
Visual Implicit Autoregressive Modeling
Pengfei Jiang, Jixiang Luo, Luxi Lin +2
Visual Autoregressive Modeling (VAR) based on next-scale prediction achieves strong generation quality, but their explicit deep stacks fix the amount of computation per scale and i…
cs.CV2025
VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
Pengfei Jiang, Hanjun Li, Linglan Zhao +4
In this study, we introduce a novel method called group-wise \textbf{VI}sual token \textbf{S}election and \textbf{A}ggregation (VISA) to address the issue of inefficient inference…
cs.CV2024
Move and Act: Enhanced Object Manipulation and Background Integrity for Image Editing
Pengfei Jiang, Mingbao Lin, Fei Chao
Current methods commonly utilize three-branch structures of inversion, reconstruction, and editing, to tackle consistent image editing task. However, these methods lack control ove…