collaborators

6 papers

cs.CV2025

Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling

Ziran Qin, Youru Lv, Mingbao Lin +4

Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR m…

cs.CV2025

Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens

Ziran Qin, Youru Lv, Mingbao Lin +4

Autoregressive (AR) visual generation has emerged as a powerful paradigm for image and multimodal synthesis, owing to its scalability and generality. However, existing AR image gen…

cs.LG2025

GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions

Jo-Ku Cheng, Zeren Zhang, Ran Chen +3

We propose GeoUni, the first unified geometry expert model capable of generating problem solutions and diagrams within a single framework in a way that enables the creation of uniq…

cs.CV2024

Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver

Zeren Zhang, Jo-Ku Cheng, Jingyang Deng +6

Mathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems that require both linguistic and visual signals. As the vision encoders of most…

cs.CV2024

Seismic Fault SAM: Adapting SAM with Lightweight Modules and 2.5D Strategy for Fault Detection

Ran Chen, Zeren Zhang, Jinwen Ma

Seismic fault detection holds significant geographical and practical application value, aiding experts in subsurface structure interpretation and resource exploration. Despite some…

cs.CV2024

SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

Zeren Zhang, Haibo Qin, Jiayu Huang +4

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models tog…