activity
20242026
collaborators

8 papers

cs.CV2026

VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding

Zhihao He, Tieyuan Chen, Kangyu Wang +6

Current Video Large Language Models (Video LLMs) typically encode frames via a vision encoder and employ an autoregressive (AR) LLM for understanding and generation. However, this…

cs.CV2025

Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens

Ziran Qin, Youru Lv, Mingbao Lin +4

Autoregressive (AR) visual generation has emerged as a powerful paradigm for image and multimodal synthesis, owing to its scalability and generality. However, existing AR image gen…

cs.CV2025

Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers

Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +5

Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feat…

cs.CV2025

Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning

Tieyuan Chen, Huabin Liu, Yi Wang +8

Video Question Answering (VideoQA) aims to answer natural language questions based on the given video, with prior work primarily focusing on identifying the duration of relevant se…

cs.LG2025

GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem Solutions

Jo-Ku Cheng, Zeren Zhang, Ran Chen +3

We propose GeoUni, the first unified geometry expert model capable of generating problem solutions and diagrams within a single framework in a way that enables the creation of uniq…

cs.CV2025

Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling

Ziran Qin, Youru Lv, Mingbao Lin +4

Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR m…