collaborators

6 papers

cs.CV2026

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +6

Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiTs) remain insufficiently und…

cs.AI2026

Accurate and Efficient Long-Term Memory for LLM Agents

Zicheng Zhao, Xinyang Guo, Luyao Lv +3

LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructured storage loses relational context need…

cs.CV2026

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Shijie Li, Yilin Gao, Siyuan Yang +7

Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuanc…

cs.CV2026

From Priors to Perception: Grounding Video-LLMs in Physical Reality

Zicheng Zhao, Chaofan Gan, Shijie Li +1

While Video Large Language Models (Video-LLMs) excel in general understanding, they exhibit systematic deficits in fine-grained physical reasoning. Existing interventions not only…

cs.CV2025

Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers

Chaofan Gan, Zicheng Zhao, Yuanpeng Tu +5

Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for visual generation. Recent observations reveal \emph{Massive Activations} (MAs) in their internal feat…

cs.CV2025

CogStream: Context-guided Streaming Video Question Answering

Zicheng Zhao, Kangyu Wang, Shijie Li +3

Despite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual…