activity
20242026
collaborators

7 papers

cs.CV2026

FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion

Zhuokun Chen, Jianfei Cai, Bohan Zhuang

Generating long-form content, such as minute-long videos and extended texts, is increasingly important for modern generative models. Block diffusion improves inference efficiency v…

cs.LG2026

R-Stitch: Dynamic Trajectory Stitching for Efficient Reasoning

Zhuokun Chen, Zeren Chen, Jiahao He +4

Chain-of-thought (CoT) enhances the problem-solving ability of large language models (LLMs) but incurs substantial inference cost due to long autoregressive trajectories. Existing…

cs.AI2025

Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning

Ziyu Ma, Chenhui Gou, Yiming Hu +4

Large Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to many-shot settings remains difficult due to limited context length and hi…

cs.CV2025

An Empirical Study on How Video-LLMs Answer Video Questions

Chenhui Gou, Ziyu Ma, Zicheng Duan +6

Taking advantage of large-scale data and pretrained language models, Video Large Language Models (Video-LLMs) have shown strong capabilities in answering video questions. However,…

cs.CV2024

MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views

Yuedong Chen, Chuanxia Zheng, Haofei Xu +4

We introduce MVSplat360, a feed-forward approach for 360° novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-p…

cs.CL2024

ME-Switch: A Memory-Efficient Expert Switching Framework for Large Language Models

Jing Liu, Ruihao Gong, Mingyang Zhang +3

LLM development involves pre-training a foundation model on massive data, followed by fine-tuning on task-specific data to create specialized experts. Serving these experts can pos…