4 papers
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention
Pengyu Wang, Chenkun Tan, Shaojun Zhou +18
Video understanding is shifting from the offline paradigm -- taking a fully recorded video as input and producing a single answer after it ends -- toward real-time interaction, in…
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
Yachun Mi, Xingyang He, Shixin Sun +6
In the digital age, advanced image editing tools pose a serious threat to the integrity of visual content, making image forgery detection and localization a key research focus. Mos…
SmartThinker: Learning to Compress and Preserve Reasoning by Step-Level Length Control
Xingyang He, Xiao Ling, Jie Liu
Large reasoning models (LRMs) have exhibited remarkable reasoning capabilities through inference-time scaling, but this progress has also introduced considerable redundancy and ine…
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads
Xingyang He, Jie Liu, Shaowei Chen
KV cache is a widely used acceleration technique for large language models (LLMs) inference. However, its memory requirement grows rapidly with input length. Previous studies have…