5 papers
MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
Qian Wang, Ziqi Huang, Ruoxi Jia +2
Despite recent advances, long-sequence video generation frameworks still suffer from significant limitations: poor assistive capability, suboptimal visual quality, and limited expr…
Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
Zhen Sun, Ziyi Zhang, Zeren Luo +10
Fine-grained detection and localization of localized image edits is crucial for assessing content authenticity, especially as modern diffusion models and image editors can produce…
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
Jingxing Fan, Jinrong Shen, Yusheng Yao +3
With the rapid advancement of artificial intelligence (AI), the proliferation of AI-generated content (AIGC) tasks has significantly accelerated developments in text-to-video gener…
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought
Tencent Hunyuan Team, Ao Liu, Botong Zhou +248
As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mam…
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding
Feiyu Yao, Qian Wang
As large language models (LLMs) continue to support increasingly longer contexts, the memory demand for key-value (KV) caches during decoding grows rapidly, becoming a critical bot…