6 papers
HunyuanImage 3.0 Technical Report
Tencent Hunyuan Foundation Model Team
We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module pub…
TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing
Peizhen Zhang, Yang Li, Xunsong Li +10
Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their enormous parameter consumption…
Co-Evolving Skill Generation and Policy Optimization
Zhiwei Zhang, Yudi Lin, Nikki Lijing Kuang +4
Skill-augmented reinforcement learning improves language agents by storing reusable procedural knowledge acquired from past experience. Existing methods typically use strong langua…
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
Xuerui Qiu, Yutao Cui, Guozhen Zhang +9
Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for gener…
Multi-Head Low-Rank Attention
Songtao Liu, Hongwu Peng, Zhiwei Zhang +2
Long-context inference in large language models is bottlenecked by Key--Value (KV) cache loading during the decoding stage, where the sequential nature of generation requires repea…
HunyuanVideo: A Systematic Framework For Large Video Generative Models
Weijie Kong, Qi Tian, Zijian Zhang +49
Recent advancements in video generation have significantly impacted daily life for both individuals and industries. However, the leading video generation models remain closed-sourc…