collaborators

6 papers

cs.CV2026

HunyuanImage 3.0 Technical Report

Tencent Hunyuan Foundation Model Team

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module pub…

cs.CV2026

TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing

Peizhen Zhang, Yang Li, Xunsong Li +10

Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their enormous parameter consumption…

cs.CL2026

Co-Evolving Skill Generation and Policy Optimization

Zhiwei Zhang, Yudi Lin, Nikki Lijing Kuang +4

Skill-augmented reinforcement learning improves language agents by storing reusable procedural knowledge acquired from past experience. Existing methods typically use strong langua…

cs.CV2026

HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization

Xuerui Qiu, Yutao Cui, Guozhen Zhang +9

Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for gener…

cs.LG2026

Multi-Head Low-Rank Attention

Songtao Liu, Hongwu Peng, Zhiwei Zhang +2

Long-context inference in large language models is bottlenecked by Key--Value (KV) cache loading during the decoding stage, where the sequential nature of generation requires repea…

cs.CV2025

HunyuanVideo: A Systematic Framework For Large Video Generative Models

Weijie Kong, Qi Tian, Zijian Zhang +49

Recent advancements in video generation have significantly impacted daily life for both individuals and industries. However, the leading video generation models remain closed-sourc…