collaborators

12 papers

cs.CV2026

HunyuanImage 3.0 Technical Report

Tencent Hunyuan Foundation Model Team

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module pub…

cs.CV2026

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching

Chang Zou, Changlin Li, Yang Li +7

While diffusion models have achieved great success in the field of video generation, this progress is accompanied by a rapidly escalating computational burden. Among the existing a…

cs.RO2026

InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation

Junhao Cai, Zetao Cai, Jiafei Cao +39

Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, b…

cs.CV2025

UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement

Tanghui Jia, Dongyu Yan, Dehao Hao +11

In this report, we introduce UltraShape 1.0, a scalable 3D diffusion framework for high-fidelity 3D geometry generation. The proposed approach adopts a two-stage generation pipelin…

cs.CV2025

TimeWalker: Personalized Neural Space for Lifelong Head Avatars

Dongwei Pan, Yang Li, Hongsheng Li +1

We present TimeWalker, a novel framework that models realistic, full-scale 3D head avatars of a person on lifelong scale. Unlike current human head avatar pipelines that capture id…

cs.CV2025

A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning

Siyang Jiang, Mu Yuan, Xiang Ji +12

Multimodal human action recognition (HAR) leverages complementary sensors for activity classification. Beyond recognition, recent advances in large language models (LLMs) enable de…