activity
20242026
collaborators

16 papers

cs.CV2026

Utonia: Toward One Encoder for All Point Clouds

Yujia Zhang, Xiaoyang Wu, Yunhan Yang +6

We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward…

cs.CV2026

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

Bin Hu, Yanwen Ma, Jiehui Huang +14

Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering traje…

cs.LG2026

RoPE-Aware Bit Allocation for KV-Cache Quantization

Fengfeng Liang, Yuechen Zhang, Jiaya Jia

Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contribution to a future attention logit decomposes into a position-…

cs.CV2026

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

Wenhu Zhang, Kun Cheng, Changyuan Wang +7

Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distrib…

cs.CV2026

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

Jiehui Huang, Yuechen Zhang, Bin Xia +7

Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches…

cs.CV2026

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

YaoYang Liu, Yuechen Zhang, Wenbo Li +3

High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of the input image. At 2K resolution…