activity
20232026
most citedTinyCLIP: CLIP Distillation via Affinity Mimicking and Weight Inheritance

4 citations · 4 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CV2026

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

Yuqian Zhou, Zhenghong Zhou, Zongze Wu +5

Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video…

cs.CV2026

Aurora: Unified Video Editing with a Tool-Using Agent

Yongsheng Yu, Ziyun Zeng, Zhiyuan Xiao +4

Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and reference images, and one set o…

cs.CV2026

Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion

Zhenghong Zhou, Xiaohang Zhan, Zhiqin Chen +8

Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine-grained control remains a key bottleneck that limits practical customizability for c…

cs.CV2025

2D Triangle Splatting for Direct Differentiable Mesh Training

Kaifeng Sheng, Zheng Zhou, Yingliang Peng +1

Differentiable rendering with 3D Gaussian primitives has emerged as a powerful method for reconstructing high-fidelity 3D scenes from multi-view images. While it offers improvement…

cs.CV2025

GaussianCAD: Robust Self-Supervised CAD Reconstruction from Three Orthographic Views Using 3D Gaussian Splatting

Zheng Zhou, Zhe Li, Bo Yu +8

The automatic reconstruction of 3D computer-aided design (CAD) models from CAD sketches has recently gained significant attention in the computer vision community. Most existing me…

cs.CV2024

Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training

Zhenghong Zhou, Jie An, Jiebo Luo

Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera…