activity
20242026
most citedLLaVA-SLT: Visual Language Tuning for Sign Language Translation

3 citations · 3 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CV2026

VINO: A Unified Visual Generator with Interleaved OmniModal Context

Junyi Chen, Tong He, Zhoujie Fu +3

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independen…

cs.CV2025

In-Context Audio Control of Video Diffusion Transformers

Wenze Liu, Weicai Ye, Minghong Cai +3

Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, thes…

cs.CV2025

A Reason-then-Describe Instruction Interpreter for Controllable Video Generation

Shengqiong Wu, Weicai Ye, Yuanxing Zhang +7

Diffusion Transformers have significantly improved video fidelity and temporal coherence, however, practical controllability remains limited. Concise, ambiguous, and compositionall…

cs.CV2025

Native 3D Editing with Full Attention

Weiwei Cai, Shuangkang Fang, Weicai Ye +7

Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimiza…

cs.GR2025

SketchVideo: Sketch-based Video Generation and Editing

Feng-Lin Liu, Hongbo Fu, Xintao Wang +4

Video generation and editing conditioned on text prompts or images have undergone significant advancements. However, challenges remain in accurately controlling global layout and g…

cs.CV2025

FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Xuan Ju, Weicai Ye, Quande Liu +6

Current video generative foundation models primarily focus on text-to-video tasks, providing limited control for fine-grained video content creation. Although adapter-based approac…