collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

VideoSketcher: Sequential Sketch Generation Using Video Model Priors

Hui Ren, Yuval Alaluf, Omer Bar Tal +3

Sketching is inherently sequential: strokes are drawn progressively to explore and refine ideas. Yet most generative approaches treat sketches as static images, ignoring the tempor…

cs.CV2026

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion

Zhenggang Tang, Yuehao Wang, Yuchen Fan +9

Recent text-to-scene generation approaches largely reduced the manual efforts required to create 3D scenes. However, their focus is either to generate a scene layout or to generate…

cs.CV2025

Spatio-Temporal LLM: Reasoning about Environments and Actions

Haozhen Zheng, Beitong Tian, Mingyuan Wu +3

Despite significant recent progress of Multimodal Large Language Models (MLLMs), current MLLMs are challenged by "spatio-temporal" prompts, i.e., prompts that refer to 1) the entir…

cs.CV2025

Hierarchical Rectified Flow Matching with Mini-Batch Couplings

Yichi Zhang, Yici Yan, Alex Schwing +1

Flow matching has emerged as a compelling generative modeling approach that is widely used across domains. To generate data via a flow matching model, an ordinary differential equa…

cs.CV2025

DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation

Chen Chen, Rui Qian, Wenze Hu +8

In this work, we empirically study Diffusion Transformers (DiTs) for text-to-image generation, focusing on architectural choices, text-conditioning strategies, and training protoco…