3 papers
cs.CV2026
Instruction-Based Video Editing by Repurposing an Image Editing Model
Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi +1
Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source…
cs.CV2026
DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer
Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi
Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this information loss, yet existing a…
cs.LG2026
Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
Vishnu Sarukkai, Asanshay Gupta, James Hong +2
Deploying LLM agents at scale typically requires choosing between quality and cost. Existing cost-reduction approaches fail to preserve agility: the ability to iterate rapidly with…