7 papers
CamPilot: A Multi-Agent Cinematic Assistant for Camera-Controlled Movie Generation
Yang Wu, Stefano Petrangeli, Ishita Dasgupta +1
The integration of large language models (LLMs) into video generation has enabled rapid text-to-video creation and improved visual quality. However, it still falls short of profess…
Agentic Planning with Reasoning for Image Styling via Offline RL
Subhojyoti Mukherjee, Stefano Petrangeli, Branislav Kveton +3
Direct prompt-based editing often fails on complex transformations because vague and subjective prompts often require nuanced understanding of what should be changed in the image.…
From Pixels to Policies: Reinforcing Spatial Reasoning in Language Models for Content-Aware Layout Design
Sha Li, Stefano Petrangeli, Yu Shen +1
We introduce LaySPA, a reinforcement learning framework that equips large language models (LLMs) with explicit and interpretable spatial reasoning for content-aware graphic layout…
Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with Large Language Models
Fadlullah Raji, Stefano Petrangeli, Matheus Gadelha +3
Generating 3D models has traditionally been a complex task requiring specialized expertise. While recent advances in generative AI have sought to automate this process, existing me…
PRISM: Learning Design Knowledge from Data for Stylistic Design Improvement
Huaxiaoyue Wang, Sunav Choudhary, Franck Dernoncourt +2
Graphic design often involves exploring different stylistic directions, which can be time-consuming for non-experts. We address this problem of stylistically improving designs base…
LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation
Sha Li, Stefano Petrangeli, Yu Shen +2
While Large Language Models (LLMs) have demonstrated impressive reasoning and planning abilities in textual domains and can effectively follow instructions for complex tasks, their…