4 papers · 1 filter
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
Bo Gao, Chang Liu, Yuyang Miao +2
Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an open challenge, where prior…
ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
Jiani Huang, Amish Sethi, Matthew Kuo +6
Multi-modal large language models (MLLMs) are making rapid progress toward general-purpose embodied agents. However, existing MLLMs do not reliably capture fine-grained links betwe…
AirSketch: Generative Motion to Sketch
Hui Xian Grace Lim, Xuanming Cui, Yogesh S Rawat +1
Illustration is a fundamental mode of human expression and communication. Certain types of motion that accompany speech can provide this illustrative mode of communication. While A…
A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models
Xuanming Cui, Jaiminkumar Ashokbhai Bhoi, Chionh Wei Peng +2
Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graph…