2 citations · 4 across the 14 of their papers we have counts for
17 papers · 1 filter
Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures
Ishani Mondal, Javad Baghirov, Jordan Boyd-Graber
Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capabili…
CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
Ishani Mondal, Yiwen Song, Mihir Parmar +4
Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene transitions. While existing gener…
SMART-Editor: A Multi-Agent Framework for Human-Like Design Editing with Structural Integrity
Ishani Mondal, Meera Bharadwaj, Ayush Roy +2
We present SMART-Editor, a framework for compositional layout and content editing across structured (posters, websites) and unstructured (natural images) domains. Unlike prior mode…
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
Zongxia Li, Yapei Chang, Yuhang Zhou +4
Evaluating open-ended long-form generation is challenging because it is hard to define what clearly separates good from bad outputs. Existing methods often miss key aspects like co…
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
Feng Gu, Zongxia Li, Carlos Rafael Colon +3
Event annotation is important for identifying market changes, monitoring breaking news, and understanding sociological trends. Although expert annotators set the gold standards, hu…
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration
Yoo Yeon Sung, Eve Fleisig, Yu Hou +2
Language models are often miscalibrated, leading to confidently incorrect answers. We introduce GRACE, a benchmark for language model calibration that incorporates comparison with…