7 papers
Visual prompt engineering for video models
Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performan…
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
Yizheng Huang, Wenjun Zeng, Aditi Kumaresan +1
Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models and benchmarks. We propose Pro…
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
Neha Kalibhat, Zi Wang, Prasoon Bajpai +4
We introduce a black-box interpretability framework that learns a verifiable constitution: a natural language summary of how changes to a prompt affect a model's specific behavior,…
ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation
Zi Wang, Xingqiao Wang, Sangah Lee +1
The rapid expansion of scholarly literature presents significant challenges in synthesizing comprehensive, high-quality academic surveys. Recent advancements in agentic systems off…
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
Belinda Z. Li, Been Kim, Zi Wang
Large language models (LLMs) have shown impressive performance on reasoning benchmarks like math and logic. While many works have largely assumed well-defined tasks, real-world que…
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
Meera Hahn, Wenjun Zeng, Nithish Kannen +4
User prompts for generative AI models are often underspecified, leading to a misalignment between the user intent and models' understanding. As a result, users commonly have to pai…