6 papers
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
Arda Uzunoglu, Alvin Zhang, Daniel Khashabi
Weak-to-strong generalization studies how to improve a strong student using supervision from a weaker teacher when reliable labels are scarce. We view this primarily as a data sele…
Instructional Text Across Disciplines: A Survey of Representations, Downstream Tasks, and Open Challenges Toward Capable AI Agents
Abdulfattah Safa, Tamta Kapanadze, Arda UzunoÄlu +1
Recent advances in large language models have demonstrated promising capabilities in following simple instructions through instruction tuning. However, real-world tasks often invol…
World-in-World: World Models in a Closed-Loop World
Jiahan Zhang, Muqing Jiang, Nanru Dai +14
Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive pe…
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
Arda Uzunoglu, Tianjian Li, Daniel Khashabi
Benchmarks shape scientific conclusions about model capabilities and steer model development. This creates a feedback loop: stronger benchmarks drive better models, and better mode…
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
Jiefu Ou, Arda Uzunoglu, Benjamin Van Durme +1
AI systems make decisions in physical environments through primitive actions or affordances that are accessed via API calls. While deploying AI agents in the real world involves nu…
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset
Arda Uzunoglu, Abdalfatah Rashid Safa, Gözde Gül Åahin
Recently, there has been growing interest within the community regarding whether large language models are capable of planning or executing plans. However, most prior studies use L…