8 citations · 14 across the 11 of their papers we have counts for
4 papers · 1 filter
Pretrained Persona Mixture Models and Tandem Models for Human Simulation
Minwoo Kang, Téa Wright, Seun Eisape +6
We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces st…
Decoupling Planning and Control for Instructable Agents
Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste +2
Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to…
: Unifying Generation and Self-Verification for Parallel Reasoners
Harman Singh, Xiuyu Li, Kusha Sareen +14
Test-time scaling for complex reasoning tasks shows that leveraging inference-time compute, by methods such as independently sampling and aggregating multiple solutions, results in…
Visually Prompted Benchmarks Are Surprisingly Fragile
Haiwen Feng, Long Lian, Lisa Dunlap +6
A key challenge in evaluating VLMs is testing models' ability to analyze visual content independently from their textual priors. Recent benchmarks such as BLINK probe visual percep…