3 papers
cs.HC2025
Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant
Joseph Matveyenko, James Liu, John David Parsons +3
Qualitative research emphasizes constructing meaning through iterative engagement with textual data. Traditionally this human-driven process requires navigating coder fatigue and i…
cs.AI2025
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
Divyansh Garg, Shaun VanWeelden, Diego Caples +15
We introduce REAL, a benchmark and framework for multi-turn agent evaluations on deterministic simulations of real-world websites. REAL comprises high-fidelity, deterministic repli…
cs.RO2024
Multi-Task Interactive Robot Fleet Learning with Visual World Models
Huihan Liu, Yu Zhang, Vaarij Betala +4
Recent advancements in large-scale multi-task robot learning offer the potential for deploying robot fleets in household and industrial settings, enabling them to perform diverse t…