3 papers
cs.AI2025
An Approach to Grounding AI Model Evaluations in Human-derived Criteria
Sasha Mitts
In the rapidly evolving field of artificial intelligence (AI), traditional benchmarks can fall short in attempting to capture the nuanced capabilities of AI models. We focus on the…
cs.CV2025
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
Aaron Foss, Chloe Evans, Sasha Mitts +3
We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world…
cs.CL2024
Towards Full Delegation: Designing Ideal Agentic Behaviors for Travel Planning
Song Jiang, Da JU, Andrew Cohen +5
How are LLM-based agents used in the future? While many of the existing work on agents has focused on improving the performance of a specific family of objective and challenging ta…