4 papers
An Approach to Grounding AI Model Evaluations in Human-derived Criteria
Sasha Mitts
In the rapidly evolving field of artificial intelligence (AI), traditional benchmarks can fall short in attempting to capture the nuanced capabilities of AI models. We focus on the…
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
Aaron Foss, Chloe Evans, Sasha Mitts +3
We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world…
Towards Full Delegation: Designing Ideal Agentic Behaviors for Travel Planning
Song Jiang, Da JU, Andrew Cohen +5
How are LLM-based agents used in the future? While many of the existing work on agents has focused on improving the performance of a specific family of objective and challenging ta…
To the Globe (TTG): Towards Language-Driven Guaranteed Travel Planning
Da JU, Song Jiang, Andrew Cohen +8
Travel planning is a challenging and time-consuming task that aims to find an itinerary which satisfies multiple, interdependent constraints regarding flights, accommodations, attr…