most citedTransformers Can Navigate Mazes With Multi-Step Prediction

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.AI2025

OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability

Karen Ullrich, Jingtong Su, Claudia Shi +7

Reliability is key to realizing the promise of autonomous UI-Agents, multimodal agents that directly interact with apps in the same manner as humans, as users must be able to trust…

cs.LG2025

What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes

Candace Ross, Florian Bordes, Adina Williams +2

Multimodal language models possess a remarkable ability to handle an open-vocabulary's worth of objects. Yet the best models still suffer from hallucinations when reasoning about s…

cs.CL2025

LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance

Patrick Haller, Mark Ibrahim, Polina Kirichenko +2

For Large Language Models (LLMs) to be reliable, they must learn robust knowledge that can be generally applied in diverse settings -- often unlike those seen during training. Yet,…

cs.CL2025

A Single Character can Make or Break Your LLM Evals

Jingtong Su, Jianyu Zhang, Karen Ullrich +2

Common Large Language model (LLM) evaluations rely on demonstration examples to steer models' responses to the desired style. While the number of examples used has been studied and…

cs.CV2024

EvalGIM: A Library for Evaluating Generative Image Models

Melissa Hall, Oscar Mañas, Reyhane Askari-Hemmat +14

As the use of text-to-image generative models increases, so does the adoption of automatic benchmarking methods used in their evaluation. However, while metrics and datasets abound…

cs.LG20241 cited

Transformers Can Navigate Mazes With Multi-Step Prediction

Niklas Nolte, Ouail Kitouni, Adina Williams +2

Despite their remarkable success in language modeling, transformers trained to predict the next token in a sequence struggle with long-term planning. This limitation is particularl…