7 papers
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11
While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas s…
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
Shomik Jain, Jack Lanchantin, Maximilian Nickel +4
Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, responses may vary in terms of pro…
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
Nikolaos Tsilivis, Eran Malach, Karen Ullrich +1
Recent advances in reasoning domains with neural networks have primarily been enabled by a training recipe that optimizes Large Language Models, previously trained to predict the n…
OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
Karen Ullrich, Jingtong Su, Claudia Shi +7
Reliability is key to realizing the promise of autonomous UI-Agents, multimodal agents that directly interact with apps in the same manner as humans, as users must be able to trust…
A Single Character can Make or Break Your LLM Evals
Jingtong Su, Jianyu Zhang, Karen Ullrich +2
Common Large Language model (LLM) evaluations rely on demonstration examples to steer models' responses to the desired style. While the number of examples used has been studied and…
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
Jingtong Su, Julia Kempe, Karen Ullrich
Transformers have achieved state-of-the-art performance across language and vision tasks. This success drives the imperative to interpret their internal mechanisms with the dual go…