activity
20242026
collaborators

7 papers

cs.AI2026

SCRuB: Social Concept Reasoning under Rubric-Based Evaluation

Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11

While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas s…

cs.CL2026

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework

Shomik Jain, Jack Lanchantin, Maximilian Nickel +4

Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, responses may vary in terms of pro…

cs.LG2025

How Reinforcement Learning After Next-Token Prediction Facilitates Learning

Nikolaos Tsilivis, Eran Malach, Karen Ullrich +1

Recent advances in reasoning domains with neural networks have primarily been enabled by a training recipe that optimizes Large Language Models, previously trained to predict the n…

cs.AI2025

OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability

Karen Ullrich, Jingtong Su, Claudia Shi +7

Reliability is key to realizing the promise of autonomous UI-Agents, multimodal agents that directly interact with apps in the same manner as humans, as users must be able to trust…

cs.CL2025

A Single Character can Make or Break Your LLM Evals

Jingtong Su, Jianyu Zhang, Karen Ullrich +2

Common Large Language model (LLM) evaluations rely on demonstration examples to steer models' responses to the desired style. While the number of examples used has been studied and…

cs.LG2025

From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers

Jingtong Su, Julia Kempe, Karen Ullrich

Transformers have achieved state-of-the-art performance across language and vision tasks. This success drives the imperative to interpret their internal mechanisms with the dual go…