activity
20172026
most citedRating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks

9 citations · 22 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

14 papers · 1 filter

cs.CL2026

Steering Instruction Hierarchies at Inference Time

Siqi Zeng, Sewoong Lee, Han Zhao +1

Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs…

cs.CL2026

ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces

Jinu Lee, Shivam Agarwal, Amruta Parulekar +3

Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the re…

cs.CL2026

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors

Yekaterina Yegorova, Argyrios Gerogiannis, Haolong Zheng +3

Speech-aware large language models often generalize poorly to out-of-domain settings. We propose SALSA (Speech-Aware LLM Adaptation via Learned Steering Activations), a lightweight…

cs.CL2026

Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures

Risham Sidhu, Julia Hockenmaier

We introduce GSU, a text-only grid dataset to evaluate the spatial reasoning capabilities of LLMs over 3 core tasks: navigation, object localization, and structure composition. By…

cs.CL20259 cited

Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks

Rajarshi Haldar, Julia Hockenmaier

As Natural Language Generation (NLG) continues to be widely adopted, properly assessing it has become quite difficult. Lately, using large language models (LLMs) for evaluating the…

cs.CL2025

ReasoningFlow: Semantic Structure of Complex Reasoning Traces

Jinu Lee, Sagnik Mukherjee, Dilek Hakkani-Tur +1

Large reasoning models (LRMs) generate complex reasoning traces with planning, reflection, verification, and backtracking. In this work, we introduce ReasoningFlow, a unified schem…