works on

From the 2 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.CL2026

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

Amruta Parulekar, Jinu Lee, Dilek Hakkani-Tür +1

Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps ar…

cs.CL2026

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah +5

The paper investigates how informing language model agents about each other's tendency to agree with users (sycophancy) can reduce error cascades in multi‑agent discussions and boo…

cs.CL2026

ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces

Jinu Lee, Shivam Agarwal, Amruta Parulekar +3

Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the re…

cs.CL2026

Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits

Soorya Ram Shimgekar, Agam Goyal, Amruta Parulekar +6

Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic…

cs.CL2025

LASER: An LLM-based ASR Scoring and Evaluation Rubric

Amruta Parulekar, Preethi Jyothi

Standard ASR evaluation metrics like Word Error Rate (WER) tend to unfairly penalize morphological and syntactic nuances that do not significantly alter sentence semantics. We intr…

cs.CL2025

AMPS: ASR with Multimodal Paraphrase Supervision

Abhishek Gupta, Amruta Parulekar, Sameep Chattopadhyay +1

Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique…