From the 2 of 8 linked papers with an AI index.
8 papers
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
Amruta Parulekar, Jinu Lee, Dilek Hakkani-Tür +1
Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps ar…
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah +5
The paper investigates how informing language model agents about each other's tendency to agree with users (sycophancy) can reduce error cascades in multi‑agent discussions and boo…
ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces
Jinu Lee, Shivam Agarwal, Amruta Parulekar +3
Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the re…
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
Soorya Ram Shimgekar, Agam Goyal, Amruta Parulekar +6
Large language models (LLMs) are increasingly deployed in conversational settings where user tone ranges from polite to adversarial or toxic, yet less is known about whether toxic…
LASER: An LLM-based ASR Scoring and Evaluation Rubric
Amruta Parulekar, Preethi Jyothi
Standard ASR evaluation metrics like Word Error Rate (WER) tend to unfairly penalize morphological and syntactic nuances that do not significantly alter sentence semantics. We intr…
AMPS: ASR with Multimodal Paraphrase Supervision
Abhishek Gupta, Amruta Parulekar, Sameep Chattopadhyay +1
Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique…