activity
20162026
most citedInteractive Model Cards: A Human-Centered Approach to Model Documentation

101 citations · 288 across the 21 of their papers we have counts for

collaborators
Showing cs.CLShow all

32 papers · 1 filter

cs.CL2026

: Benchmarking AI Agents for Long-Term Planning and Consistent Execution

Muyu He, Adit Jain, Anand Kumar +4

As LLM agents tackle increasingly complex tasks, a critical question is whether they can maintain strategic coherence over long horizons: planning under uncertainty, learning from…

cs.CL2025

The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models

Muyu He, Muhammad Ali Shafique, Anand Kumar +2

Distilling the thinking traces of a Large Language Model (LLM) with reasoning capabilities into a smaller model has been proven effective. Yet, there is a scarcity of work done on…

cs.CL2025★ 1 cited

Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models

Meghana Rajeev, Rajkumar Ramamurthy, Prapti Trivedi +5

We investigate the robustness of reasoning models trained for step-by-step problem solving by introducing query-agnostic adversarial triggers - short, irrelevant text that, when ap…

cs.CL2024

VERITAS: A Unified Approach to Reliability Evaluation

Rajkumar Ramamurthy, Meghana Arakkal Rajeev, Oliver Molenschot +2

Large language models (LLMs) often fail to synthesize information from their context to generate an accurate response. This renders them unreliable in knowledge intensive settings…

cs.CL2024

Self-rationalization improves LLM as a fine-grained judge

Prapti Trivedi, Aditya Gulati, Oliver Molenschot +7

LLM-as-a-judge models have been used for evaluating both human and AI generated content, specifically by providing scores and rationales. Rationales, in addition to increasing tran…

cs.CL2022★ 1 cited

Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations

Swarnadeep Saha, Peter Hase, Nazneen Rajani +1

Recent work on explainable NLP has shown that few-shot prompting can enable large pretrained language models (LLMs) to generate grammatical and factual natural language explanation…