activity
20232026
most citedLanguage Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

16 citations · 28 across the 33 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents

Amir Saeidi, Zehua Zhang, Rishitosh Singh +6

Large language model (LLM) agents are increasingly deployed in long-horizon, interactive, and stateful environments. In these settings, a single wrong action, such as refunding the…

cs.CL2026

DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings

Wenya Xie, Shengming Zhou, Zelin Li +9

LLM agents increasingly act as personal assistants that must remember a user's profile over months: who they are (attributes), what they routinely do (habits), and what they prefer…

cs.CL2026

FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments

Amir Saeidi, Venkatesh Mishra, Souradeep Mukhopadhyay +4

Large Language Models are being increasingly deployed as the decision-making core of autonomous agents capable of effecting change in external environments. Yet, in conversational…

cs.CL2025

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on -bench

Venkatesh Mishra, Amir Saeidi, Satyam Raj +5

Recent advances in reasoning and planning capabilities of large language models (LLMs) have enabled their potential as autonomous agents capable of tool use in dynamic environments…

cs.CL2025

Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm

Baixiang Huang, Zhen Tan, Haoran Wang +6

Agents based on Large Language Models (LLMs) have demonstrated strong capabilities across a wide range of tasks. However, deploying LLM-based agents in high-stakes domains comes wi…

cs.CL2025

AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control

Ruosen Li, Ziming Luo, Quan Zhang +4

Large reasoning models (LRMs) achieve impressive reasoning capabilities by generating lengthy chain-of-thoughts, but this "overthinking" incurs high latency and cost without commen…