activity
20232026
most citedReasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions

2 citations · 5 across the 27 of their papers we have counts for

collaborators
Showing cs.CLShow all

23 papers · 1 filter

cs.CL2026

ABSOL: Aggregated Bayesian Subsampling Orchestrated with LLMs

Jackson Hassell, Chen Shen, Estevam Hruschka

Large language models are increasingly used as natural-language interfaces to structured data, yet they remain unreliable when answers require consistent evidence conditioning, dep…

cs.CL2026

Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance

Chen Shen, Estevam Hruschka

Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files su…

cs.CL2026

Reflective Prompt Tuning through Language Model Function-Calling

Farima Fatahi Bayat, Moin Aminnaseri, Pouya Pezeshkpour +1

Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for adapting models without par…

cs.CL2026

Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

Naoki Otani, Nikita Bhutani, Hannah Kim +2

Explicit planning is a critical capability for LLM-based agents solving complex data-centric tasks, which require precise tool calling over external data sources. Existing strategi…

cs.CL2026

AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs

Pouya Pezeshkpour, Estevam Hruschka

Verification is becoming central to both reinforcement-learning-based training and inference-time control of large language models (LLMs). Yet current verifiers face a fundamental…

cs.CL2026

A Dynamic Self-Evolving Extraction System

Moin Amin-Naseri, Hannah Kim, Estevam Hruschka

The extraction of structured information from raw text is a fundamental component of many NLP applications, including document retrieval, ranking, and relevance estimation. High-qu…