From the 2 of 37 linked papers with an AI index.
2 citations · 2 across the 19 of their papers we have counts for
9 papers · 1 filter
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework…
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation
Yan Wang, Yi Han, Lingfei Qian +11
Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatil…
Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
Yi Han, Yan Wang, Lingfei Qian +12
Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remains unclear. Unlike reactive tasks…
GAM: Hierarchical Graph-based Agentic Memory for LLM Agents
Zhaofen Wu, Hanrong Zhang, Fulin Lin +9
To sustain coherent long-term interactions, Large Language Model (LLM) agents must navigate the tension between acquiring new information and retaining prior knowledge. Current uni…
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
Hanrong Zhang, Shicheng Fan, Henry Peng Zou +12
Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool is a single, self-contained func…
Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation
Bowei He, Yankai Chen, Xiaokun Zhang +4
Knowledge distillation from Large Language Models (LLMs) to smaller models has emerged as a critical technique for deploying efficient AI systems. However, current methods for dist…