Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation
Yunan Lu, Ryan Shea, Yusen Zhang +1
Evaluation remains a critical bottleneck for interactive agent development. Existing evaluation methods often rely on static benchmarks, which fail to capture the dynamic, multi-st…
cs.CL2025
SAGE: A Top-Down Bottom-Up Knowledge-Grounded User Simulator for Multi-turn AGent Evaluation
Ryan Shea, Yunan Lu, Liang Qiu +1
Evaluating multi-turn interactive agents is challenging due to the need for human assessment. Evaluation with simulated users has been introduced as an alternative, however existin…
cs.CL2024
LocalRQA: From Generating Data to Locally Training, Testing, and Deploying Retrieval-Augmented QA Systems
Xiao Yu, Yunan Lu, Zhou Yu
Retrieval-augmented question-answering systems combine retrieval techniques with large language models to provide answers that are more accurate and informative. Many existing tool…