From the 1 of 22 linked papers with an AI index.
9 papers · 1 filter
TRACER: Early Failure Detection for Task-Oriented Dialogue
Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1
Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present…
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
Ke Yang, Zixi Chen, Xuan He +6
Long-term memory is essential for large language model (LLM) agents operating in complex environments, yet existing memory designs are either task-specific and non-transferable, or…
A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems
ÄorÄe Klisura, Astrid R Bernaga Torres, Anna Karen Gárate-Escamilla +4
Privacy policies inform users about data collection and usage, yet their complexity limits accessibility for diverse populations. Existing Privacy Policy Question Answering (QA) sy…
CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories
Yilong Lai, Yipin Yang, Jialong Wu +5
Recent years have witnessed the rapid development of LLM-based agents, which shed light on using language agents to solve complex real-world problems. A prominent application lies…
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
Yiran Liu, Ke Yang, Zehan Qi +3
We present a novel statistical framework for analyzing stereotypes in large language models (LLMs) by systematically estimating the bias and variation in their generation. Current…
ADEPT: A DEbiasing PrompT Framework
Ke Yang, Charles Yu, Yi Fung +2
Several works have proven that finetuning is an applicable approach for debiasing contextualized word embeddings. Similarly, discrete prompts with semantic meanings have shown to b…