From the 1 of 3 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
VAmoS Bench: Voice Agent Simulation Bench
Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5
The paper presents VAmoS Bench, a simulation-based benchmark that evaluates complete voice‑agent systems on end‑to‑end customer‑support tasks, checking both conversational behavior…
cs.AI2026
POaaS: Minimal-Edit Prompt Optimization as a Service to Lift Accuracy and Cut Hallucinations on On-Device sLLMs
Jungwoo Shim, Dae Won Kim, Sun Wook Kim +4
Small language models (sLLMs) are increasingly deployed on-device, where imperfect user prompts--typos, unclear intent, or missing context--can trigger factual errors and hallucina…