7 citations · 7 across the 17 of their papers we have counts for
4 papers · 1 filter
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
Xilun Chen, Zhaleh Feizollahi, Ross Goodwin +5
Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric requires expressing the structure o…
Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
Siddhant Arora, Haidar Khan, Kai Sun +14
End-to-end speech-in speech-out dialogue systems are emerging as a powerful alternative to traditional ASR-LLM-TTS pipelines, generating more natural, expressive responses with sig…
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
Shervin Ghasemlou, Ashish Katiyar, Aparajita Saraf +5
In this paper, we investigate the problem of "generation supervision" in large language models, and present a novel bicameral architecture to separate supervision signals from thei…
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
Zekun Li, Zhiyu Zoey Chen, Mike Ross +7
Large language models (LLMs) are increasingly prevalent in conversational systems due to their advanced understanding and generative capabilities in general contexts. However, thei…