28 citations · 67 across the 9 of their papers we have counts for
7 papers · 1 filter
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Jiarui Lu, Thomas Holleis, Yizhe Zhang +9
Recent large language models (LLMs) advancements sparked a growing research interest in tool assisted LLMs solving real-world challenges, which calls for comprehensive evaluation o…
Improving Factual Consistency of Abstractive Summarization via Question Answering
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu +7
A commonly observed problem with the state-of-the art abstractive summarization models is that the generated summaries can be factually inconsistent with the input documents. The f…
Towards Clinical Encounter Summarization: Learning to Compose Discharge Summaries from Prior Notes
Han-Chin Shing, Chaitanya Shivade, Nima Pourdamghani +4
The records of a clinical encounter can be extensive and complex, thus placing a premium on tools that can extract and summarize relevant information. This paper introduces the tas…
Entity-level Factual Consistency of Abstractive Text Summarization
Feng Nan, Ramesh Nallapati, Zhiguo Wang +5
A key challenge for abstractive summarization is ensuring factual consistency of the generated summary with respect to the original document. For example, state-of-the-art models t…
Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction
Yifan Gao, Henghui Zhu, Patrick Ng +7
In open-domain question answering, questions are highly likely to be ambiguous because users may not know the scope of relevant topics when formulating them. Therefore, a system ne…
End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering Systems
Siamak Shakeri, Cicero Nogueira dos Santos, Henry Zhu +5
We propose an end-to-end approach for synthetic QA data generation. Our model comprises a single transformer-based encoder-decoder network that is trained end-to-end to generate bo…