2 citations · 3 across the 8 of their papers we have counts for
6 papers · 1 filter
LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
Ming Yin, Dinghan Shen, Silei Xu +11
Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, provider-specific tool definitions, the Mo…
LogicIF: Towards Complex Logic Instruction Following
Mian Zhang, Shujian Liu, Sixun Dong +9
Instruction following has catalyzed the recent era of Large Language Models (LLMs) and is the foundational skill underpinning more advanced capabilities such as reasoning and agent…
Developing a Reliable, Fast, General-Purpose Hallucination Detection and Mitigation Service
Song Wang, Xun Wang, Jie Mei +6
Hallucination, a phenomenon where large language models (LLMs) produce output that is factually incorrect or unrelated to the input, is a major challenge for LLM applications that…
LMGQS: A Large-scale Dataset for Query-focused Summarization
Ruochen Xu, Song Wang, Yang Liu +5
Query-focused summarization (QFS) aims to extract or generate a summary of an input document that directly answers or is relevant to a given query. The lack of large-scale datasets…
Summarization with Precise Length Control
Lesly Miculicich, Yujia Xie, Song Wang +1
Many applications of text generation such as summarization benefit from accurately controlling the text length. Existing approaches on length-controlled summarization either result…
An End-to-End Dialogue Summarization System for Sales Calls
Abedelkadir Asi, Song Wang, Roy Eisenstadt +4
Summarizing sales calls is a routine task performed manually by salespeople. We present a production system which combines generative models fine-tuned for customer-agent setting,…