5 papers
Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
Shicheng Fan, Mingdai Yang, Duohao Wang +9
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in…
Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents
Ziyi Wang, Yuxuan Lu, Yimeng Zhang +12
Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized settings with general, fixed, and…
Unifying Inference-Time Planning Language Generation
Prabhu Prakash Kagitha, Bo Sun, Ishan Desai +5
A line of work in planning uses LLM not to generate a plan, but to generate a formal representation in some planning language, which can be input into a symbolic solver to determin…
LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World
Sina J. Semnani, Pingyue Zhang, Wanyue Zhai +6
This paper presents LEMONADE, a large-scale conflict event dataset comprising 39,786 events across 20 languages and 171 countries, with extensive coverage of region-specific entiti…
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
Yuji Zhang, Sha Li, Cheng Qian +8
Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which err…