10 papers
HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices
Wei Liu, Siya Qi, Linhai Zhang +2
Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisation. The emergence of LLM ag…
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
Wei Liu, Siya Qi, Yali Du +1
Large language models (LLMs) make it plausible to build systems that improve through self-evolving loops, but many existing proposals are better understood as self-play and often p…
Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
Wei Liu, Peijie Yu, Michele Orini +2
The agency expected of Agentic Large Language Models goes beyond answering correctly, requiring autonomy to set goals and decide what to explore. We term this investigatory intelli…
ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour
Jack Contro, Simrat Deol, Yulan He +1
This paper introduces ChatbotManip, a novel dataset for studying manipulation in Chatbots. It contains simulated generated conversations between a chatbot and a (simulated) user, w…
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
Siya Qi, Lin Gui, Yulan He +1
The rapid advancement of Large Language Models (LLMs) has brought a pressing challenge: how to reliably assess hallucinations to guarantee model trustworthiness. Although Automatic…
Beyond Quantification: Navigating Uncertainty in Professional AI Systems
Sylvie Delacroix, Diana Robinson, Umang Bhatt +12
The growing integration of large language models across professional domains transforms how experts make critical decisions in healthcare, education, and law. While significant res…