7 papers
FedPref: Federated Preference Learning for Structured Radiology Report Extraction
Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong +1
Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema. Learning this extraction requires labe…
Why Do We Suffer for Fun? Ordeal Pleasure in Souls-like Games
Flint Xiaofeng Fan
Souls-like games exemplify how digital play can produce radical forms of pleasure through sustained challenge: players voluntarily invest tens or hundreds of hours in experiences d…
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
Ofer Meshi, Krisztian Balog, Sally Goldman +5
The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but…
Information Fidelity in Tool-Using LLM Agents: A Martingale Analysis of the Model Context Protocol
Flint Xiaofeng Fan, Cheston Tan, Roger Wattenhofer +1
As AI agents powered by large language models (LLMs) increasingly use external tools for high-stakes decisions, a critical reliability question arises: how do errors propagate acro…
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
Zhenglin Wan, Xingrui Yu, David Mark Bossens +5
Imitation learning (IL) has shown promise in various applications (e.g. robot locomotion) but is often limited to learning a single expert policy, constraining behavior diversity a…
FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong +2
In the era of increasing privacy concerns and demand for personalized experiences, traditional Reinforcement Learning with Human Feedback (RLHF) frameworks face significant challen…