6 papers
1000 Rallies: An Event-Camera Dataset and Real-Time Learned Ball-State Estimation for Robotic Table Tennis
Raphaela Kreiser, Asude Aydin, Yin Bi +3
Robotic table tennis has emerged as a compelling benchmark for real-time robotic perception due to its fast ball dynamics and stringent timing requirements. Accurate, high-frequenc…
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
Claudio Fanconi, Nicolás Astorga, Mihaela van der Schaar
Teaching large language models (LLMs) to reason during post-training typically relies on reinforcement learning with explicit outcome- or process-based reward functions. However, i…
Tiny Autoregressive Recursive Models
Paulius Rauba, Claudio Fanconi, Mihaela van der Schaar
Tiny Recursive Models (TRMs) have recently demonstrated remarkable performance on ARC-AGI, showing that very small models can compete against large foundation models through a two-…
Cascaded Language Models for Cost-effective Human-AI Decision-Making
Claudio Fanconi, Mihaela van der Schaar
A challenge in human-AI decision-making is to balance three factors: the correctness of predictions, the cost of knowledge and reasoning complexity, and the confidence about whethe…
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
Katarzyna Kobalczyk, Claudio Fanconi, Hao Sun +1
As large language models (LLMs) become increasingly embedded in everyday applications, ensuring their alignment with the diverse preferences of individual users has become a critic…
Discovering Preference Optimization Algorithms with and for Large Language Models
Chris Lu, Samuel Holt, Claudio Fanconi +4
Offline preference optimization is a key method for enhancing and controlling the quality of Large Language Model (LLM) outputs. Typically, preference optimization is approached as…