3 papers
cs.LG2026
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion mode…
cs.RO2025
NeRF-Aug: Data Augmentation for Robotics with Neural Radiance Fields
Eric Zhu, Mara Levy, Matthew Gwilliam +1
Training a policy that can generalize to unknown objects is a long standing challenge within the field of robotics. The performance of a policy often drops significantly in situati…
cs.AI2025
Magentic-UI: Towards Human-in-the-loop Agentic Systems
Hussein Mozannar, Gagan Bansal, Cheng Tan +17
AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-l…