11 papers
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Lingao Xiao, Yalun Dai, Yangyu Huang +17
Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multi…
Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies
Zhuoran Li, Hai Zhong, Xun Wang +3
Online Multi-Agent Reinforcement Learning (MARL) is a prominent framework for efficient agent coordination. Crucially, enhancing policy expressiveness is pivotal for achieving supe…
Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling
Zhuoran Li, Ruishuo Chen, Hai Zhong +1
Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live streaming, and data center managemen…
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
Ruishuo Chen, Yu Chen, Zhuoran Li +1
Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without…
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
Ruishuo Chen, Xun Wang, Rui Hu +2
Generative Flow Networks (GFlowNets) excel at sampling diverse, high-reward objects. In many practical applications where active reward queries are infeasible, these models must be…
Offline-to-Online Multi-Agent Reinforcement Learning with Offline Value Function Memory and Sequential Exploration
Hai Zhong, Xun Wang, Zhuoran Li +1
Offline-to-Online Reinforcement Learning has emerged as a powerful paradigm, leveraging offline data for initialization and online fine-tuning to enhance both sample efficiency and…