6 papers
Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning
Benjamin Poole, Minwoo Lee
Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feed…
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
Benjamin Poole, Andrew Quinn, Li Yang +1
Data rehearsal has emerged as a leading approach for mitigating catastrophic forgetting in Continual Reinforcement Learning (CRL). However, existing work remains confined to policy…
From Static Risk to Dynamic Trajectories: Toward World-Model-Inspired Clinical Prediction
Pujun Feng, Xiaoyu Guo, Seyed Ehsan Saffari +10
Clinical decision-making is a feedback system where risk estimates influence treatment, which in turn changes disease trajectories, and both shape clinicians' measurement practices…
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
Ziyan Wang, Enmao Diao, Qi Le +6
Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectures. However, the dominant local…
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
Ziyan Wang, Enmao Diao, Qi Le +5
Reasoning LLMs (RLMs) such as OpenAI o1, DeepSeek-R1, and Qwen3 deliver strong multi-step reasoning through chain-of-thought generation, but their large model sizes and lengthy dec…
Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques
Raju Challagundla, Mohsen Dorodchi, Pu Wang +1
As privacy regulations become more stringent and access to real-world data becomes increasingly constrained, synthetic data generation has emerged as a vital solution, especially f…