collaborators

6 papers

cs.AI2026

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

Benjamin Poole, Minwoo Lee

Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feed…

cs.LG2026

Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning

Benjamin Poole, Andrew Quinn, Li Yang +1

Data rehearsal has emerged as a leading approach for mitigating catastrophic forgetting in Continual Reinforcement Learning (CRL). However, existing work remains confined to policy…

cs.AI2026

From Static Risk to Dynamic Trajectories: Toward World-Model-Inspired Clinical Prediction

Pujun Feng, Xiaoyu Guo, Seyed Ehsan Saffari +10

Clinical decision-making is a feedback system where risk estimates influence treatment, which in turn changes disease trajectories, and both shape clinicians' measurement practices…

cs.CL2026

From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models

Ziyan Wang, Enmao Diao, Qi Le +6

Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectures. However, the dominant local…

cs.CL2025

Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models

Ziyan Wang, Enmao Diao, Qi Le +5

Reasoning LLMs (RLMs) such as OpenAI o1, DeepSeek-R1, and Qwen3 deliver strong multi-step reasoning through chain-of-thought generation, but their large model sizes and lengthy dec…

cs.LG2025

Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques

Raju Challagundla, Mohsen Dorodchi, Pu Wang +1

As privacy regulations become more stringent and access to real-world data becomes increasingly constrained, synthetic data generation has emerged as a vital solution, especially f…