13 papers
Continuous-Time Probabilistic Correctors for Uncertainty-Aware Physics-Based Spacecraft Trajectory Forecasting
Muhammad Bilal Shahid, Zhanhong Jiang, Soumik Sarkar +1
Long-horizon spacecraft trajectory forecasting suffers from error accumulation due to the absence of corrective observations in the forecast regime, making reliable uncertainty est…
Lighting-aware Unified Model for Instance Segmentation
Qisai Liu, Alloy Das, Zhanhong Jiang +4
Foundation models like the Segment Anything Model (SAM) demonstrate impressive zero-shot generalization but frequently degrade under diverse real-world illumination, particularly f…
Distributed Direct Preference Optimization
Zhanhong Jiang
Preference-based reinforcement learning (RL) is a key paradigm for aligning policies with human judgments, yet its theoretical behavior in distributed settings where preference dat…
TabQL: In-Context Q-Learning with Tabular Foundation Models
Qisai Liu, Zhanhong Jiang, Timilehin Ayanlade +4
We propose Tabular Q-Learning (TabQL), a reinforcement learning framework that replaces the conventional parametric Q-network in Deep Q-Learning (DQN) with a tabular foundation mod…
COOPO: Cyclic Offline-Online Policy Optimization Algorithm
Qisai Liu, Zhanhong Jiang, Joshua Russell Waite +3
Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL demands prohibitive environment in…
ADKO: Agentic Decentralized Knowledge Optimization
Lucas Nerone Rillo, Zhanhong Jiang, Nastaran Saadati +4
We present Agentic Decentralized Knowledge Optimization (ADKO), a framework for collaborative black-box optimization across autonomous agents that achieves sample efficiency, priva…