3 papers
cs.LG2026
Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning
Yu Li, Shu Hong, Tian Lan
Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy sel…
cs.NI2026
ZODIAC: Zero-shot Offline Diffusion for Inferring Multi-xApps Conflicts in Open Radio Access Networks
Zeyu Fang, Shu Hong, Huu Trung Thieu +2
Open Radio Access Network (O-RAN) enables network control through multi-vendor xApps operating both within and across layers, subnets, and domains, whose concurrent execution can t…
cs.LG2025
Global Optimization on Graph-Structured Data via Gaussian Processes with Spectral Representations
Shu Hong, Yongsheng Mei, Mahdi Imani +1
Bayesian optimization (BO) is a powerful framework for optimizing expensive black-box objectives, yet extending it to graph-structured domains remains challenging due to the discre…