Publications (11)
State commitment learning: training language models to distinguish computation from memory
Fei Ding, Yongkang Zhang, Runhao Liu +3
Reasoning language models do not distinguish tokens used for computation from tokens that constitute persistent state: once generated, all hidden thoughts remain in context and inf…
Probabilistic Method for Optimizing Submarine Search and Rescue Strategy Under Environmental Uncertainty
Runhao Liu, Ziming Chen, Peng Zhang
When coping with the urgent challenge of locating and rescuing a deep-sea submersible in the event of communication or power failure, environmental uncertainty in the ocean can not…
Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs
Fei Ding, Yongkang Zhang, Runhao Liu +3
Post-training of large language models optimizes only parameters, while inference-time procedural scaffolds are typically designed independently of parameter training. This disconn…
LATTICE: Constraint-Directed Scheduling, Memory Planning, and Pipeline Refinement for NPUs
Runhao Liu, Minman Pei, Peng Zheng +7
General-purpose NPUs execute fine-grained command DAGs across heterogeneous compute and memory-transfer engines backed by finite, explicitly managed on-chip memories. This executio…
Exploring the Challenge and Value of Deep Learning in Automated Skin Disease Diagnosis
Runhao Liu, Ziming Chen, Guangzhen Yao +1
Skin cancer is one of the most prevalent and deadly forms of cancer worldwide, highlighting the critical importance of early detection and diagnosis in improving patient outcomes.…
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
Fei Ding, Yongkang Zhang, Runhao Liu +4
The central challenge of reinforcement learning for reasoning lies not only in the sparsity of outcome-level supervision, but more fundamentally in how to transform feedback provid…