2 papers
cs.LG2026
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
Tianle Zhong, Neiwen Ling, Yifan Pi +5
Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation…
cs.RO2024
TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications
Neiwen Ling, Guojun Chen, Lin Zhong
Large Language Models (LLMs) such as GPT-4 and Llama3 can already comprehend complex commands and process diverse tasks. This advancement facilitates their application in controlli…