3 papers
cs.CL2026
Efficient LLM-based Advertising via Model Compression and Parallel Verification
Wenxin Dong, Chang Gao, Guanghui Yu +9
Large language models (LLMs) have shown remarkable potential in advertising scenarios such as ad creative generation and targeted advertising. However, deploying LLMs in real-time…
cs.CL2026
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
Wenxin Dong, Mingqing Hu, Guanghui Yu +7
When large language models (LLMs) serve real-time inference in commercial online advertising systems, end-to-end latency must be strictly bounded to the millisecond range. Yet ever…
cs.LG2026
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
Huimin Xu, Shuai Zhao, Xiaobao Wu +1
Reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning ability of large language models. However, widely used RLVR algor…