6 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CL2025
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
Leitian Tao, Ilia Kulikov, Swarnadeep Saha +5
Post-training for reasoning of large language models (LLMs) increasingly relies on verifiable rewards: deterministic checkers that provide 0-1 correctness signals. While reliable,…
cs.CV2023
Activate and Reject: Towards Safe Domain Generalization under Category Shift
Chaoqi Chen, Luyao Tang, Leitian Tao +4
Albeit the notable performance on in-domain test points, it is non-trivial for deep neural networks to attain satisfactory accuracy when deploying in the open world, where novel do…
cs.LG2023★ 6 cited
Non-Parametric Outlier Synthesis
Leitian Tao, Xuefeng Du, Xiaojin Zhu +1
Out-of-distribution (OOD) detection is indispensable for safely deploying machine learning models in the wild. One of the key challenges is that models lack supervision signals fro…