3 papers
cs.LG2025
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
Pingzhi Li, Morris Yu-Chao Huang, Zhen Tan +6
Knowledge Distillation (KD) accelerates training of large language models (LLMs) but poses intellectual property protection and LLM diversity risks. Existing KD detection methods b…
cs.LG2025
Can GRPO Help LLMs Transcend Their Pretraining Origin?
Kangqi Ni, Zhen Tan, Zijie Liu +2
Reinforcement Learning with Verifiable Rewards (RLVR), primarily driven by the Group Relative Policy Optimization (GRPO) algorithm, is a leading approach for enhancing the reasonin…
cs.AI2025
Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
Tianyu Hu, Zhen Tan, Song Wang +2
With advancements in reasoning capabilities, Large Language Models (LLMs) are increasingly employed for automated judgment tasks. While LLMs-as-Judges offer promise in automating e…