1 citations · 2 across the 11 of their papers we have counts for
6 papers · 1 filter
One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
Rui Ming, Haoyuan Wu, Shoubo Hu +2
Supervised fine-tuning (SFT) is the predominant method for adapting large language models (LLMs), yet it often struggles with generalization compared to reinforcement learning (RL)…
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
Ziru Liu, Cheng Gong, Xinyu Fu +7
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a powerful paradigm for facilitating the self-improvement of large language models (LLMs), particularl…
A Systematic Evaluation of On-Device LLMs: Quantization, Performance, and Resources
Qingyu Song, Rui Liu, Wei Lin +11
Deploying Large Language Models (LLMs) on edge devices enhances privacy but faces performance hurdles due to limited resources. We introduce a systematic methodology to evaluate on…
ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving
Haoyuan Wu, Xueyi Chen, Rui Ming +4
Large language models (LLMs) demonstrate significant reasoning capabilities, particularly through long chain-of-thought (CoT) processes, which can be elicited by reinforcement lear…
TorchResist: Open-Source Differentiable Resist Simulator
Zixiao Wang, Jieya Zhou, Su Zheng +5
Recent decades have witnessed remarkable advancements in artificial intelligence (AI), including large language models (LLMs), image and video generative models, and embodied AI sy…
Architect of the Bits World: Masked Autoregressive Modeling for Circuit Generation Guided by Truth Table
Haoyuan Wu, Haisheng Zheng, Shoubo Hu +2
Logic synthesis, a critical stage in electronic design automation (EDA), optimizes gate-level circuits to minimize power consumption and area occupancy in integrated circuits (ICs)…