3 papers
cs.LG2026
Q-Steer: Action-Value Guidance for Molecular Policy Optimization
Xinyu Wang, Jinbo Bi, Minghu Song
Oracle-limited molecular optimization gives reward only after a complete molecule is generated, while each rollout requires many local next-token decisions. This delayed-feedback i…
cs.CV2026
LARV: Data-Free Layer-wise Adaptive Rescaling Veneer for Model Merging
Xinyu Wang, Ke Deng, Fei Dou +2
Model merging aims to combine multiple fine-tuned models into a single multi-task model without access to training data. Existing task-vector merging methods such as TIES, TSV-M, a…
cs.LG2025
Leveraging Partial SMILES Validation Scheme for Enhanced Drug Design in Reinforcement Learning Frameworks
Xinyu Wang, Jinbo Bi, Minghu Song
SMILES-based molecule generation has emerged as a powerful approach in drug discovery. Deep reinforcement learning (RL) using large language model (LLM) has been incorporated into…