15 papers
Path-Space Mirror Descent for On-Policy Reinforcement Learning under the Generalized Schrödinger Bridge
Yuehu Gong, Zeyuan Wang, Yulin Chen +3
Classical on-policy algorithms such as PPO and mirror descent policy optimization provide stable proximal policy updates through tractable action likelihoods, but are typically ins…
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
Zeyuan Wang, Da Li, Yulin Chen +6
Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and have tractable entropy, but str…
The Unlearnability Phenomenon in RLVR for Language Models
Yulin Chen, He He, Chen Zhao
Reinforcement Learning with Verifiable Reward (RLVR) has proven effective in improving Large Language Model's (LLM) reasoning ability. However, the learning dynamics of RLVR remain…
Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
Jucheng Hu, Zhangquan Chen, Yulin Chen +9
Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical external signals (e.g., a cat's…
TACT: Mitigating Overthinking and Overacting in Coding Agents via Activation Steering
Yuan Sui, Yulin Chen, Yibo Li +6
When language model agents tackle complex software engineering tasks, they often degrade over long trajectories, which we define as *agent drift*. We focus on two recurring failure…
Understanding Parents' Desires in Moderating Children's Interactions with GenAI Chatbots through LLM-Generated Probes
John Driscoll, Yulin Chen, Viki Shi +3
This paper studies how parents want to moderate children's interactions with Generative AI chatbots, with the goal of informing the design of future GenAI parental control tools. W…