4 papers
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
Siyao Song, Cong Ma, Zhihao Cheng +5
Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcom…
Trans-Glasso: A Transfer Learning Approach to Precision Matrix Estimation
Boxin Zhao, Cong Ma, Mladen Kolar
Precision matrix estimation is essential in various fields; yet it is challenging when samples for the target study are limited. Transfer learning can enhance estimation accuracy b…
LISM: Long-range Integrative State space Models via Input-Latent State Interactions
Cong Ma, Kayvan Najarian, Hovhannes Baghdasaryan +1
State space models (SSMs) are an emerging paradigm that achieves linear-time scaling, however, they intrinsically suffer from "curse of memory". The memory of SSMs, including Mamba…
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Minghao Li, Ying Zeng, Zhihao Cheng +2
The advent of Deep Research agents has substantially reduced the time required for conducting extensive research tasks. However, these tasks inherently demand rigorous standards of…