2 papers
cs.LG2026
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Xin-Qiang Cai, Wei Wang, Feng Liu +3
Reinforcement Learning with Verifiable Rewards (RLVR) replaces costly human labeling with automated verifiers. To reduce verifier hacking, many RLVR systems binarize rewards to $\{…
cs.CV2024
Mind the Gap Between Prototypes and Images in Cross-domain Finetuning
Hongduan Tian, Feng Liu, Zhanke Zhou +3
In cross-domain few-shot classification (CFC), recent works mainly focus on adapting a simple transformation head on top of a frozen pre-trained backbone with few labeled data to p…