2 papers
cs.LG2026
Quantifying Empirical Compute-Supervision Tradeoffs in RLVR
Ryo Mitsuhashi, Patrick Chen, Isabelle Tseng +2
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are rarely perfect. Recent theore…
cs.CV2026
Guess the Unified Model: How Much Can We Recover from Generated Images?
Jasin Cekinmez, Ryo Mitsuhashi, Addison J. Wu +1
With unified model-generated images now widespread online, attributing their model of origin offers a path toward transparency and deeper insight into the characteristic behaviors…