2 papers
cs.CV2026
Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models
Michael Jerge, Joseph Pelczar, Justin Downes
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning ability of vision-language models (VLMs), and diversifying the rollouts within each optimization group…
cs.LG2026
NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning
Michael Jerge, David Evans
This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal representations. Unlike fine-tuning me…