Showing 2026Show all
2 papers · 1 filter
cs.AI2026
Interaction Locality in Hierarchical Recursive Reasoning
Yosuke Miyanishi, Tetsuro Morimura
Spatial reasoning requires both location-bound computation and location-invariant structure: agents must make local moves while preserving route, object, or constraint-level plans.…
cs.LG2026
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura +3
Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available. However, such a highly reliable reward model is no…