3 papers
cs.AI2025
Automated Reward Design for Gran Turismo
Michel Ma, Takuma Seno, Kaushik Subramanian +3
When designing reinforcement learning (RL) agents, a designer communicates the desired agent behavior through the definition of reward functions - numerical feedback given to the a…
math.AC2025
A new proof of non-Cohen-Macaulayness of Bertin's example
Takuma Seno
Bertin's example is famous as the first known Noetherian UFD that is not Cohen-Macaulay. In the example, she employed a ring of invariants and proved that the ring is not Cohen-Mac…
cs.LG2025
Hyperspherical Normalization for Scalable Deep Reinforcement Learning
Hojoon Lee, Youngdo Lee, Takuma Seno +3
Scaling up the model size and computation has brought consistent performance improvements in supervised learning. However, this lesson often fails to apply to reinforcement learnin…