3 papers
cs.AI2026
Evidence-State Rewards for Long-Context Reasoning
Ya Gao, Pekka Marttinen
Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods usually reward final answers o…
cs.AI2026
Edit Knowledge, Not Just Facts via Multi-Step Reasoning over Background Stories
Ya Gao, Kalle Kujanpää, Pekka Marttinen +2
Enabling artificial intelligence systems, particularly large language models, to update knowledge and flexibly apply it during reasoning remains a central challenge. Existing knowl…
cs.LG2025
Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization
Minttu Alakuijala, Ya Gao, Georgy Ananov +4
As the general capabilities of artificial intelligence (AI) agents continue to evolve, their ability to learn to master multiple complex tasks through experience remains a key chal…