Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use
Abhijit Kumar, Zoey Wu, Mohit Suley
Humans know when to reach for help e.g. warrants a calculator while does not. Language models do not. Prompt-based approaches can instruct a model when to inv…
cs.LG2026
Execution-Grounded Credit Assignment for GRPO in Code Generation
Abhijit Kumar, Natalya Kumar, Shikhar Gupta
Critic-free reinforcement learning with verifiable rewards (RLVR) improves code generation by optimizing unit-test pass rates, but GRPO-style updates suffer from coarse credit assi…