3 papers
cs.AI2026
IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
Varun Gumma, Navonil Majumder, Soumitra Sinhahajari +1
Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they gen…
cs.DL2026
On the Limits of LLM-as-Judge for Scientific Novelty Assessment
Soumitra Sinhahajari, Navonil Majumder, Soujanya Poria
LLMs are increasingly used to generate and judge scientific ideas. This makes novelty evaluation a central problem. Full idea evaluation is difficult because it often requires judg…
cs.LG2026
Policy Gradient Methods for Non-Markovian Reinforcement Learning
Avik Kar, Siddharth Chandak, Rahul Singh +4
We study policy gradient methods for reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards depend on the entire interaction history. To…