7 papers
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
Rounak Saha, Gurusha Juneja, Dayita Chaudhuri +3
A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction…
Learning the ARTS of Search for Automated Discovery
Gurusha Juneja, Arnav Kumar Jain, Deepak Nathani +2
Scientific discovery can be formulated as an iterative search process over the space of hypotheses and experiments. Contemporary methods navigate this space using heuristics such a…
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents
Gurusha Juneja, Dylan Lu, Saaket Agashe +7
Theory of Mind (ToM), the ability to track others epistemic state, makes humans efficient collaborators. AI agents need the same capacity in multi agent settings, yet existing benc…
Adversarial Training for Process Reward Models
Gurusha Juneja, Deepak Nathani, William Yang Wang
Process Reward Models (PRMs) enhance reasoning ability of LLMs by providing step-level supervision. However, their widespread adoption is limited due to expensive manual step-level…
MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
Gurusha Juneja, Jayanth Naga Sai Pasupulati, Alon Albalak +2
A core challenge for autonomous LLM agents in collaborative settings is balancing robust privacy understanding and preservation alongside task efficacy. Existing privacy benchmarks…
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
Gurusha Juneja, Alon Albalak, Wenyue Hua +1
The proliferation of LLM-based agents has led to increasing deployment of inter-agent collaboration for tasks like scheduling, negotiation, resource allocation etc. In such systems…