From the 1 of 4 linked papers with an AI index.
4 papers
Stabilized Best-of- Training for Neural Combinatorial Optimization
Melveena Jolly, Midhun Xavier
Leader Reward modifies POMO training to emphasize the best trajectory produced by repeated inference. We test a narrow extension: replace its binary leader/non-leader distinction w…
Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of- Objective
Melveena Jolly, Midhun Xavier
The paper proposes an unbiased estimator for the expected maximum reward of a size‑K Plackett‑Luce draw without replacement, using rank‑conditioned Horvitz‑Thompson estimation to e…
IndustriConnect: MCP Adapters and Mock-First Evaluation for AI-Assisted Industrial Operations
Melwin Xavier, Melveena Jolly, Vaisakh M A +1
AI assistants can decompose multi-step workflows, but they do not natively speak industrial protocols such as Modbus, MQTT/Sparkplug B, or OPC UA, so this paper presents INDUSTRICO…
Agentproof: Static Verification of Agent Workflow Graphs
Melwin Xavier, Vaisakh M A, Melveena Jolly +1
Agent frameworks increasingly encode tool-using behavior as explicit workflow graphs, yet safety enforcement remains a runtime concern. These frameworks expose analyzable graph str…