7 papers
CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models
Yash Shah, Abhijit Chakraborty, Vivek Gupta
Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming…
Synapse: Federated Tool Routing via Typed Compendium Artifacts
Abhijit Chakraborty, Yash Shah, Vivek Gupta
The unit of collaboration in federated learning determines what guarantees are even expressible. Flat units like weights, prompts, raw examples, carry no type signature on which pr…
SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code
Natalia Tarasova, Enrique Balp-Straffon, Aleksei Iancheruk +10
Building infrastructure-as-code (IaC) in cloud computing is a critical task, underpinning the reliability, scalability, and security of modern software systems. Despite the remarka…
GamED.AI: A Hierarchical Multi-Agent Framework for Automated Educational Game Generation
Shiven Agarwal, Yash Shah, Ashish Raj Shekhar +2
We introduce GamEDAI, a hierarchical multi-agent framework that transforms instructor-provided questions into fully playable, pedagogically grounded educational games validated thr…
OSCAR: Orchestrated Self-verification and Cross-path Refinement
Yash Shah, Abhijit Chakraborty, Naresh Kumar Devulapally +2
Diffusion language models (DLMs) expose their denoising trajectories, offering a natural handle for inference-time control; accordingly, an ideal hallucination mitigation framework…
DoPE: Decoy Oriented Perturbation Encapsulation Human-Readable, AI-Hostile Documents for Academic Integrity
Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi +3
Multimodal Large Language Models (MLLMs) can directly consume exam documents, threatening conventional assessments and academic integrity. We present DoPE (Decoy-Oriented Perturbat…