15 papers
Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
Chih-Hsuan, Yang, Tanwi Mallick +5
Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principled ways to connect system-level evaluation…
When More Cores Hurts: The Vector Database Scaling Paradox in HPC
Seth Ockerman, Song Young Oh, Amal Gueroudji +12
Vector databases have been designed and optimized for cloud environments; however, emerging scientific AI workloads (e.g., molecular search, meteorological trajectory detection, an…
No Test Cases, No Problem: Distillation-Driven Code Generation for Scientific Workflows
Siddeshwar Raghavan, Tanwi Mallick
Existing multi-agent Large Language Model (LLM) frameworks for code generation typically use execution feedback and improve iteratively using Input/Output (I/O) test cases. However…
Toward Reliable, Safe, and Secure LLMs for Scientific Applications
Saket Sanjeev Chaturvedi, Joshua Bergerson, Tanwi Mallick
As large language models (LLMs) evolve into autonomous "AI scientists," they promise transformative advances but introduce novel vulnerabilities, from potential "biosafety risks" t…
Statistical Early Stopping for Reasoning Models
Yangxinyu Xie, Tao Wang, Soham Mallick +6
While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given…
SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation
Homaira Huda Shomee, Rochana Chaturvedi, Yangxinyu Xie +1
Large language models (LLMs) are increasingly used to support question answering and decision-making in high-stakes, domain-specific settings such as natural hazard response and in…