5 papers
Explainable AI in Big Data Fraud Detection
Ayush Jain, Rahul Kulkarni, Siyi Lin
Big Data has become central to modern applications in finance, insurance, and cybersecurity, enabling machine learning systems to perform large-scale risk assessments and fraud det…
Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors
Henger Li, Shuangjie You, Flavio Di Palo +2
Tool calling enables large language models (LLMs) to interact with external environments through tool invocation, providing a practical way to overcome the limitations of pretraini…
LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature
Magdalena Lederbauer, Siddharth Betala, Xiyao Li +16
The development of synthesis procedures remains a fundamental challenge in materials discovery, with procedural knowledge scattered across decades of scientific literature in unstr…
Train on Validation (ToV): Fast data selection with applications to fine-tuning
Ayush Jain, Andrea Montanari, Eren Sasoglu
State-of-the-art machine learning often follows a two-stage process: ~pre-training on large, general-purpose datasets; ~fine-tuning on task-specific data. In fine-tuning…
Scaling laws for learning with real and surrogate data
Ayush Jain, Andrea Montanari, Eren Sasoglu
Collecting large quantities of high-quality data can be prohibitively expensive or impractical, and a bottleneck in machine learning. One may instead augment a small set of dat…