5 papers
FLARE: Few-shot Learning-based Adaptive Reflective Engine
Dhanasekar Sundararaman, Bharat Gandhi, Aashna Garg +1
Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-of-the-art optimizers like G…
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
Aashna Garg, Siddharth Singha Roy, Jinu Jang +2
Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences. Existing routers make binary strong-vs-weak decisions and c…
Synthetic Hallucinations, Real Gains: Hard Negatives from Frontier Models for FIM Hallucination Mitigation
Mahdi Erfanian, Nelson Daniel Troncoso, Aashna Garg +4
Small open-source code models that power IDE autocomplete still emit hallucinated Fill-in-the-Middle (FIM) completions: syntactically natural calls to methods, parameters, variable…
Delulu: A Verified Multi-Lingual Benchmark for Code Hallucination Detection in Fill-in-the-Middle Tasks
Mahdi Erfanian, Nelson Daniel Troncoso, Aashna Garg +4
Large Language Models for code generation frequently produce hallucinations in Fill-in-the-Middle (FIM) tasks -- plausible but incorrect completions such as invented API methods, i…
LOCUS: A System and Method for Low-Cost Customization for Universal Specialization
Dhanasekar Sundararaman, Keying Li, Wayne Xiong +1
We present LOCUS (LOw-cost Customization for Universal Specialization), a pipeline that consumes few-shot data to streamline the construction and training of NLP models through tar…