1 citations · 1 across the 11 of their papers we have counts for
4 papers · 1 filter
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
Shailja Thakur, Sungeun An, Chad DeLuca +1
A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it d…
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
Sungeun An, Swanand Ravindra Kadhe, Shailja Thakur +2
Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill ga…
Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
Kellen Tan Cheng, Anna Lisa Gentile, Chad DeLuca +1
The pervasiveness of large language models (LLMs) in enterprise settings has also brought forth a significant amount of risks associated with their usage. Guardrails technologies a…
SAUCE: Truncated Sparse Document Signature Bit-Vectors for Fast Web-Scale Corpus Expansion
Muntasir Wahed, Daniel Gruhl, Alfredo Alba +5
Recent advances in text representation have shown that training on large amounts of text is crucial for natural language understanding. However, models trained without predefined n…