16 papers
MathDuels: A Self-Play Benchmark That Grows
Zhiqiu Xu, Shibo Jin, Shreya Arya +1
As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to differentiate model capabilities, lar…
Detecting Safety Violations Across Many Agent Traces
Adam Stein, Davis Brown, Hamed Hassani +2
To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex, and sometimes even adversar…
Do We Need Frontier Models to Verify Mathematical Proofs?
Aaditya Naik, Guruprerana Shabadi, Rajeev Alur +1
Advances in training, post-training, and inference-time methods have enabled frontier reasoning models to win gold medals in math competitions and settle challenging open problems.…
QLCoder: A Query Synthesizer For Static Analysis of Security Vulnerabilities
Claire Wang, Ziyang Li, Saikat Dutta +1
Static analysis tools provide a powerful means to detect security vulnerabilities by specifying queries that encode vulnerable code patterns. However, writing such queries is chall…
CAMEL: An ECG Language Model for Forecasting Cardiac Events
Neelay Velingker, Alaia Solko-Breslin, Mayank Keoliya +9
Electrocardiograms (ECG) are electrical recordings of the heart that are critical for diagnosing cardiovascular conditions. ECG language models (ELMs) have recently emerged as a pr…
On Improving Neurosymbolic Learning by Exploiting the Representation Space
Aaditya Naik, Efthymia Tsamoura, Shibo Jin +2
We study the problem of learning neural classifiers in a neurosymbolic setting where the hidden gold labels of input instances must satisfy a logical formula. Learning in this sett…