34 papers
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
Jiyuan Tan, Vasilis Syrgkanis
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research lo…
Order-Explicit Linearization of High-Dimensional -Statistics
David M. Ritzwoller, Vasilis Syrgkanis
The paper derives explicit large‑deviation bounds for high‑dimensional U‑statistics, showing that their deviation from the Hájek projection scales as O_p(ϕ b n⁻¹ log²(dn)) and appl…
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
Hanlin Zhang, Jikai Jin, Vasilis Syrgkanis +1
Machine learning model performance improvements tend to arise from competition and application. For deployment, we consider prescriptive scaling laws: given a pre-training compute…
CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation
Ayush Sawarni, Jiyuan Tan, Vasilis Syrgkanis
Many benchmarks for automated causal inference evaluate a system's performance based on a single numerical output, such as an Average Treatment Effect (ATE). This approach conflate…
The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice
Jikai Jin, Vasilis Syrgkanis
Offline evaluation of language models from usage logs is biased when model choice is confounded: the same user-side factors that influence which model is used can also influence ho…
Partial Identification of Policy-Relevant Treatment Effects with Instrumental Variables via Optimal Transport
Jiyuan Tan, Jose Blanchet, Vasilis Syrgkanis
Policy-Relevant Treatment Effects (PRTEs) are generally not point-identified under standard Instrumental Variable (IV) assumptions when the instrument generates limited support in…