collaborators

13 papers

cs.LG2026

SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant

Adel Javanmard, David P. Woodruff, Vahab Mirrokni

Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD,…

cs.LG2026

Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness

Aria Masoomi, Mahsa Bazzaz, Adel Javanmard +1

Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps. While much attention has been paid to th…

cs.LG2026

Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data

Kareem Amin, Rudrajit Das, Alessandro Epasto +4

The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. Ho…

cs.AI2026

Aletheia tackles FirstProof autonomously

Tony Feng, Junehyuk Jung, Sang-hyun Kim +14

We report the performance of Aletheia (Feng et al., 2026b), a mathematics research agent powered by Gemini 3 Deep Think, on the inaugural FirstProof challenge. Within the allowed t…

cs.CL2026

Accelerating Scientific Research with Gemini: Case Studies and Common Techniques

David P. Woodruff, Vincent Cohen-Addad, Lalit Jain +33

Recent advances in large language models (LLMs) have opened new avenues for accelerating scientific research. While models are increasingly capable of assisting with routine tasks,…

cs.LG2026

Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models

Adel Javanmard, Baharan Mirzasoleiman, Vahab Mirrokni

Large Language Models (LLMs) are pretrained on massive datasets and later instruction-tuned via supervised fine-tuning (SFT) or reinforcement learning (RL). Best practices emphasiz…