8 papers
Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs
Mohammad Jalali, Azim Ospanov, Amin Gohari +1
Generative models guided by text prompts are widely evaluated for fidelity and prompt alignment, yet their ability to produce outputs remains underexplored. Existing diversity metr…
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
Azim Ospanov, Zijin Feng, Jiacheng Sun +3
Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments. However, purely informal reasoni…
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
Farzan Farnia, Mohammad Jalali, Azim Ospanov
Deep generative models have achieved great success in producing high-quality samples, making them a central tool across machine learning applications. Beyond sample quality, an imp…
miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
Azim Ospanov, Farzan Farnia, Roozbeh Yousefzadeh
We perform a thorough analysis of the formal and informal statements in the miniF2F benchmark from the perspective of an AI system that is tasked to participate in a math Olympiad…
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning
Azim Ospanov, Farzan Farnia, Roozbeh Yousefzadeh
Formal reasoning and automated theorem proving constitute a challenging subfield of machine learning, in which machines are tasked with proving mathematical theorems using formal l…
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
Azim Ospanov, Mohammad Jalali, Farzan Farnia
The use of CLIP embeddings to assess the fidelity of samples produced by text-to-image generative models has been extensively explored in the literature. While the widely adopted C…