6 papers
IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO
Mostapha Benhenda
Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financial tasks. However, i…
Smooth Pseudo-Rotations Measure-Theoretically Isomorphic to Circle Rotations of Rationally Independent Angle
Mostapha Benhenda
Let M be a smooth compact connected manifold, on which there exists an effective smooth circle action preserving a positive smooth volume. We show that on M, the smooth closure of…
YC Bench: a Live Benchmark for Forecasting Startup Outperformance in Y Combinator Batches
Mostapha Benhenda
Forecasting startup success is notoriously difficult, partly because meaningful outcomes, such as exits, large funding rounds, and sustained revenue growth, are rare and can take y…
Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance
Mostapha Benhenda
We introduce Look-Ahead-Bench, a standardized benchmark measuring look-ahead bias in Point-in-Time (PiT) Large Language Models (LLMs) within realistic and practical financial workf…
A smooth Gaussian-Kronecker diffeomorphism
Mostapha Benhenda
We build a smooth Gaussian-Kronecker diffeomorphism, answering a question raised by Anatole Katok in his list of ''Five Most Resistant Problems in Dynamics''.
FinRL-DeepSeek: LLM-Infused Risk-Sensitive Reinforcement Learning for Trading Agents
Mostapha Benhenda
This paper presents a novel risk-sensitive trading agent combining reinforcement learning and large language models (LLMs). We extend the Conditional Value-at-Risk Proximal Policy…