Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
VCBench: Benchmarking LLMs in Venture Capital
Rick Chen, Joseph Ternasky, Afriyie Samuel Kwesi +7
Benchmarks such as SWE-bench and ARC-AGI demonstrate how shared datasets accelerate progress toward artificial general intelligence (AGI). We introduce VCBench, the first benchmark…
cs.AI2025
LLM-AR: LLM-powered Automated Reasoning Framework
Rick Chen, Joseph Ternasky, Aaron Ontoyin Yin +3
Large language models (LLMs) can already identify patterns and reason effectively, yet their variable accuracy hampers adoption in high-stakes decision-making applications. In this…