activity
20242026
most citedOpenAI-o1 AB Testing: Does the o1 model really do good reasoning in math problem solving?

2 citations · 2 across the 3 of their papers we have counts for

collaborators

6 papers

q-fin.PM2026

Uncertainty-Adjusted Sorting for Asset Pricing with Machine Learning

Yan Liu, Ye Luo, Zigan Wang +1

Machine learning is central to empirical asset pricing, but portfolio construction still relies on point predictions and largely ignores asset-specific estimation uncertainty. We p…

cs.MA2025

AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems

Yang Li, Siqi Ping, Xiyu Chen +4

With the rapid progress of large language models (LLMs), LLM-powered multi-agent systems (MAS) are drawing increasing interest across academia and industry. However, many current M…

q-fin.PM2025

Hierarchical AI Multi-Agent Fundamental Investing: Evidence from China's A-Share Market

Chujun He, Zhonghao Huang, Xiangguo Li +5

We present a multi-agent, AI-driven framework for fundamental investing that integrates macro indicators, industry-level and firm-specific information to construct optimized equity…

cs.AI2025

Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM Agent

Chunlong Wu, Ye Luo, Zhibo Qu +1

Large language model (LLM) agents achieve impressive single-task performance but commonly exhibit repeated failures, inefficient exploration, and limited cross-task adaptability. E…

econ.EM2025

Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks

Qiang Chen, Tianyang Han, Jin Li +5

Can AI effectively perform complex econometric analysis traditionally requiring human expertise? This paper evaluates AI agents' capability to master econometrics, focusing on empi…

cs.AI20242 cited

OpenAI-o1 AB Testing: Does the o1 model really do good reasoning in math problem solving?

Leo Li, Ye Luo, Tingyou Pan

The Orion-1 model by OpenAI is claimed to have more robust logical reasoning capabilities than previous large language models. However, some suggest the excellence might be partial…