Research

Advanced Financial Reasoning at Scale: A Comprehensive Evaluation of Large Language Models on CFA Level III

Pranam Shetty, Abhisek Upadhayaya, Parth Mitesh Shah, Srikanth Jagabathula, Shilpi Nayak, Anna Joo Fee

The abstract states that financial institutions are increasingly adopting Large Language Models (LLMs), making rigorous domain-specific evaluation critical. This paper presents a comprehensive benchmark evaluating 23 state-of-the-art LLMs on the Chartered Financial Analyst (CFA) Level III exam, which is the gold standard for advanced financial reasoning. The evaluation reveals that leading models demonstrate strong capabilities, with composite scores such as 79.1% (o4-mini) and 77.3% (Gemini 2.5 Flash) on CFA Level III. These results indicate significant progress in LLM capabilities for high-stakes financial applications and provide crucial guidance for practitioners on model selection, while also highlighting remaining challenges in cost-effective deployment and the need for nuanced interpretation of performance against professional benchmarks.