APEX Accounting Benchmark

Product · 1 update · last updated Oct 2

Where it stands

As of Oct 2, Mercor's APEX study reported by The Decoder found Claude Opus 5.5 leading at 61.8% on the 160-task accounting benchmark, though no model fully solved nearly 60% of tasks.

Timeline · newest first
  1. Oct 2
    Reported by The Decoder

    Frontier models beat licensed accountants on speed and accuracy on APEX, but can't close the books alone, per The Decoder

    A Mercor study reported by The Decoder had 12 licensed CPAs average about 37% on simplified APEX Accounting Benchmark tasks 18 months ago; today's best models solve those tasks almost flawlessly, faster and cheaper. On the full 160-task benchmark built with 40+ professionals, Claude Opus 5.5 leads at 61.8% of grading criteria met, ahead of Fable 5.1 at 61.0% and GPT-6 Astra at 57.9%, yet no model fully solved nearly 60% of tasks. Mercor says models still can't close the books without oversight because the study omits client communication, collaboration and accumulated context.

    Source: The Decoder
How we label
Official
An institution announces something that has already happened, through its own channel or a founder or executive account.
Statement by …
A named person's views, predictions, plans or comments, including remarks reported by media when the original is unavailable.
Reported by …
Only media or another third party is saying it. We name the outlet.
Unconfirmed
Circulating online (posts, comments, leaks). No company, named person or outlet has confirmed it.

Items older than 72 hours never go in Top stories and show their original date.

DailyCitedRSSAboutPrivacy