As of Oct 2, Mercor's APEX study reported by The Decoder found Claude Opus 5.5 leading at 61.8% on the 160-task accounting benchmark, though no model fully solved nearly 60% of tasks.
A Mercor study reported by The Decoder had 12 licensed CPAs average about 37% on simplified APEX Accounting Benchmark tasks 18 months ago; today's best models solve those tasks almost flawlessly, faster and cheaper. On the full 160-task benchmark built with 40+ professionals, Claude Opus 5.5 leads at 61.8% of grading criteria met, ahead of Fable 5.1 at 61.0% and GPT-6 Astra at 57.9%, yet no model fully solved nearly 60% of tasks. Mercor says models still can't close the books without oversight because the study omits client communication, collaboration and accumulated context.
Items older than 72 hours never go in Top stories and show their original date.