AI vs. Licensed CPAs: The 2026 Accounting Benchmark Results

October 3, 20266 min readBy The Crossing Report

Published: October 3, 2026 | By: The Crossing Report

AI vs. Licensed CPAs: The 2026 Accounting Benchmark Results

Eighteen months ago, the best AI models scored below the average CPA on accounting benchmark tasks. As of October 1, 2026, that's changed — and the first AI vs. licensed CPA accounting benchmark tells you exactly where the line now sits.

The APEX-Accounting benchmark, published October 1 by researchers at Mercor (arXiv), tested frontier AI against 12 licensed CPAs averaging 5.5 years of experience. On medium-length, well-defined accounting tasks: Claude scored 100% accuracy. Licensed CPAs ranged from 0% to 90% on the same tasks. AI completed each task in under 10 minutes; CPAs took 30 to 180 minutes. Cost per task: 10x lower for AI.

Before you read the rest, here's what this is not: it is not evidence that you should replace your staff. On the full 160-task benchmark, AI fails nearly 60% of tasks. The study's own language — "AI still can't close the books without supervision" — is accurate. What this benchmark provides is precision: for the first time, you have academic data on exactly which tasks AI wins on, and which it doesn't.


The AI vs. Licensed CPA Accounting Benchmark: What It Tested

The APEX-Accounting benchmark was designed to answer a specific question: on the accounting tasks a junior staff member handles daily, how does frontier AI compare to a licensed CPA?

The researchers tested 160 tasks across accounting functions. The tasks ranged from routine, input-defined work (reconciliations, coding, journal entries) to judgment-intensive work (complex exception handling, multi-step financial determinations). The 12 licensed CPAs in the study had an average of 5.5 years of experience — comparable to a solid mid-level hire at a 10-person firm.

The headline finding is real. On what the study calls "medium-length, well-defined" tasks — tasks with a structured input, a defined process, and a verifiable output — Claude scored 100%. The CPAs in the study scored between 0% and 90% on the same tasks. The performance gap is not marginal.

The critical qualifier: on the full 160-task benchmark, AI failed nearly 60% of tasks. This isn't a caveat. It's the boundary that defines where AI belongs in your workflow.


The Five Task Types Where AI Now Wins

The "medium-length, well-defined" profile has a practical definition. A task qualifies when it has:

  • A clear, structured input (a bank statement, a transaction list, a general ledger export)
  • A defined process (a reconciliation rule, a coding schema, a journal entry standard)
  • A verifiable output (a reconciled balance, a coded transaction file, a formatted summary)

No professional judgment about what the numbers mean. No client context. No exception reasoning. Accuracy, completeness, and process execution.

The five task types that consistently fit this profile in the benchmark:

  1. Bank reconciliations — matching transactions to bank statements against defined rules
  2. Transaction coding and categorization — classifying transactions to the chart of accounts
  3. Journal entry preparation — creating entries from source documents and defined rules
  4. General ledger posting — recording transactions to the appropriate ledger accounts
  5. Month-end close summaries — compiling and formatting period-close documentation

If your firm runs any of these as recurring workflows — weekly, monthly, or per-client — you now have academic evidence that AI outperforms licensed junior staff on them at a fraction of the time and cost.


The 60% That Still Belongs to a CPA

The tasks where licensed CPAs still win decisively share a common trait: they require judgment the AI doesn't have.

Interpreting unusual or multi-step transactions. Managing client conversations around discrepancies. Advising on tax strategy with client-specific context. Making professional determinations that carry regulatory weight. Signing off on work product with your name and license attached.

The Circular 230 guidance issued June 24, 2026 by the IRS and Treasury makes the liability framework explicit: CPAs remain personally liable for AI-assisted tax advice. "Over-reliance on AI without independent professional judgment" is a Circular 230 violation. The benchmark findings don't change that — they clarify where AI can operate within it.

This is the correct structure: AI on the defined, verifiable layer; licensed CPAs owning the judgment layer. The benchmark tells you which is which.


What This Means for a 10-Person CPA Firm

The practical question the benchmark answers for a small firm: which of your recurring workflows fit the "medium-length, well-defined" profile where AI now outperforms your junior staff?

For a 10-person firm with two or three staff accountants, the allocation model looks like this:

  • AI handles: bank reconciliations, transaction coding, journal entry prep, GL posting, month-end close summaries — the repeatable, input-defined work that currently consumes most of junior staff's hours
  • Staff accountants handle: exception review, client communication, audit-ready documentation, and the judgment calls AI flags for review
  • Partners and seniors handle: strategy, advisory work, complex returns, relationship management, and sign-off on AI-generated work product

This is not theoretical. Small CPA firms that automated their bank reconciliation and transaction coding workflows in 2025 and 2026 — including firms profiled by the AICPA's Journal of Accountancy in August 2026 — have been operating with this structure. The APEX-Accounting benchmark is the academic rationale for what those early adopters found in practice.

The freed capacity doesn't disappear. SPI Research's 2026 Maturity Benchmark (509 professional services organizations, $63 billion in combined revenue) found that firms deploying AI deeply — not just experimenting, but operationally deployed — earned 23.8% EBITDA margins, compared to 10.2% for firms using AI peripherally. The gap is nearly 14 percentage points.

The EBITDA gap isn't about whether you use AI. It's about whether you've made the specific task allocation decisions the benchmark now quantifies. Staff accountants who are no longer spending three hours per client per month on bank reconciliations have three hours available for advisory conversations, additional client load, or the exception work that actually requires their license.


One Thing to Do This Week

Pull up your firm's most recent monthly close workflow — the actual task list, not the summary. Identify the recurring tasks that have a structured input, a defined process, and a verifiable output.

Ask: of these, which ones would I trust a precise rule-follower to execute without judgment calls?

Those are your medium-length, well-defined candidates. The benchmark says AI now outperforms your licensed staff on them — completing each task in under 10 minutes at 100% accuracy. The firms already at 23.8% EBITDA made this call.

The question isn't whether to make it. It's which task you start with.


The Crossing Report covers the AI transition in professional services every Monday. The weekly issue goes deeper — subscribe here to get the full briefing.

Frequently Asked Questions

Can AI replace CPAs at my accounting firm?

Not fully — the APEX-Accounting benchmark shows AI fails nearly 60% of the full 160-task test range. But on specific, well-defined tasks — bank reconciliations, transaction coding, journal entry preparation, GL posting, and month-end close summaries — AI now scores 100% accuracy versus licensed CPAs averaging 0-90%, at 10x lower cost and in a fraction of the time. The right frame is task allocation, not replacement: deploy AI where it wins, keep licensed CPAs focused on the 60% of tasks that still require professional judgment.

What accounting tasks can AI do better than a licensed CPA in 2026?

The APEX-Accounting benchmark identifies five task types where AI now outperforms licensed CPAs with 5+ years of experience: (1) bank reconciliations, (2) transaction coding and categorization, (3) journal entry preparation, (4) general ledger posting, and (5) month-end close summaries. These are medium-length, well-defined tasks with a clear input, a defined process, and a verifiable output. AI completes each in under 10 minutes at 100% accuracy; CPAs in the study averaged 30-180 minutes with 0-90% accuracy on the same tasks.

How do I know which accounting tasks to automate at my small CPA firm?

Apply the 'medium-length, well-defined' test. Ask: does this task have a clear input (a bank statement, a transaction list), a defined process (a reconciliation rule, a coding schema), and a verifiable output (a reconciled balance, a coded file)? If yes, it fits the profile where AI now outperforms licensed staff. If the task requires professional judgment — interpreting unusual transactions, advising on tax strategy, evaluating audit exceptions — it stays with your licensed team. The five task types from the APEX-Accounting benchmark are the most common medium-length, well-defined workflows in a small CPA firm's recurring cycle.

What is the APEX-Accounting benchmark?

The APEX-Accounting benchmark is the first peer-reviewed academic study to test frontier AI models head-to-head against licensed CPAs on real accounting tasks. Published October 1, 2026 on arXiv (APEX-Accounting, Mercor), it tested Claude and other AI models against 12 licensed CPAs averaging 5.5 years of experience on 160 accounting tasks. Key finding: on medium-length, well-defined tasks, Claude scored 100% accuracy versus CPA accuracy ranging from 0% to 90%, while completing each task in under 10 minutes at 10x lower cost. On the full 160-task benchmark, AI failed nearly 60% of tasks — the supervision requirement for complex accounting work remains.

Get the weekly briefing

AI adoption intelligence for accounting, law, and consulting firms. Free to start.

Related Reading

This is the kind of intelligence premium subscribers get every week.

Deep analysis, cross-sector patterns, and the frameworks that help professional services firms make the crossing.