McKinsey Found That 30% of Firms Got Less Productive After Deploying AI Agents. Here's Why.
Here is something the AI hype cycle doesn't say enough: artificial intelligence can make your firm slower.
McKinsey's Technology Trends Outlook 2026 found that productivity actually declined in nearly 30% of companies after their teams began using agentic AI tools. Not stalled — declined. Nearly one in three organizations that deployed AI agents came out the other side getting less done.
The same study found that 66% of agentic AI users reported measurable productivity gains. So the technology works. It's working well for most firms that deploy it. The gap between the 70% and the 30% isn't the tool. It's the implementation.
This distinction matters enormously if you're a professional services firm owner thinking about deploying AI agents this quarter — and the McKinsey data makes a specific, actionable claim about why the 30% failed.
The Three Failure Modes
McKinsey's analysis of the firms that saw productivity decline points to a consistent pattern: they deployed AI agents without the three conditions that separate outcomes.
Failure Mode 1: Unclear task scope. The agent was deployed without a defined boundary. It was given access to a workflow — email, scheduling, document processing, intake — but no one specified exactly where the agent's responsibility started and stopped. The result: the agent handled some things well and some things badly, and the team spent time managing the errors rather than using the output. What had been handled by a person who knew the full context was now handled by a system that had to guess.
The fix: before deploying any agent, write a one-paragraph scope definition. "This agent receives client intake forms, extracts the key facts, and produces a structured summary for attorney review. It does not contact clients directly and does not make recommendations. Every output is reviewed by [specific person] before use." If you can't write that paragraph, you're not ready to deploy.
Failure Mode 2: No oversight checkpoint. The firms that lost productivity often had agents operating in an end-to-end mode — receiving input, processing it, and producing a final output that went directly into a system, a client-facing document, or a downstream workflow without human review. When the agent made an error (and it will make errors), that error compounded before anyone caught it.
This is particularly dangerous in professional services. An AI agent that misclassifies a tax document, drafts a contract with the wrong jurisdiction, or routes the wrong staffing candidate to a client is creating rework that costs more time than the agent saved. The error shows up days later, the team has to backtrack, and the trust damage with the client is often disproportionate.
The fix: design a checkpoint into the workflow before any agent output reaches a client or an irreversible system action. Not a checkbox — a named person who reviews a specific output before it moves forward. "Before any draft engagement letter from the agent goes to a client, Maria reviews and approves it" is a checkpoint. "Someone should review this occasionally" is not.
Failure Mode 3: No feedback loop. The firms that got into sustained productivity decline weren't just getting bad outputs — they had no mechanism to detect and correct them. There was no process for catching agent errors, no way to report a pattern of problems, no structured system for improving the agent's behavior over time.
Compare this to the firms that saw strong gains: they built in a regular review cycle. Every two weeks or every month, someone spent an hour reviewing the last period of agent output, identifying patterns in the errors, and adjusting the agent's prompt, scope, or configuration based on what they found. The agent got better over time. The errors became less frequent. The oversight burden decreased as confidence increased.
The fix: when you deploy an agent, also deploy a review calendar. Schedule a 30-minute review at 30 days and 60 days post-launch. The agenda: what errors occurred, what patterns do we see, what one adjustment would have the most impact. Keep a simple log. Assign it to one person.
Why This Pattern Mirrors Early RPA
This isn't a new failure pattern. It appeared almost identically in the early robotic process automation wave of the 2010s, when professional services firms bought RPA licenses, pointed them at their most painful workflows, and then found themselves managing broken bots more than they were managing the workflows the bots were supposed to replace.
The firms that succeeded with RPA were the ones that spent the most time before launch defining the task, building in exceptions, and establishing human review. The firms that failed were the ones that deployed quickly, assuming the technology would adapt itself to their messy, exception-filled processes.
Agentic AI is more capable and more flexible than early RPA. That makes it more powerful. It also makes it easier to overextend. A more capable agent can get further into a broken workflow before the error becomes visible.
The Measurement Problem Underneath
There is another layer to the McKinsey finding that connects directly to how professional services firms should think about AI deployment in 2026.
Thomson Reuters' 2026 AI in Professional Services Report found that only 18% of firms currently measure AI's business value in any systematic way. That means 82% of firms using AI have no baseline against which to judge whether AI is helping or hurting.
If you can't measure, you can't detect a 30% productivity decline. You would feel it — more errors, more rework, more time spent managing outputs — but you wouldn't have the data to attribute it correctly. You might blame the tool, or the specific deployment, or the person overseeing it, rather than identifying the structural implementation gap that caused it.
The measurement fix is not complicated. Before you deploy any agentic AI:
- Record how long the workflow currently takes (per week, per month)
- Record your current error rate or rework rate
- Record the time spent on the task
After 30 days, measure the same three things. The comparison is your ROI number. If it's negative, you have an implementation problem, not a technology problem, and you can fix it.
What Good Implementation Looks Like
The 70% of firms that saw productivity gains from agentic AI share a consistent profile.
They deployed one agent doing one clearly defined task. They assigned one person to own the agent's output. They built in a review checkpoint before any output became irreversible. And they scheduled time at 30 and 60 days to evaluate performance and make adjustments.
For a 15-person accounting firm, that might look like this: one agent drafts client status update emails for accounts in the closing cycle. One staff accountant reviews the drafts each morning before they go out. Every two weeks, the owner reads the last two weeks of sent emails and identifies any that needed major revisions.
For a 12-person law firm: one agent processes intake questionnaire responses and produces a structured matter summary. One paralegal reviews each summary before it's added to the matter file. At 30 days, the supervising attorney reviews a sample and adjusts the agent's extraction criteria based on what was missed or wrong.
The pattern is the same across firm types. Specific workflow. Named owner. Human checkpoint. Feedback loop.
That's the difference between the 70% and the 30%. Not the platform. Not the model. Not the budget.
Your Next Step
Before you deploy your next AI agent, write down three things:
- Scope: The exact task this agent does and the exact point where it hands off to a human.
- Owner: The name of the person who reviews the agent's output before it becomes irreversible.
- Measurement baseline: What you will measure to know if this is working at 30 days.
If you can't write all three in under five minutes, you're not ready to deploy. Do the scoping work first. The agents will still be there when you're ready.
Get the weekly briefing
AI adoption intelligence for accounting, law, and consulting firms. Free to start.
Related Reading
- Why Your Firm's AI Still Isn't Working (It's Not the Tool's Fault)
- Why 61% of Professional Services Firms Abandon AI — And What the Other 39% Did Differently
- Your Firm Is in the 83%. Here's the Governance Framework to Get Into the 25%.
- Thomson Reuters 2026 AI in Professional Services Report: What the Data Means for Small Firms
- What AI Governance Actually Costs a Small Professional Services Firm in 2026
This is the kind of intelligence premium subscribers get every week.
Deep analysis, cross-sector patterns, and the frameworks that help professional services firms make the crossing.