Almost every company has AI now. Far fewer have results.
McKinsey’s November 2025 survey found 88 percent of respondents saying their organization regularly uses AI in at least one function, while only 39 percent reported any EBIT impact at the enterprise level, most of them under 5 percent. Deloitte asked 1,854 senior executives, all at organizations already running AI in production, how long payback took. Six percent said under a year.
That six percent is the number I’d sit with. Everyone in that survey had already shipped something.
So access is no longer the constraint. Value is.
I get asked about this a lot, and the diagnosis is simpler than people want it to be. Most organizations bought a tool and kept everything else the same. As I put it to Newsweek earlier this month, people treat it like a plug-and-play tool. The technology showed up and the work didn’t change.
McKinsey has the receipts. When it tested 25 organizational attributes against whether companies actually saw EBIT impact, the one that mattered most was whether they had redesigned the workflow. Only 21 percent of the companies already using it had.
There are two halves to closing that gap. This piece is the first: what has to exist around a model before it does real work. The second is about what happens to your team once it does.
What tool thinking costs
Two well-documented cases, in both of which a model went in and nothing else moved.
In February 2024, Klarna announced its AI assistant had handled two thirds of customer service chats in its first month, the equivalent work of 700 full-time agents, with $40 million in projected profit improvement. Fourteen months later the CEO told Bloomberg that cost “seems to have been a too predominant evaluation factor when organizing this,” and that what you end up having is lower quality. They began hiring human agents again.
The tool worked; the redesign around it didn’t.
The same year, British Columbia’s Civil Resolution Tribunal decided Moffatt v. Air Canada. A grieving customer asked the airline’s chatbot about bereavement fares, and the bot described a refund policy that did not exist. The tribunal found negligent misrepresentation and awarded damages.
Its reasoning is the part worth reading. The airline, the tribunal wrote, “suggests the chatbot is a separate legal entity that is responsible for its own actions.” It called that “a remarkable submission,” noting that a chatbot “is still just a part of Air Canada’s website.” Small claims, so treat it as directional.
Liability survives the demo.
Was the technology at fault in either case? No. In both, the capability was fine, and what was missing sat around it.
Three ways to think about AI
Tool. You go to it. ChatGPT in a browser tab is a destination, and you have to remember to walk over there. The gain is faster tasks.
Assistant. It lives inside the software you already use, like Copilot in Outlook or Claude in Excel. The gain is faster work.
System. Nobody opens anything. Something happens in the business (an invoice arrives, a permit gets filed, a meeting ends) and the work runs. The gain is scaled output.
Most of the value sits at that third level, which is also the level almost nobody has reached. Adoption is not a switch, either.
| Level | Name | What it looks like | Maps to |
|---|---|---|---|
| 0 | Not started | No real usage | |
| 1 | Prompting | Copying and pasting into a chat window | Tool |
| 2 | Embedded | Using the AI built into tools you already pay for | Assistant |
| 3 | Orchestrated | Agents running defined workflows | System |
| 4 | Systematized | It’s simply how operations run | Operating model |
Most organizations we meet sit between 1 and 2 and believe they’re further along.
What an agent actually is
“Agent” and “agentic” are getting stretched past usefulness. At the core, an agent is four things.
- A goal. Something concrete and finishable. “Review this lease.” “Draft the journal entry.”
- Multiple steps. It chains actions instead of doing one thing and stopping. Decide, do, check, adjust, until it’s done.
- Tools and data. A PDF reader, your CRM, a database, an API. Whatever the work requires.
- An output. A record, a summary, or a decision handed to a human. Structured, and landed somewhere.
Structured work, executed automatically.
The harness
On their own, agents aren’t enough, and the failure modes are consistent. They lose context between steps and sessions, and give different answers to the same input. They can reason but can’t reliably reach the systems where work has to land. And while they understand a goal and an output, they usually don’t know what good looks like for either.
Which of those does a better model fix? None of them. They’re fixed by the harness, which is everything you build around the model.
Context and memory. What the model sees each turn: instructions, the relevant documents, the history, knowledge retrieved on the fly. It arrives with general knowledge only, so somebody decides what else it gets.
Tools. The model’s hands. Read and write files, query a database, hit an API, send the email, run the code.
The loop. Think, act, observe, repeat, until the job is actually done. This is what makes something agentic rather than conversational.
Guardrails. Permissions, validation, retries, and a human in the loop before anything irreversible; plus an audit log of every action taken.
For one workflow, that’s four decisions and you make all four. Hand it the policy document. Let it query the ledger. Stop it before it posts anything, and log what it touched.
Not one of those is something the model provides. They’re specifications, and writing them is the work.
Agents are the brains; harnesses are the operating system.
If you’ve ever watched an impressive AI demo fail to survive contact with your actual business, the harness is almost always what was missing.
And the boring half of the harness
There’s a second missing piece, less interesting and just as decisive. Call it glue. So where does adoption actually stall? Usually at the last mile, because the systems don’t talk to each other and somebody ends up stitching the work together by hand.
At most medium-to-large organizations that glue already exists, in the connectors to your line-of-business systems (Microsoft Dynamics, Power Automate and Copilot, or Google BigQuery, Looker and Gemini). It’s the boring layer that closes the distance between a chat window and where work actually lives.
Why the harness is a reliability question, not a polish question
AI is a skill amplifier, and sometimes it’s a trap. In a 2023 Harvard Business School and BCG study, 758 consultants were randomized across a set of consulting tasks. The group using AI finished about 25 percent faster and produced more than 40 percent higher quality output, and the below-average performers improved most, by 43 percent.
On one task deliberately placed outside what the model could reliably do, AI users were 19 percentage points less likely to get the right answer.
Task-level results, not firm-level productivity, and a 2023-era model. But the shape holds, and my read of that last number is that people kept trusting the output on exactly the task where they shouldn’t have. Guardrails, validation and a human check on the calls that matter are what turn an amplifier into infrastructure.
So build the harness. Then the more interesting problem arrives.
Because a harness that works takes real work off real people, and those freed hours do not turn into money by themselves. That’s to come soon, in part two.
Wells Stringham has spent 18+ years building for Nike, Disney, lululemon, TELUS, and the NFL. He's Co-Founder & Partner, Experience at better&co, a digital collective working on strategic clarity, AI-native product development, experience innovation, and operational evolution. betterand.co