Thomas Nedjar
All apps

AI agents in business: from the demo to real use

Published on 23 July 2026

Everyone has seen the demo. An agent opens a browser, fills in a CRM, books a flight, chains twelve actions without anyone touching the keyboard. It is spectacular, it travels beautifully, and the labs have all pivoted to it: Anthropic with computer use and the Model Context Protocol, OpenAI with its agents, everyone running the same way. Now the awkward question: how many companies that you know have an agent running, on a Monday morning, on a real process?

Very few, in my experience. The models are more than good enough. It is the direct continuation of what I wrote about AI adoption: between a demo that impresses and a tool you open every day, there is a chasm. With agents that chasm is wider still, because an agent does not just answer. It acts.

Why AI agents stay stuck at the demo stage

A demo is built to succeed. The path is mapped out, the data is clean, the presenter has rehearsed. A real process, on the other hand, is made of exceptions: the customer with no VAT number, the duplicate invoice, the field someone filled in by hand in 2019. The demo shows the normal case, the business lives in the edge cases.

And then there is the arithmetic nobody wants to do. An agent that is 90 % reliable per step, over a chain of twelve steps, ends up at 28 % success end to end. That is a multiplication, nothing more. As long as you stack autonomous steps with no checkpoint, you are building a machine that fails quietly.

The claim that 95 per cent of AI projects fail deserves more than a repost

You have certainly come across it: the MIT Media Lab report on the state of AI in business, and its headline number, 95 % of generative AI projects with no measurable effect on the bottom line. The sentence went round LinkedIn in forty-eight hours.

It deserves some rigour. The sample is thin, the definition of failure is vague, and a pilot that was stopped is not necessarily a pilot that failed: sometimes it is simply a company that learned something and cut at the right moment. Taking that 95 % as a measurement is the same mistake as taking a model leaderboard for a strategy.

That said, the intuition behind the number is sound, and everyone working on the ground shares it: a great many projects stop between the successful POC and going live. The figure is debatable, the phenomenon holds.

What separates an agent in production from a POC

Four things, and none of them is a model problem.

A narrow scope. An agent that does one precise thing on a bounded process reaches production. A general-purpose assistant agent stays a demo forever, because nobody can say when it has done its job well.

A human checkpoint. The agent prepares, a human approves, the agent executes. That is what makes the tool deployable in a real company, with real accountability.

Traceability. Knowing what the agent did, when, on which data, and being able to show it. Without that, no serious manager will sign off.

Reversibility. Every action must be undoable, or at the very least must destroy nothing. An agent that writes before being reviewed is a bad idea, whatever model sits behind it.

Which processes to hand to an AI agent first

A good first candidate ticks four boxes: it is repetitive, its output is verifiable at a glance, a mistake costs little, and the volume is high enough for the gain to show. Data reconciliation, file preparation, consistency checks, structured monitoring, formatting: thankless, and it holds in production.

The bad first candidate is almost always the one people pick: customer-facing work. It is visible, it makes a nice announcement, and it is precisely where a mistake costs the most. Starting there means betting your internal credibility on the hardest case.

Give an agent tools, not powers

This is the mental shift that changes everything. An agent does not need access, it needs tools, with a scope written down in black and white: what it may read, what it may propose, what it may never do alone. That is exactly what the MCP standardises: you do not plug an intelligence into a system, you expose named and bounded capabilities to it.

That is how I build my own business tools. The agent analyses, proposes a fix or an internal link, and nothing ships until a human has approved it. This safeguard is structural: it is the only setup in which anyone agrees to wire this into their production site.

An agent that acts raises the question of where it runs

As long as AI merely answered, where it ran was a conference topic. An agent that acts on your systems, reads your customer data and writes into your tools turns that into an operational question. Who sees what, under which jurisdiction, with what ability to pull the plug? It is now a line in a contract. I will come back to it.

From the demo to real use: what actually triggers adoption

Agents are hitting trust, not a technical ceiling. And trust is won by something other than twelve chained actions: it is won by showing one action, done correctly, a hundred times in a row, with the ability to check.

So when someone asks whether your company is doing agents, the right answer is a process, and to say how long it has been running without anyone having to step in.

Read next: AI adoption in business: why model power is not enough · The UCP protocol, six months on: what the data already says about agentic commerce

Thomas Nedjar
Thomas Nedjar
SEO/GEO & automation expert

Thomas Nedjar has spent fourteen years in digital acquisition and e-commerce. An entrepreneur, he founded and ran several e-commerce companies between France and Switzerland, with solid experience in international and intra-EU trade. He is now Senior SEO/GEO, AI & Tech Expert at Suisseo, and builds his own marketing applications: ED (automated community management) and SEO Cartograph (local SEO audit crawler). A speaker (CVCI) and trainer, he shares his hands-on insights on SEO, GEO, automation and e-commerce here.

All articles · LinkedIn

Latest articles

You might also like

France
Suisse