Skip to content
Free diagnostic

BlogFramework and control

Can an AI agent make mistakes? Errors, hallucinations and safeguards

Yes, an AI agent can get things wrong. The five safeguards that make errors rare, visible and recoverable: real cases, hand-over, approval, records.

By Stéphane Barrio Published on 2 min readFor SMEsMid-sized companies

"And what if it gets it wrong?" It is the right question. An AI agent can make mistakes, and denying it would be a lie. What matters is making errors rare, visible and recoverable.

The possible errors

  • Misreading: an amount wrongly extracted from a crookedly scanned invoice.
  • Inventing: a plausible but wrong answer, when information is missing — what is called a hallucination.
  • Misapplying a rule: treating an exception as a routine case.
  • Drifting over time: a tool changes, an AI model evolves, and what worked starts answering beside the point.

Even the solutions of major software publishers are no exception: on a benchmark published by Salesforce, a customer relations agent succeeded at 58% of simple tasks and 35% of multi-step tasks. Hence the importance of designing for error, rather than denying it.

The five safeguards

1. Check on your real cases, before going live

Not on made-up examples: on your invoices, your emails, your requests, with their exceptions. For a complete process, we check the agent on 15 to 30 of your real cases before putting it into service.

2. Hand over when it does not know

A good agent knows how to say "I don’t know". The rules for handing over to a person are written during scoping: beyond a threshold of doubt, for a given type of case, the agent stops and passes it on, with a summary.

3. Have whatever commits the company approved

Each action is classified by what is at stake:

Mode For what Example
The agent acts alone what is reversible and low-stakes circulating meeting notes
The agent prepares, a person approves what goes outside the company the reply to a complaint
The agent prepares, a person decides and executes what cannot be undone a payment

Paying, signing, setting a price, publishing without review: never by an agent alone.

4. Keep a record of every decision

A log of what the agent has handled, and a record of its decisions: to understand an error, fix it, and not repeat it.

5. Monitor over time

Tools change, and so do rules. Regular monitoring, with fixes and a value report, prevents slow drift. See Why an AI agent needs ongoing monitoring.

Reducing hallucinations at the source

An AI model invents mainly when it lacks information. The best remedy is to give it your sources — your rules, your prices, your procedures — and forbid it from answering outside them. That is the role of the knowledge written down for agents: Why your AI agents need a shared memory and shared rules.

What we deliver with every agent

The log, the record of decisions, access limited to what is needed, the recovery documentation in case of error, and the approval rules: all included in every project. The details: Governance.

Sources

  • Salesforce AI Research, CRMArena benchmark: 58% success on simple tasks, 35% on multi-step tasks, for an off-the-shelf agent.

Frequently asked questions

What is a hallucination?

It is a wrong answer delivered with confidence: an invented figure, reference or fact. It happens mainly when the AI model lacks information. The remedy: give it your sources, and teach it to say "I don’t know".

What error rate is acceptable?

It depends on the cost of an error. A reminder sent a day too early can be put right; a payment sent to the wrong account cannot. That is why each action is classified by what is at stake before deciding what the agent does alone.

Who is responsible if the agent gets it wrong?

The company remains responsible for what it does, with or without an agent. That is why actions that commit it remain approved by a person, and why every decision of the agent leaves a record.

Share on LinkedIn All articles

Read next

In the same category

Framework and control 3 min read

AI and the GDPR: what to check before deploying an agent

Necessary data, legal basis, retention periods, people’s rights, processors: the GDPR checklist to go through before entrusting data to an AI agent.

SMEsMid-sized companies

Flash Diagnostic · free

Start with a one-hour interview, free of charge.

A questionnaire that takes under 10 minutes, a one-hour interview, then within 72 hours a written report: what you can stop doing by hand, the time saved and the order of magnitude of the budget. If an off-the-shelf tool is enough, we will tell you.