Every week we talk to companies that have already done “something with AI”. An internal chatbot, a copilot for the sales team, a pilot with a big vendor. Almost all of them tell the same story: the demo was impressive, the pilot looked promising — and six months later nobody uses it.
The model was never the problem
Today's models are more than enough for most operational tasks. What's missing is everything else: real access to the systems where the work lives, clear rules about what the AI can touch, and someone who answers when something turns out differently than expected.
A pilot running in a sandbox with test data proves nothing. The right question isn't “can the model do this?” but “can it do it inside our ERP, with our permissions, leaving a trace that audit will accept?”.
The demo measures the model's capability. Production measures the organization's capability to govern it.
What the ones that make it do differently
In the projects that do reach production we see three repeated decisions:
- They pick a critical process, not a curiosity. Where there's volume, people-hours and a measurable outcome. If the process doesn't matter, neither does the pilot.
- They connect the real systems from day one. The agent works against the actual ERP with read-only permissions at first, and earns capabilities as it shows judgment.
- They define governance before functionality. Roles, limits, human approvals for sensitive actions and traces of every action. That's what turns a demo into a working tool.
The measure of success: it becomes invisible
An agent in production goes unnoticed. Appointments show up booked, quotes go out in minutes, the daily summary arrives at 7 PM. Nobody talks about AI: they talk about the work they no longer have to do.
That's the bar. If three months in the team is still “trying the pilot”, one of the three decisions above was missing. If nobody remembers there's an agent in the middle, it reached production.