We're handing this edition of Nybble Bytes over to someone special today: Gustavo Castenetto, CEO and Founder of Nybble Group.
A few weeks ago, he sat on a panel with three other people who've actually shipped AI into production, and came away with more questions than tidy takeaways. Through this article, he will share his experience.
AI Looks Different From the Inside
By Gustavo Castenetto | CEO, Nybble Group
Earlier this year I sat on a panel in Atlanta with four people who have actually shipped AI into production. A CPO with experience running large-scale technology programs at global consumer goods companies. A technology executive who has led digital transformation inside financial services. A partner at a large law firm where every AI decision carries professional liability.
Nobody was there to sell anything. Nobody had a finished playbook. That made the conversation more useful than most.
Here is what I took away.
The pilot isn't the proof.
This came up early and stayed in the room all night. Pilots succeed on clean data, controlled conditions, and motivated teams. Production is messier. Real volume. Real edge cases. Real people who weren't in the room when the thing was designed.
One of the panelists, who has run large-scale transformation inside financial services, said it plainly: the MVP works well, but when you deliver in production the results are often awful. You need to be prepared for that. It is going to fail in some way. The question is whether you've designed for that reality or assumed your way past it.
Another panelist had specifics. A supply chain chatbot at a major consumer goods company delivered strong pilot results. They knew going in that without rebuilding the underlying architecture and scaling across regions, not just one or two countries, the ROI wouldn't be there. The pilot worked. The business case required a much bigger bet. Those are two different decisions and they don't always get made at the same time.
Then there was the diaper sensor. A startup approached a large consumer goods company with a reusable sensor that could detect when a baby needed a change. Real use case. Genuine customer problem. The machine learning algorithm never got above 50% accuracy. They needed 80-plus to ship. They never got there.
That story doesn't end with a postmortem. It ends with a question nobody in most organizations has answered yet: at what accuracy threshold does an AI-assisted decision become trustworthy enough to act on? And who decides?
The technology usually isn't the problem.
What slows AI adoption isn't the model. It's the organization.
The 80 or 90 percent of people who sit in the middle, not early adopters, not resistors, are watching. They're watching whether the people who tried things early got supported or exposed.
They're watching whether leadership is honest about what worked and what didn't. That signal sets the tone for everything that follows.
In professional services this is more explicit. One panelist, a partner at a large law firm, put it directly: when you're a lawyer, your professional reputation is built over decades. Asking someone to trust an output they didn't produce, on work where they're personally liable, is not a training problem. It's a trust problem. You solve it differently.
The same dynamic exists everywhere. Surgeons. Engineers. Underwriters. Any professional whose identity is tied to the quality of their judgment will encounter AI as a challenge to that identity before they encounter it as a tool. The organizations moving fastest are the ones that acknowledged this early.
Governance is not a compliance exercise.
Before AI, technology failures were traceable. Something broke, you followed the chain, you found the cause.
AI makes probabilistic decisions in places you didn't anticipate, at a speed that makes real-time human review impossible. That changes what governance means.
There's a well-documented case of an airline whose AI chatbot hallucinated a refund policy and communicated it to a customer as fact. The airline was legally compelled to honor it. Nobody owned the output of that agent. Nobody had defined the human-in-the-loop requirement for customer-facing commitments.
That is not a story about a bad model. It is a story about an absent governance architecture.
If you haven't defined, per workflow, in writing, before deployment, which decisions an agent can make autonomously and which require human confirmation, you will define it in a crisis. And in a crisis the definition costs more.
Someone has to be accountable. Name them.
I'll tell you what I told the room.
That morning I was driving to the event. My navigation app suggested a route I didn't recognize. I ignored it. Sat in traffic for ten minutes. Eventually took the route. It was faster. The app was right.
Everyone in that room had done the same thing. We override recommendations we don't trust even when they're better than our judgment. That friction is manageable when the stakes are a commute.
Now scale it. Agents making decisions across workflows. Humans second-guessing outputs out of habit, professional identity, or genuine uncertainty about who is accountable if something goes wrong.
The line between autonomous action and required human oversight has to be drawn explicitly. Per workflow. Per decision type. Before deployment. Because leaving it undefined is itself a governance decision, just a bad one.
One of the panelists put it simply: the agent always has a human as a boss. Whoever owns the workflow owns everything the agent does inside it. That accountability doesn't distribute across a team. It concentrates in a person.
The lawyer on the panel closed it from where he works every day: human intelligence always has to trump artificial intelligence. That's how you lead through all of this. You don't hand over your professionalism to it.
The cost of waiting
One observation from the panel that stayed with me: companies that waited to see if they needed a website didn't survive the internet. Companies that delayed the move to cloud paid for it. This wave is moving faster and steeper than either of those.
The models will commoditize. That's already happening. The investment that holds its value is the application layer: the workflows, the governance, the organizational fluency. Build there.
And underneath all of it: we are all beginners. Every company, every leader, every team. The honest response to that isn't false confidence. It's the willingness to try things, be wrong, and correct course without it becoming a crisis of credibility.
These are the conversations we're having every day at Nybble.
If the pilot-to-production gap sounds familiar, if you have agents in flight but accountability still undefined, if your teams are technically ready but organizationally somewhere behind that — those are not unique problems. They are the operational reality for most mid-market companies trying to move at the speed this moment requires.
Nybble Labs exists because we believe you can't advise clients on what works if you're not actively building and stress-testing it yourself. We're not watching from the outside. We test, experiment, build, and learn alongside the organizations we work with.
If that's the conversation you want to have, we should talk. Let’s connect!