I Ran The Same AI Agents Twice. They Gave Me Different Answers

0
3

I have spent fifteen years using data to predict what customers will do. CRM programs, personalization engines, campaign forecasting, retention models, across pharma, luxury, travel and e-commerce. In that time, I have sat through more vendor demos than I care to count, and lately every single one of them says the same word: agents. Software that profiles your data, makes the decision, checks its own work, and hands you the answer. Which offer to send. Which customer is about to leave. What next quarter looks like.

I wanted to know if it held up. Not in a demo. In the boring, repetitive, hundreds-of-decisions way that real customer and commercial work happens.

So, I built one myself — a small team of AI agents wired together the way I would staff an actual analytics team, because after fifteen years I don’t know any other way to think about it. One agent inspects the data, like the junior analyst you trust to spot the gaps. One produces the forecast and the scenarios around it. One plays the skeptical reviewer, the person in the room who asks, “Would you actually bet budget on this number?” I pointed the whole thing at a public retail dataset, hundreds of forecasting tasks, and let it run.

Then I did the thing nobody does with a pilot. I ran it again. Same tasks, same data, same everything.

If I only showed you the summary page, you would sign off on this system. Overall accuracy across the two runs was close enough that you would call it stable. And that summary page is precisely what gets shown to steering committees. It is what I would have shown a client ten years ago, if I am honest.

Then I started comparing individual tasks. Same product, same history, same instructions. Different answers.

Not wildly different, most of the time. But the ranges moved. Some judgments flipped. Enough that the advice you got for any single product depended on which day the pipeline happened to run.

Now translate that into customer experience terms, because the same class of system is being sold into your CX stack right now. An agent deciding which customers get the retention offer. An agent choosing the next best action for a journey. If it reaches a different conclusion about the same customer on Tuesday than it did on Monday, with nothing about that customer having changed, your personalization is no longer a strategy. It is a lottery with good branding. And unlike my forecasting experiment, your customer feels the inconsistency directly.

“If it reaches a different conclusion about the same customer on Tuesday than it did on Monday, with nothing about that customer having changed, your personalization is no longer a strategy. It is a lottery with good branding.”

Here is the thing that took me a while to accept: this is not a bug I failed to fix. The models underneath these agents are not fully deterministic, and at least one major provider has quietly removed the settings that are used to let you force consistent behavior. I went looking for the knob to turn it off. There isn’t one anymore. You either design your process around the variability, or you pretend it isn’t there.

The reviewer agent was the part I was proudest of, going in. A built-in challenge function. It genuinely worked, in one sense: it revised a large share of the output that came to it, and some of those revisions were sensible.

But its final verdict, task after task after task, was “use with caution.” Hundreds of decisions. Nearly identical stamp on all of them.

I have managed human reviewers who did exactly this. The one who signs everything off with a mild disclaimer so that nothing is ever their fault. With a person, you notice after a few weeks and you have a conversation. The agent version of that person will do it forever, politely, at scale, and it will never occur to it that a warning applied to everything is a warning applied to nothing. I only caught it because I counted. Most organizations wiring agents into customer journeys are not counting.

“A warning applied to everything is a warning applied to nothing.”

If I were sitting on the buying side of the table, three questions.

Run it twice and show me what moved. Not the accuracy score. The individual answers, per customer, per product, per decision. Any team can produce one good run. If they look uncomfortable at the request, that discomfort is your answer.

Show me behavior at the level my customers experience it. Aggregate uplift is how these systems get sold. A single customer’s journey is how they get used to it. If assurance only happens at the portfolio level, nobody has looked at the thing your customer will actually touch.

Show me what the agent’s self-assessment correlates with. My reviewer worked hard and told me nothing. Both things were true at once, and I would never have separated them without measuring each behavior on its own.

Most of my time on this project did not go into anything clever. It went into what happens when a batch dies halfway through and the system has kept no record of where it got to. Into cleaning invisible characters out of model output before they broke everything downstream. At one point I lost days to a dataset sitting in the wrong cloud region, which had nothing to do with AI at all and everything to do with me moving too fast.

That is what production looks like. Not a smarter prompt. Plumbing, restarts, and a steady background rate of mundane failure that you either engineer for or get ambushed by. If that is true in a controlled experiment, it is doubly true in a live CX environment where the failure lands in a customer’s inbox.

I am not writing this to warn anyone about agentic AI. The technology is real, and being able to encode an entire analyst workflow, including a challenge step, into software is something I could not have done even three years ago. Parts of what I built impressed me.

But I have watched this industry evaluate new technology the same way for fifteen years: one flattering run, reported at the aggregate, judged on a single metric, decided in a room where nobody asks whether the demo would give the same answer tomorrow.

“One flattering run, reported at the aggregate, judged on a single metric, decided in a room where nobody asks whether the demo would give the same answer tomorrow.”

That is not a reason to walk away. It is a reason to ask cheaper questions earlier. Does it agree with itself? Does its confidence mean anything? Who notices when it fails on a Tuesday night, and which customers were at the receiving end when it did?

I got my answers for the price of running the experiment twice. That is the best money I have spent on AI so far.

For more details, visit: https://theleadershipchronicle.com/

If you would like to get featured or want to know more, connect with us at contact@theleadershipchronicle.com

LinkedInn : https://www.linkedin.com/company/the-leadership-chronicle/

البحث
الأقسام
إقرأ المزيد
أخرى
Global Chlorinated Paraffins for Rubber and Textile Market Growing at 1.2% CAGR Through 2032
According to a new report from Intel Market Research, the global Chlorinated Paraffins for Rubber...
بواسطة Subhayan Mayra 2026-07-24 10:45:40 0 175
أخرى
RF Testing Innovations Accelerating Anechoic Chamber Industry Development
The demand for advanced electromagnetic testing environments is steadily increasing as industries...
بواسطة Pratiksha Mkam 2026-05-22 09:35:20 0 282
Networking
Next-Generation Forex and Remittance Solutions for Digital Payments
The global financial landscape is rapidly evolving with the rise of digital payment technologies....
بواسطة Shaheedra Salkhan 2026-08-06 08:26:42 0 17
Music
Fence Paint Online Slot Games with Unique Designs and Special Bonuses
Fence paint is far more than a decorative finish; it's an important protective coating that...
بواسطة Hamza Khatri 2026-07-18 08:47:49 0 129
Theater
Online Slot Games: A total Guidebook for you to Participating in Correctly along with Replacing the same with Possibilities
On-line port online games are getting to be the most common varieties of digital camera leisure...
بواسطة Mashr Beda 2026-08-04 09:36:20 0 141
BuzzingAbout https://www.buzzingabout.com