I Ran The Same AI Agents Twice. They Gave Me Different Answers

0
4

I have spent fifteen years using data to predict what customers will do. CRM programs, personalization engines, campaign forecasting, retention models, across pharma, luxury, travel and e-commerce. In that time, I have sat through more vendor demos than I care to count, and lately every single one of them says the same word: agents. Software that profiles your data, makes the decision, checks its own work, and hands you the answer. Which offer to send. Which customer is about to leave. What next quarter looks like.

I wanted to know if it held up. Not in a demo. In the boring, repetitive, hundreds-of-decisions way that real customer and commercial work happens.

So, I built one myself — a small team of AI agents wired together the way I would staff an actual analytics team, because after fifteen years I don’t know any other way to think about it. One agent inspects the data, like the junior analyst you trust to spot the gaps. One produces the forecast and the scenarios around it. One plays the skeptical reviewer, the person in the room who asks, “Would you actually bet budget on this number?” I pointed the whole thing at a public retail dataset, hundreds of forecasting tasks, and let it run.

Then I did the thing nobody does with a pilot. I ran it again. Same tasks, same data, same everything.

If I only showed you the summary page, you would sign off on this system. Overall accuracy across the two runs was close enough that you would call it stable. And that summary page is precisely what gets shown to steering committees. It is what I would have shown a client ten years ago, if I am honest.

Then I started comparing individual tasks. Same product, same history, same instructions. Different answers.

Not wildly different, most of the time. But the ranges moved. Some judgments flipped. Enough that the advice you got for any single product depended on which day the pipeline happened to run.

Now translate that into customer experience terms, because the same class of system is being sold into your CX stack right now. An agent deciding which customers get the retention offer. An agent choosing the next best action for a journey. If it reaches a different conclusion about the same customer on Tuesday than it did on Monday, with nothing about that customer having changed, your personalization is no longer a strategy. It is a lottery with good branding. And unlike my forecasting experiment, your customer feels the inconsistency directly.

“If it reaches a different conclusion about the same customer on Tuesday than it did on Monday, with nothing about that customer having changed, your personalization is no longer a strategy. It is a lottery with good branding.”

Here is the thing that took me a while to accept: this is not a bug I failed to fix. The models underneath these agents are not fully deterministic, and at least one major provider has quietly removed the settings that are used to let you force consistent behavior. I went looking for the knob to turn it off. There isn’t one anymore. You either design your process around the variability, or you pretend it isn’t there.

The reviewer agent was the part I was proudest of, going in. A built-in challenge function. It genuinely worked, in one sense: it revised a large share of the output that came to it, and some of those revisions were sensible.

But its final verdict, task after task after task, was “use with caution.” Hundreds of decisions. Nearly identical stamp on all of them.

I have managed human reviewers who did exactly this. The one who signs everything off with a mild disclaimer so that nothing is ever their fault. With a person, you notice after a few weeks and you have a conversation. The agent version of that person will do it forever, politely, at scale, and it will never occur to it that a warning applied to everything is a warning applied to nothing. I only caught it because I counted. Most organizations wiring agents into customer journeys are not counting.

“A warning applied to everything is a warning applied to nothing.”

If I were sitting on the buying side of the table, three questions.

Run it twice and show me what moved. Not the accuracy score. The individual answers, per customer, per product, per decision. Any team can produce one good run. If they look uncomfortable at the request, that discomfort is your answer.

Show me behavior at the level my customers experience it. Aggregate uplift is how these systems get sold. A single customer’s journey is how they get used to it. If assurance only happens at the portfolio level, nobody has looked at the thing your customer will actually touch.

Show me what the agent’s self-assessment correlates with. My reviewer worked hard and told me nothing. Both things were true at once, and I would never have separated them without measuring each behavior on its own.

Most of my time on this project did not go into anything clever. It went into what happens when a batch dies halfway through and the system has kept no record of where it got to. Into cleaning invisible characters out of model output before they broke everything downstream. At one point I lost days to a dataset sitting in the wrong cloud region, which had nothing to do with AI at all and everything to do with me moving too fast.

That is what production looks like. Not a smarter prompt. Plumbing, restarts, and a steady background rate of mundane failure that you either engineer for or get ambushed by. If that is true in a controlled experiment, it is doubly true in a live CX environment where the failure lands in a customer’s inbox.

I am not writing this to warn anyone about agentic AI. The technology is real, and being able to encode an entire analyst workflow, including a challenge step, into software is something I could not have done even three years ago. Parts of what I built impressed me.

But I have watched this industry evaluate new technology the same way for fifteen years: one flattering run, reported at the aggregate, judged on a single metric, decided in a room where nobody asks whether the demo would give the same answer tomorrow.

“One flattering run, reported at the aggregate, judged on a single metric, decided in a room where nobody asks whether the demo would give the same answer tomorrow.”

That is not a reason to walk away. It is a reason to ask cheaper questions earlier. Does it agree with itself? Does its confidence mean anything? Who notices when it fails on a Tuesday night, and which customers were at the receiving end when it did?

I got my answers for the price of running the experiment twice. That is the best money I have spent on AI so far.

For more details, visit: https://theleadershipchronicle.com/

If you would like to get featured or want to know more, connect with us at contact@theleadershipchronicle.com

LinkedInn : https://www.linkedin.com/company/the-leadership-chronicle/

Suche
Kategorien
Mehr lesen
Spiele
A Beginner's Guide to Unlocking bet365 Welcome Offers Successfullyg
Joining an online gaming platform for the first time can be exciting, especially when welcome...
Von Lotus 365 2026-06-29 07:07:41 0 677
Networking
QuickBooks Time Mobile Login: Quick Access Guide
Managing employee schedules, tracking work hours, and approving timesheets becomes easier when...
Von Quickbooks Upportnet 2026-06-30 16:31:46 0 454
Andere
Europe Liquid Chromatography Devices Market Size, Trends Analysis and Forecast by 2032
According to the latest report published by Data Bridge Market Research, the Europe...
Von Ankita Patil 2026-07-17 10:27:38 0 163
Food
Chicory Root Inulin Market Set to Reach USD 3.87 Billion by 2036
As global food manufacturers accelerate sugar reduction initiatives and governments tighten...
Von Satyam Harishchan 2026-07-07 10:03:33 0 255
Health
Outsource Doctor Credentialing to Cut Admin Costs
Administrative bloat is quietly draining the profitability of modern healthcare practices....
Von Salman Ahmad 2026-08-10 17:21:01 0 10
BuzzingAbout https://www.buzzingabout.com