Skip to main content
Copertina articolo: A/B test for AI agents: Compare experiences, not just answers
Articles/Experimentation

A/B test for AI agents: Compare experiences, not just answers

/

Version A: The user sees a traditional guide. Version B: The user receives help from an AI agent. After a week, the B version shows more interactions. The team is tempted to declare victory.

Wait.

With AI agents, the more interactions they automatically mean more value. Sometimes the user talks more because the agent is unclear. Sometimes he asks for confirmation because he doesn’t trust. Sometimes he explores it out of curiosity and then abandons it.

What are you really testing?

An A/B test on an agent does not compare only two screens. Compare two ways to complete a job.

The hypothesis must be behavioural and measurable:

  • “AI agent reduces the time needed to complete the setup without increasing configuration errors.”, “AI agent increases the percentage of users who interpret the report correctly and perform an action within 24 hours.”

These assumptions are more solid than “agent increases engagement.” Engagement alone can be a false friend.

Triggering: who enters the test?

Triggering is serious. You only have to include users who have really had the chance to use the agent.

If there are users in group B who have never seen the point where the agent appears, the analysis gets dirty. You are comparing unexposed people and the result loses meaning.

For example:

  • onboarding agent: enter the test who arrives at the setup;, dashboard agent: enter the report opener;, support agent: enter the application launcher.

Primary metrics and guardrail

The primary metric measures the desired result. Guardrrails protect against unwanted side effects.

Examples for a support agent:

  • primary: resolution at first contact;, guardrail: reopening, incorrect escalation, negative feedback, use of obsolete sources.

Examples for a marketing agent:

  • primary: time from insight to approved test;, guardrail: duplicate tests, incomplete briefs, heavy human corrections.

Humbly perform

If the agent wins, ask yourself who won for: new users? Experts? Mobile? Desktop? Small customers? Enterprise?

If it loses, it does not mean that AI agents do not work. It may indicate a case of wrong use, a missing context, a uninviting UI or a metric that does not capture the value.

An A/B test is not used to defend AI, but to find out where AI deserves space in the product.

How to apply it without complicated work

Do not start with the newest tool. Start with the point where the team is wasting time, discussing without data or making decisions with incomplete information. Here you can see whether the theme has operational value or is just a nice slide idea.

An agent is not a brilliant chat. He must have clear inputs, limited tools, controlled memory and an explicit rule to pass the decision on to a person when the risk grows.

A useful sequence:

  1. Define which data the agent can read and which not; 2. Write the expected result in verifiable form, not as a generic intention; 3. Decide when it needs human revision before sending or saving output; 4. Measure time saved, avoided errors and cases where the agent stops.

What to measure to see if it works

The right question is not “have we used AI?” or “have we added a new dashboard?” The question is: what decision has become faster, clearer or safer?

Measure at least three levels:

  • operational time saved;, quality of the result;, confidence of the team in the process.

Time alone can deceive: a faster but less controllable flow is not an improvement. Quality alone can deceive: a perfect system but too slow does not enter everyday work.

The point a A/B test for AI agents is a discipline of decision under uncertainty. it is not enough to see more interactions or clicks. it is necessary to define precise hypotheses, choose who enters the test, measure relevant metrics and interpret data with humility. only then it is discovered where AI brings real value and where instead it risks being only a technical decoration.

Related articles

Backlog experiments: How not to turn it into a cemetery of ideas
June 14, 20261 min read
Read
Product Experiences: They are not races, they are questions
June 14, 20261 min read
Read
Fake door test: Validate the question without fooling people
June 14, 20261 min read
Read