Skip to main content

A/B Testing

An A/B test splits live prediction traffic between two or more versions of a model, records which one served each prediction, and compares them. Use it to decide, on real traffic, whether a new model should replace the current one. A/B tests cover traditional ML models in the Model Registry.

Why Use A/B Testing?​

  • Validate a change. Try a retrained or new model on real traffic before it replaces the current one.
  • Measure the difference. Compare success rate, latency, or an outcome you score, between variants, with statistical significance.
  • Reduce risk. Send a small share of traffic to the new model first.
  • Optimize automatically. Let a multi-armed bandit route more traffic to whichever variant earns the best reward.

Concepts​

  • Variant. One deployed registry model in the test (running, or idle and woken by its first prediction), with an id you choose (for example control or challenger). Only deployed models can be picked: deploy a model in the Model Registry before adding it. One variant is the control: the current model the others are compared with.
  • Strategy. How the test picks a variant for each prediction.
  • Experiment. A statistical comparison of the control with the other variants on one metric, over the predictions made while the experiment runs.

The A/B Tests page​

Open MLOps > Model Registry and select A/B Testing in the header. The cards count Total A/B Tests and those Running and Stopped. The list shows each test's Name, Strategy, Status, Variants, Total Requests and when it was last updated, with View Details, Start, Stop and Delete. Filter by status, strategy or owner. Create A/B Test starts a new one.

In this section​

  • Routing strategies: weighted random, feature-based, multi-armed bandit and canary routing, and sticky routing.
  • Create and run a test: create, deploy, send predictions, and the test's page.
  • Recording outcomes: actuals and rewards for the predictions a test routed.
  • Experiments: statistical comparisons of the variants, and how a test is concluded.

Access​

The owner of a test and administrators can change, deploy, stop and delete it. People it is shared with can view it and its results.

Best Practices​

  • Decide the metric first. Create the experiment before you judge the variants.
  • Start small. Give a new model a small weight or a low canary percentage, then increase it.
  • Wait for enough samples. Judge on the experiment's significance, not on early numbers.
  • Record outcomes. Rewards drive bandits and reward experiments; actuals measure real accuracy.
  • Watch the Traffic tab. Rising errors or latency on a new variant are the first sign of trouble.
  • After the test, switch your application to the winning model and stop the test.