Skip to main content
A/B Testing compares your current agent, the Control, with a duplicated Variant using live calls. Phonely routes the configured share of calls to the variant and evaluates both agents against the same success criterion. Tests can run on the calls your agent answers, on the calls it places, or on both. Use Unit Testing for repeatable generated scenarios before exposing changes to live traffic.

Create a test

Open Testing > A/B Testing and create a test in Planned. The setup has four parts: The category—Voice, Workflow, Agent Settings, Knowledge Base, or Other—labels what you intend to compare. It does not restrict which variant settings you can edit.

Choose which calls the test runs on

A Workflow test asks which calls it should run on. Choose one: Every other category splits calls in both directions, as do workflow tests created before this choice existed. When a test includes outbound calls, use Include voicemail calls in results to decide whether outbound calls that reach voicemail count toward the result. It is off by default, so those calls are excluded until you enable it. Success-criteria options follow that choice for workflow tests: outcome tags come from the selected direction’s calls only, and an inbound test does not offer the Voicemail ended reason. When choosing successful outcomes or ended reasons, search to narrow the options. Select all N results selects the matching options and keeps any selections outside the search. If all matches are already selected, selecting it again clears only those matches. Clear the search and review the full selection before starting the test.

Edit the variant

Creating a test duplicates the control agent. Make the change you want to measure, then keep unrelated settings aligned so the result remains interpretable.
  • Most tests open the variant in the agent editor through Edit variant agent.
  • An outbound workflow test opens the variant’s call flow instead. Edit variant flow takes you to the campaign page’s call-flow canvas, where the Campaigns view is hidden and a chip marks the surface as the variant.
Select Back to tests in the notice that stays on screen while you edit to return to Testing.

Start and manage the test

Review the variant before selecting Begin Test. While a test is in progress, eligible calls are assigned to the Control or Variant according to the traffic split. The test appears in one of three sections:
  • Planned has been configured but is not routing calls.
  • In Progress is currently routing live calls.
  • Completed has reached its end criterion or was terminated.
You can adjust the traffic split while a test is in progress. Terminating a test stops new calls from being assigned to it; it does not undo calls already completed.
A/B Testing affects live traffic. Test the variant independently and begin with a traffic share appropriate for the risk of the change.

Review the result

Open a test to compare completed calls, successful calls, and success rates for the Control and Variant. You can also open the calls assigned to either arm and inspect them in Call History. Phonely calculates success from the criterion selected during setup:
  • Call Outcome counts calls with one of the selected outcomes.
  • Call Ended Reason counts calls with one of the selected ended reasons.
  • Call Duration compares whether shorter or longer calls better match the goal.
The result may also show the difference between success rates, the probability that the variant wins, a p-value, and a confidence interval. Treat these as evidence from the calls collected so far, not a guarantee about future calls. A small sample or an interval that crosses zero means the result remains uncertain. If you decide to keep the variant, apply it deliberately and complete an end-to-end test of the resulting control agent. Do not assume every variant difference should be promoted simply because one metric improved.
Confirm that the agent is receiving eligible live calls and review the configured traffic split. Random assignment can be uneven in a small sample; allow enough calls before interpreting the result.
Review the selected success criterion and the affected calls. Confirm that their outcome, ended reason, or duration matches the values configured for the test.
Review the completed-call count and confidence interval. Continue the test when appropriate, or conclude that the measured change did not produce a reliable difference within the collected sample.
Last modified on September 7, 2026