Use ABlyftResults & analysis

Statistical methods

How ABlyft calculates results, which methods you can choose and which rules decide a winner.

ABlyft compares every variation with the baseline and calculates how sure it can be that the difference is real and not random. You choose between three methods: Frequentist, Bayesian and Legacy. On top of the method, three settings decide when a variation may be called a winner or loser: the confidence level, the minimum conversions and the minimum runtime.

The baseline

The baseline is the variation that all others are compared with. Normally this is Original, the unchanged version. It is marked with "Baseline" on the results page. Improvement, certainty and statistical state are only calculated for the other variations.

Two kinds of metrics

ABlyft uses a method for each kind of goal:

KindGoalsCompared numbers
Binary metricsGoals that count conversions, such as pageview, click and custom goals with measurement type Conversion.Conversion rates
Continuous metricsGoals that measure a value, such as revenue goals and custom goals with measurement type Value.Value per participant

For both kinds you select a method separately, so a revenue goal can be evaluated differently than a click goal.

The three methods

Frequentist

The default for new projects. It asks: if there were no real difference, how likely would we see a difference this large?

  • Conversion goals: a two-sample z-test on the conversion rates.
  • Value goals: Welch's t-test on the value per participant. Participants without a value count as zero.
  • Goals with counting method Every, where the rate can exceed 100%: a Poisson rate z-test.
  • Certainty is calculated from the p-value (two-sided). The result also shows the p-value and the z- or t-score.

Bayesian

It asks: how likely is it that this variation is better than the baseline?

  • Conversion goals: a Beta-Binomial model.
  • Value goals: a model for means with a weak prior.
  • Both use 10,000 simulations, so the numbers can differ slightly after a recalculation.
  • Certainty is the probability that the variation is better than the baseline. For goals with winning direction Decrease, it is the probability that the variation is lower. The table shows Credible Interval instead of Confidence Interval.

Bayesian with counting method Every

If the conversion rate of a goal is above 100% (possible with the counting method Every), the Bayesian method cannot calculate a result for conversion goals. The certainty stays at 50% and the variation stays inconclusive. Choose Frequentist for such goals, or use the counting method One.

Legacy

The method of earlier ABlyft versions, kept so that existing experiments stay comparable. It uses a classic z-test with the standard errors of both groups. New experiments should use Frequentist or Bayesian.

Choose the method

WhereWhat
Project → Settings → Statistics & Analytics → Default Statistical MethodsDefaults for new experiments: Statistical method for binary metrics and Statistical method for continuous metrics.
Experiment → Edit settings → Statistics → Statistical MethodsThe same two options for one experiment.
Results → Filters and Settings → Statistical methodA temporary view with another method. The method badges below the summary turn orange while an override is active.

The summary on the results page shows the methods in use, for example "Statistical method: Frequentist (binary) / Frequentist (numeric)".

Decision rules

A variation is declared winner or loser only if all three conditions are met. The results page repeats this rule below the summary.

SettingMeaningDefault
Confidence levelThe certainty a result must reach, from 50% to 99.9%. Available steps: 50, 70, 75, 80, 85, 90, 92, 95, 96, 97, 98, 99, 99.5, 99.8, 99.9.Set per project
Minimum conversionsNumber of conversions the variation needs before a decision.100
Minimum runtimeNumber of days the experiment must have run before a decision.2 days

If one condition is missing, the state is inconclusive, even when the certainty is high.

  • Winner: the result is confident and the change goes in the winning direction of the goal.
  • Loser: the result is confident and the change goes against the winning direction.
  • Inconclusive: not confident, or minimum runtime or minimum conversions not reached.

Minimum runtime counts from the day the goal was added to the experiment.

Change the settings

Set the defaults for the project under Project → Settings → Statistics & Analytics: Confidence level, Minimum conversions and Minimum runtime. For a single experiment, open Edit settings, go to the Statistics section and switch on Override project defaults. Then the three values of the experiment replace the project defaults.

A higher confidence level lowers the risk of a false result but needs more participants and a longer runtime. Change the settings before the experiment starts. Changing them after you have seen the results makes them less reliable.

Further options in Statistics & Analytics:

OptionDescription
SRM checkDisplays a warning in the results if the traffic split is off, see Check the traffic split (SRM).
Auto-pause on significancePauses a running experiment automatically, see Primary and secondary goals.

Next steps

On this page