Statistical methods
How ABlyft calculates results, which methods you can choose and which rules decide a winner.
ABlyft compares every variation with the baseline and calculates how sure it can be that the difference is real and not random. You choose between three methods: Frequentist, Bayesian and Legacy. On top of the method, three settings decide when a variation may be called a winner or loser: the confidence level, the minimum conversions and the minimum runtime.
The baseline
The baseline is the variation that all others are compared with. Normally this is Original, the unchanged version. It is marked with "Baseline" on the results page. Improvement, certainty and statistical state are only calculated for the other variations.
Two kinds of metrics
ABlyft uses a method for each kind of goal:
| Kind | Goals | Compared numbers |
|---|---|---|
| Binary metrics | Goals that count conversions, such as pageview, click and custom goals with measurement type Conversion. | Conversion rates |
| Continuous metrics | Goals that measure a value, such as revenue goals and custom goals with measurement type Value. | Value per participant |
For both kinds you select a method separately, so a revenue goal can be evaluated differently than a click goal.
The three methods
Frequentist
The default for new projects. It asks: if there were no real difference, how likely would we see a difference this large?
- Conversion goals: a two-sample z-test on the conversion rates.
- Value goals: Welch's t-test on the value per participant. Participants without a value count as zero.
- Goals with counting method Every, where the rate can exceed 100%: a Poisson rate z-test.
- Certainty is calculated from the p-value (two-sided). The result also shows the p-value and the z- or t-score.
Bayesian
It asks: how likely is it that this variation is better than the baseline?
- Conversion goals: a Beta-Binomial model.
- Value goals: a model for means with a weak prior.
- Both use 10,000 simulations, so the numbers can differ slightly after a recalculation.
- Certainty is the probability that the variation is better than the baseline. For goals with winning direction Decrease, it is the probability that the variation is lower. The table shows Credible Interval instead of Confidence Interval.
Bayesian with counting method Every
If the conversion rate of a goal is above 100% (possible with the counting method Every), the Bayesian method cannot calculate a result for conversion goals. The certainty stays at 50% and the variation stays inconclusive. Choose Frequentist for such goals, or use the counting method One.
Legacy
The method of earlier ABlyft versions, kept so that existing experiments stay comparable. It uses a classic z-test with the standard errors of both groups. New experiments should use Frequentist or Bayesian.
Choose the method
| Where | What |
|---|---|
| Project → Settings → Statistics & Analytics → Default Statistical Methods | Defaults for new experiments: Statistical method for binary metrics and Statistical method for continuous metrics. |
| Experiment → Edit settings → Statistics → Statistical Methods | The same two options for one experiment. |
| Results → Filters and Settings → Statistical method | A temporary view with another method. The method badges below the summary turn orange while an override is active. |
The summary on the results page shows the methods in use, for example "Statistical method: Frequentist (binary) / Frequentist (numeric)".
Decision rules
A variation is declared winner or loser only if all three conditions are met. The results page repeats this rule below the summary.
| Setting | Meaning | Default |
|---|---|---|
| Confidence level | The certainty a result must reach, from 50% to 99.9%. Available steps: 50, 70, 75, 80, 85, 90, 92, 95, 96, 97, 98, 99, 99.5, 99.8, 99.9. | Set per project |
| Minimum conversions | Number of conversions the variation needs before a decision. | 100 |
| Minimum runtime | Number of days the experiment must have run before a decision. | 2 days |
If one condition is missing, the state is inconclusive, even when the certainty is high.
- Winner: the result is confident and the change goes in the winning direction of the goal.
- Loser: the result is confident and the change goes against the winning direction.
- Inconclusive: not confident, or minimum runtime or minimum conversions not reached.
Minimum runtime counts from the day the goal was added to the experiment.
Change the settings
Set the defaults for the project under Project → Settings → Statistics & Analytics: Confidence level, Minimum conversions and Minimum runtime. For a single experiment, open Edit settings, go to the Statistics section and switch on Override project defaults. Then the three values of the experiment replace the project defaults.
A higher confidence level lowers the risk of a false result but needs more participants and a longer runtime. Change the settings before the experiment starts. Changing them after you have seen the results makes them less reliable.
Further options in Statistics & Analytics:
| Option | Description |
|---|---|
| SRM check | Displays a warning in the results if the traffic split is off, see Check the traffic split (SRM). |
| Auto-pause on significance | Pauses a running experiment automatically, see Primary and secondary goals. |
Next steps
- Interpret results
- Metrics
- Pre-test calculator to plan runtime