# Statistical methods

URL: https://docs.ablyft.com/guides/results/statistics-methods/

> How ABlyft calculates results, which methods you can choose and which rules decide a winner.



ABlyft compares every variation with the **baseline** and calculates how sure it can be that the difference is real and
not random. You choose between three methods: **Frequentist**, **Bayesian** and **Legacy**. On top of the method, three
settings decide when a variation may be called a winner or loser: the confidence level, the minimum conversions and the
minimum runtime.

## The baseline

The baseline is the variation that all others are compared with. Normally this is **Original**, the unchanged version.
It is marked with "Baseline" on the results page. Improvement, certainty and statistical state are only calculated for the other variations.

## Two kinds of metrics

ABlyft uses a method for each kind of goal:

| Kind                   | Goals                                                                                                        | Compared numbers      |
| ---------------------- | ------------------------------------------------------------------------------------------------------------ | --------------------- |
| **Binary metrics**     | Goals that count conversions, such as pageview, click and custom goals with measurement type **Conversion**. | Conversion rates      |
| **Continuous metrics** | Goals that measure a value, such as revenue goals and custom goals with measurement type **Value**.          | Value per participant |

For both kinds you select a method separately, so a revenue goal can be evaluated differently than a click goal.

## The three methods

### Frequentist

The default for new projects. It asks: *if there were no real difference, how likely would we see a difference this large?*

* Conversion goals: a two-sample z-test on the conversion rates.
* Value goals: Welch's t-test on the value per participant. Participants without a value count as zero.
* Goals with counting method **Every**, where the rate can exceed 100%: a Poisson rate z-test.
* **Certainty** is calculated from the p-value (two-sided). The result also shows the p-value and the z- or t-score.

### Bayesian

It asks: *how likely is it that this variation is better than the baseline?*

* Conversion goals: a Beta-Binomial model.
* Value goals: a model for means with a weak prior.
* Both use 10,000 simulations, so the numbers can differ slightly after a recalculation.
* **Certainty** is the probability that the variation is better than the baseline. For goals with winning direction
  **Decrease**, it is the probability that the variation is lower. The table shows **Credible Interval** instead of **Confidence Interval**.

> **Bayesian with counting method Every:** If the conversion rate of a goal is above 100% (possible with the counting method **Every**), the Bayesian method
> cannot calculate a result for conversion goals. The certainty stays at 50% and the variation stays inconclusive.
> Choose Frequentist for such goals, or use the counting method **One**.

### Legacy

The method of earlier ABlyft versions, kept so that existing experiments stay comparable. It uses a classic z-test with
the standard errors of both groups. New experiments should use Frequentist or Bayesian.

## Choose the method

| Where                                                                         | What                                                                                                                       |
| ----------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| **Project → Settings → Statistics & Analytics → Default Statistical Methods** | Defaults for new experiments: **Statistical method for binary metrics** and **Statistical method for continuous metrics**. |
| **Experiment → Edit settings → Statistics → Statistical Methods**             | The same two options for one experiment.                                                                                   |
| **Results → Filters and Settings → Statistical method**                       | A temporary view with another method. The method badges below the summary turn orange while an override is active.         |

The summary on the results page shows the methods in use, for example "Statistical method: Frequentist (binary) / Frequentist (numeric)".

## Decision rules

A variation is declared **winner** or **loser** only if all three conditions are met. The results page repeats this rule below the summary.

| Setting                 | Meaning                                                                                                                                  | Default         |
| ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | --------------- |
| **Confidence level**    | The certainty a result must reach, from 50% to 99.9%. Available steps: 50, 70, 75, 80, 85, 90, 92, 95, 96, 97, 98, 99, 99.5, 99.8, 99.9. | Set per project |
| **Minimum conversions** | Number of conversions the variation needs before a decision.                                                                             | 100             |
| **Minimum runtime**     | Number of days the experiment must have run before a decision.                                                                           | 2 days          |

If one condition is missing, the state is **inconclusive**, even when the certainty is high.

* **Winner**: the result is confident and the change goes in the winning direction of the goal.
* **Loser**: the result is confident and the change goes against the winning direction.
* **Inconclusive**: not confident, or minimum runtime or minimum conversions not reached.

Minimum runtime counts from the day the goal was added to the experiment.

### Change the settings

Set the defaults for the project under **Project → Settings → Statistics & Analytics**: **Confidence level**,
**Minimum conversions** and **Minimum runtime**. For a single experiment, open **Edit settings**, go to the **Statistics**
section and switch on **Override project defaults**. Then the three values of the experiment replace the project defaults.

> A higher confidence level lowers the risk of a false result but needs more participants and a longer runtime.
> Change the settings before the experiment starts. Changing them after you have seen the results makes them less reliable.

Further options in **Statistics & Analytics**:

| Option                         | Description                                                                                                                                                         |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **SRM check**                  | Displays a warning in the results if the traffic split is off, see [Check the traffic split (SRM)](https://docs.ablyft.com/guides/results/interpret-results/#check-the-traffic-split-srm). |
| **Auto-pause on significance** | Pauses a running experiment automatically, see [Primary and secondary goals](https://docs.ablyft.com/guides/goals-and-measurement/configure-goals/#primary-and-secondary-goals).           |

## Next steps

* [Interpret results](https://docs.ablyft.com/guides/results/interpret-results/)
* [Metrics](https://docs.ablyft.com/guides/results/metrics/)
* [Pre-test calculator](https://docs.ablyft.com/guides/results/pre-test-calculator/) to plan runtime

