How to handle A/B tests when multiple metrics point in different directions
A/B testing becomes complex when different metrics from the same experiment produce conflicting results, such as higher signups but lower revenue. Testing multiple metrics simultaneously raises the risk of false positives, and standard statistical corrections like Bonferroni require significantly longer test durations that many sites cannot afford. The author argues that p-values answer the wrong question, and proposes a Bayesian approach using Beta-Binomial posteriors to calculate the expected loss of shipping a variant rather than relying solely on statistical significance. Two key outputs — probability that variant B is better and expected loss if it is not — together provide a more actionable decision framework. For skewed metrics like revenue per visitor, the author recommends bootstrapping instead of standard distribution models, which tend to underestimate uncertainty in zero-inflated, heavy-tailed data.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in