By Dilhan · · 5 min read

A/B Testing with Low Traffic: An Onboarding Experiment Plan

Estimate whether an onboarding A/B test can finish with your traffic, define a meaningful conversion lift, and choose useful work when it cannot.

Check whether the experiment can answer your question

A/B testing with low traffic starts with a feasibility calculation: how many eligible users reach the point you want to change, how large a conversion improvement would matter, and how long can you keep the test stable?

There is no universal traffic threshold. A large change to a frequently reached step may be measurable sooner than a small change to a rarely reached paywall. Count eligible new users or accounts, not sessions, page views, or total site visits.

Microsoft's guidance on experiment planning describes the importance of statistical power and trustworthy assignment. Use that planning discipline even when your product is small.

A worked example: 30% to 36% activation

Suppose 200 eligible new users reach onboarding each week. Your current activation rate is 30%, and you want to detect an increase to 36%.

That is a six-percentage-point increase, or a 20% relative improvement. Those quantities are different; entering “20” as a percentage-point lift would describe a much larger change.

For a two-variant test with equal allocation, a two-sided 5% significance level, and 80% power, a standard normal approximation gives about 963 users per variant. That is roughly 1,926 eligible users overall, or 9.6 weeks of recruitment at 200 per week.

text
Expected recruitment time = total required eligible users
                            ÷ eligible users per week

1,926 ÷ 200 = about 9.6 weeks

This is an illustrative planning estimate, not a guarantee that a test will reach significance. The calculation assumes independent assignments and a binary conversion outcome. It uses the pooled variance under the null and separate variances under the alternative, consistent with the approach documented in statsmodels' two-proportion power function.

If activation is measured through day seven, allow another week for the last recruited users to finish that window. At 50 eligible users per week, recruitment alone would take about 38.5 weeks. The same hypothesis may no longer be useful on that timeline.

Count traffic at the point of eligibility

A website with 1,000 weekly visitors does not necessarily have 1,000 users available for an onboarding test. If 100 sign up and 60 reach the changed step, the test may recruit only 60 eligible users a week.

Decide where random assignment happens. If you assign at signup, analyse every eligible signup in its original variant, including users who later leave. If you assign at the paywall, the result describes paywall entrants. Do not compare those two experiments as if they had the same denominator.

For team products, randomising whole workspaces can prevent teammates from receiving conflicting experiences. Sample size must reflect that assignment unit; treating correlated teammates as independent users can overstate precision.

Choose one outcome and write the decision before launch

Start with a claim you can test: “Showing a sample project before asking for an integration will increase the share of new accounts that create their first useful report within seven days.”

Define the primary outcome, eligibility rules, attribution to variants, minimum effect worth detecting, sample target, and observation window. Pick guardrails such as errors, abandonment, refunds, or support requests where appropriate.

An onboarding completion event is convenient, but a change could increase completion without increasing delivered value. Choose the event deliberately using the activation guide, and keep downstream paid outcomes visible when they mature.

Use two variants when traffic is scarce. Four variants divide the available users into smaller groups and introduce additional comparison decisions. Extra designs are useful only if the traffic and analysis plan support them.

When the planned test is too slow

You still have useful work available. Choose a method that answers a smaller question honestly.

QuestionUseful next step
Can a user complete this task?Watch representative users attempt it and record the obstacle
Is a step technically broken?Reproduce errors and inspect event delivery
Where do users leave?Inspect a consistently defined funnel with counts
Is the offer understood?Ask users to explain the price and outcome in their own words
Does the redesign cause a conversion lift?Run a planned experiment when recruitment is feasible

Qualitative research can explain confusion, but a handful of interviews cannot establish a conversion lift. Fixing a reproducible crash also does not require an experiment to establish that the crash exists.

Use the funnel calculator to choose the step to investigate. The drop-off instrumentation guide helps distinguish missing events from real abandonment.

Treat before-and-after results as observations

If you ship an improvement without a randomised test, compare releases using the same eligibility and outcome window. Record campaign changes, traffic sources, pricing, seasonality, and other releases that could affect the result.

An observed improvement is useful evidence for a next decision. It is weaker evidence for causality because users arriving before and after the release may differ. Say “conversion increased after the change” rather than claiming the change caused the entire increase.

Our Splitr case study is a concrete example of investigating a specific handoff and comparing subsequent results. Read the method as well as the percentage.

Avoid declaring a winner every morning

A conventional fixed-sample significance test assumes the agreed analysis plan. Repeatedly checking results and stopping on the first favourable reading can make a positive result misleading. Use an appropriate sequential method if you need continuous decision-making.

Also inspect whether the allocated groups match the planned split, whether tracking works in both variants, and whether all included users have completed their outcome window. A statistically impressive number from a broken assignment or missing event is not a useful product result.

OnRamp can help establish the funnel baseline and compare product outcomes. Experiment assignment and statistical analysis need their own implementation; a funnel comparison alone is not an A/B testing engine. Try OnRamp when your immediate need is to find and measure the onboarding step worth changing.