Holdout testing
We run holdout tests for CRO: we set aside a control group that is NOT shown a new feature/change/personalization, and measure the true causal effect of the change on conversion from the difference with those who see it. So you know a change really works rather than 'it seems better'. Honestly upfront: this is an experiment with a probabilistic result for the test conditions, not a growth guarantee; the control group is deliberately 'un-shown' the change, the price of measurement cleanliness.
Holdout testing — overview

Holdout testing for CRO is an experimental project: we design a test with a control (holdout) group at the user/segment level to estimate the incremental effect of a specific change, feature or personalization; we define metrics, compute the needed size and duration for statistical power, run it and analyze lift with a confidence interval and significance. Honestly about the nature, this is key: even a correct holdout gives a PROBABILISTIC result with an interval, not an exact number; the conclusion is valid for the test's audience and period and need not repeat in another context; a 'no effect' or 'worse' result is possible — that is valid and saves you resources. Honestly about the method's cost: a holdout means deliberately NOT giving the change to part of the audience during the test — a 'forgone' effect there for reliability. Honestly about requirements: a sufficient audience volume for significance and the ability to technically isolate a control are needed; on low traffic the test will not work. Honestly about discipline: size and duration are set in advance, no 'peeking'. Honestly about data: correct conversions and a clean split (garbage in, garbage out). Honestly about the essence: this is effect measurement, NOT implementation and NOT a growth guarantee — decisions are yours. An important boundary: this is holdout in the CRO context (effect of a site change), adjacent to holdout/incrementality in analytics (for ads/CRM — separate services 698/696); not A/B (two versions at once — separate) and not building the feature itself. If the audience is small or a control cannot be isolated, the test is premature. Picture this: instead of 'we launched personalization and argue whether it helped' you see the clean difference vs control. The base price starts from 30,000 ₽ per test; it depends on design and scale.
Problems we solve
- You launched a change/feature and do not know if it had an effect.
- 'It seems better' — but without a control that is not proof.
- You assess the effect of personalization and changes by gut feel.
- There is no control group for an honest comparison.
What's included in the Holdout testing service
- Holdout test design (users/segment)
- Defining metrics and computing size/duration for power
- Clean isolation of the control group
- Running and controlling test cleanliness
- Analyzing incremental effect (lift) with an interval and significance
- An honest conclusion (including 'no effect'/'worse')
- A report with limitations and validity conditions
- Reviewing results with you
What you get
- A causal estimate of the change's effect on conversion (with an interval)
- You understand whether a change actually works
- Savings on non-working changes
- A base for decisions (implementation and growth — separately)
How the work goes: steps
- We clarify the change, audience, metrics, scale; plan the design
- We isolate the control, run it, do not 'peek'
- We analyze lift and significance, compile a report, review with you
Why PDV Expert
- Fixed price and timeline — no surprises on the invoice.
- Report and recommendations in plain language — clear without a technical background.
- In touch at every step and answering questions about the result.
FAQ
Does a holdout guarantee a conversion increase?
No. It is a way to measure a change's effect, not a guarantee. A 'no effect' or 'worse' result is possible — valid and resource-saving. The conclusion is probabilistic, with an interval, and valid for the test conditions.
Is the control group a loss?
Partly: part of the audience is deliberately not given the change during the test — a 'forgone' effect there for measurement cleanliness. In return you stop investing in what does not work. A sufficient volume is needed for significance.
How is it different from A/B?
A/B compares two versions at once; a holdout compares 'with the change' against 'without the change' (control), which is convenient for assessing the effect of features and personalization. The methods often complement each other; the choice depends on the task.
About the provider
The «Holdout testing» service is provided by PDV Expert — a team specialising in «Conversion & analytics». We work under contract and deliver a written report with recommendations.