On this page

Sales after a rebate launch include purchases that would have happened anyway. A customer holdout can help estimate the additional purchases caused by offering the reward, provided the groups are created before exposure and remain interpretable through the observation period.

The experiment tests the offer’s effect on a defined population. It does not test whether customers who claimed a reward are better customers than those who did not. Claiming happens after the offer and can reflect many differences the experiment did not randomly assign.

Randomize before the offer is shown

Define the eligible experimental population before the offer is shown. Then randomly assign a stable customer-level unit to the proposed offer or the comparison experience. The treatment must be sufficiently specific to reproduce: reward amount, displayed explanation, eligible products, and other relevant conditions.

In a hypothetical pilot, 10,000 eligible customer identities are assigned equally before exposure. Treatment customers can see the approved purchase reward offer. Control customers see the otherwise comparable shopping experience without that new offer. Existing rewards from earlier purchases remain intact in both groups.

The merchant must determine that the proposed experiment and customer presentation are appropriate for its actual commercial and program conditions. If a public offer has already been promised to everyone in the population, do not revoke it from a newly created control group. Choose a future, properly scoped experiment instead.

Microsoft’s experimentation platform overview describes controlled experimentation as a measurement discipline. The protocol below is an original merchant example, not a claim that RebateCardX supplies an automatic randomization service.

Protocol item Hypothetical decision
Assignment unit Eligible customer identity defined before exposure
Allocation 5,000 treatment and 5,000 control assignments
Treatment The approved purchase reward offer and its associated explanation
Comparison Same planned shopping experience without the new offer
Primary outcome At least one retained purchase within 28 days of assignment
Existing entitlements Preserved in both groups
Analysis population Every assigned eligible customer, under the predefined exclusions
Decision owner Merchant experiment owner, with program and analytics review

Preserve assignment through purchase and qualification

Keep assignment stable across ordinary returns to the store. If customers can repeatedly enter until they receive the preferred offer, the comparison becomes partly self-selected. Document expected limits of identity persistence and exposure rather than assuming every shopper uses one browser and one device.

Preserve the original assignment in the analysis even when a treatment customer does not notice the offer. The business question is the effect of offering this experience to the assigned population, including the fact that some people will not engage. An analysis restricted to people who clicked the reward panel answers a selected, later-stage question.

Track qualification separately from assignment. A treatment customer who buys an ineligible product is still a treatment-assigned customer for the main outcome, if the outcome definition includes that purchase. Do not retrospectively turn product choice into an experimental eligibility rule unless it was determined before exposure.

A difficult case is a control customer who receives the offer from a friend. The merchant should honor the applicable published promise and approved program rules rather than deny a legitimate benefit merely to preserve a clean spreadsheet. Record mixed exposure and assess its effect on interpretation. The contamination analysis is a separate task from deciding the customer’s entitlement.

Before results are read, check assignment integrity, outcome coverage, and consistent follow-up. A random split is useful only if the data pipeline and analysis preserve it. Missing records that disproportionately affect one group can undermine the comparison even when the initial allocation was correct.

Estimate the effect using the assigned population

Suppose the fictional pilot ends with 600 retained purchasers among 5,000 treatment assignments and 500 among 5,000 controls. Treatment purchase conversion is 12%; control conversion is 10%. The observed difference is 2 percentage points, equivalent to a 20% relative increase from the control rate.

At the treatment group’s size, the point estimate corresponds to 100 additional purchasers: 5,000 multiplied by 0.02. This is an estimate from the experimental comparison, not a claim that the merchant can identify exactly which 100 individuals were incremental.

The readout also needs uncertainty. Under a simple independent two-proportion approximation, the standard error is about 0.00625, giving an approximate 95% interval for the difference of 0.77 to 3.23 percentage points. That calculation assumes the protocol’s unit and independence conditions are appropriate; repeated assignments or cluster effects require a suitable analysis.

Do not turn this purchase result into an immediate profit conclusion. Reward costs, retained revenue, support burden, and other predefined guardrails need their own review. A positive purchase effect may still fail the merchant’s commercial decision criteria.

Keep the primary result alongside protocol deviations. If a product went out of stock in one experience or a parallel promotion reached only one group, explain the changed interpretation. The experiment estimates the experience actually delivered, which may differ from the experience the team intended to test.

Finish with a decision record stating the population, offer, outcome window, point estimate, uncertainty, and remaining limitations. A well-run holdout produces a defensible estimate of offering the reward under those conditions. It does not prove that every future campaign, audience, or reward amount will produce the same result.

The commercial threshold also belongs in the protocol. Suppose the merchant had agreed that a gain smaller than one percentage point would not justify the operational work. The fictional interval extends below and above that threshold, even though it lies above zero under the simple approximation. Statistical evidence of a positive effect therefore does not by itself settle the scale decision. State whether the evidence resolves the merchant’s useful-effect threshold, and identify which additional cost or guardrail result will change the decision.

Discuss the reporting and operating requirements for a measured rebate pilot with Rebate Card X. Discuss program fit.

Source references

Back to contents

General information only

This guide is general information, not financial, legal, tax or regulatory advice. Eligibility, card availability, permitted use and responsibilities depend on the applicable offer and card terms.