How to Run a Sales SMS Holdout Test While Reps Keep Dialing

Updated 19 min read How we research

TL;DR: A sales SMS holdout test randomly withholds the tested message from part of an eligible audience so you can measure what the SMS actually caused instead of what it got credit for. Randomize at the customer level, not the phone number and not the opportunity, because one buyer with two records lands in both groups and quietly flattens the result. Pick one primary outcome tied to money, such as purchases, qualified meetings, incremental revenue, or contribution profit, and set the attribution window from your buying cycle rather than from the day clicks stop arriving. Size the test on your baseline conversion rate and the smallest lift worth acting on. The NIST/SEMATECH handbook’s two-proportion test pools both observed rates into one estimate and leans on a normal approximation that holds only when the samples are reasonably large, with Fisher’s exact test as the small-sample fallback, and its sample-size derivation puts the difference you want to detect inside a squared denominator, which is why halving the effect you care about costs you far more than double the audience. Analyze by original assignment. ICH E9 calls that the intention-to-treat principle, meaning subjects are followed up and analyzed as members of the group they were allocated to regardless of whether they complied with the treatment, and that is exactly why carrier filtering and undelivered texts do not wreck the estimate. Two things break sales holdouts that never break marketing holdouts. Reps keep calling the holdout, so suppression has to cover the one-off text a rep fires from the call screen and not just the automated sequence. And opt-outs are law rather than a setting, because 47 CFR 64.1200 lets a called party revoke consent by any reasonable method, makes stop, quit, end, revoke, opt out, cancel, and unsubscribe reasonable per se in a reply text, requires you to honor other wording a reasonable person would read as revocation, forbids designating one exclusive channel for it, and gives you a reasonable time not to exceed ten business days to act. Its solicitation hours run 8 a.m. to 9 p.m. in the called party’s local time, not yours, so one nationwide send time is a different treatment in every time zone. Report both group sizes, both rates, absolute lift, relative lift, and an uncertainty interval, then clear sample-ratio mismatch, duplicates, contamination, and delayed conversions before anyone scales the campaign.

Your SMS campaign reported 240 purchases. Good. How many of those people were going to buy anyway?

That is the only question a holdout test answers, and it is the question attribution reporting cannot answer at all. Click-through and reply rate tell you the message got noticed. They do not tell you the message moved revenue, because the buyer who was already halfway to a decision also clicks. A holdout test separates the two by the only method that reliably works, which is withholding the message from a randomly chosen slice of the same eligible audience and watching what that slice does without it.

The design is simple. The execution is where sales teams lose it, because a sales floor is not a marketing list. Reps are still dialing. Deals are still moving. Somebody is going to text a holdout contact from the call screen on Thursday afternoon and nobody is going to log it.

What a sales SMS holdout test actually measures

A sales SMS holdout test is a randomized experiment. You take one eligible population, split it at random, send the defined message or sequence to the treatment group, withhold it from the holdout, and compare business outcomes over a window you set in advance. If the randomization is clean and the execution holds, the difference between the groups estimates what the SMS caused during that test.

Define the treatment precisely, and write it down before anyone builds the audience. Sender, message, offer, timing, cadence, audience, conversion window. A three-message sequence is not the same treatment as a single promotional text even when both go to the same list, and if you swap the offer on day four you no longer have one test, you have two underpowered ones.

When to use a holdout and when to run an A/B test

An A/B test compares two variants. Different copy, different call to action, different send time. So which one do you need? It depends entirely on the question you are actually asking. The A/B test tells you which version performed better. It does not tell you whether sending either version beat sending nothing, and that gap is where a lot of SMS budget hides.

So pick by the question. If the question is what did SMS cause, run a holdout. If the question is which message works better, run the variant test, and the same discipline that makes A/B testing your sales scripts useful applies here. You can also do both at once by splitting an eligible audience across two or three SMS variants plus a no-message control, as long as you size the test for the number of comparisons instead of pretending it is still a two-arm experiment.

Write the test hypothesis and pick one primary outcome

Write the plan before launch. A useful hypothesis names the audience, the treatment, the expected direction, the measurement window, and the decision rule you will follow when the number comes back.

Among eligible prospects, the defined three-message automated SMS follow-up sequence will increase qualified meetings booked within 14 days compared with sending no campaign SMS.

Then pick one primary metric. Which one? The one your CFO would recognize. Purchase conversion, qualified meetings, completed applications, incremental revenue, contribution profit. Clicks and replies are diagnostics. They tell you whether the message landed and whether the copy worked, which is worth knowing, but a lift in replies that does not show up in booked meetings is a finding about your reps, not about your SMS.

Decide now how you will treat repeat purchases, cancellations, returns, duplicate opportunities, and conversions that land after the window closes. Deciding later, with the results on the screen, is how a flat test becomes a winning one.

Decide who is eligible before the test randomizes anyone

Build the eligible population first, then randomize it. Same criteria on both sides, no exceptions. Typical exclusions are contacts you are not permitted to message, existing suppressions, employees, suspected fraud, geographies you cannot support, recent purchasers, and anyone already enrolled in a conflicting sequence.

A frosted glass hopper filled with small glass beads empties through a perforated glass plate on a pale violet background, with some beads diverted sideways into a small separate tray and the remainder falling into two equal side-by-side glass trays.

Document the rules. Then stop touching them. If eligibility genuinely has to change mid-campaign, keep an audit trail showing when each record was added, removed, or suppressed and why, because a silent mid-test audience change is indistinguishable from cheating when someone reviews the result in three months.

Consent, opt-out handling, quiet hours, recordkeeping, and message content requirements vary by jurisdiction, message type, carrier and provider policy, and your existing relationship with the recipient. Have qualified legal or compliance counsel review the campaign. An experiment is not an exemption from anything.

Size the test on your baseline rate, not a holdout percentage

So what is the right holdout percentage? There is no such thing. Ten percent is not a rule, it is a habit. The required sample size depends on your baseline outcome rate, the minimum effect worth detecting, the power you want, the significance threshold you set, and how you split the audience between treatment and holdout.

The NIST/SEMATECH handbook’s derivation for testing proportions puts the difference you want to detect in the denominator of a squared term. Read that as an operating constraint rather than as math. If you decide a 2-point lift is the smallest result you would act on, and then someone asks whether you could catch a 1-point lift instead, the honest answer is that it costs roughly four times the audience. Detecting a small change in a rare purchase takes far more contacts than detecting a large change in something that happens all the time. Get an analyst to run the calculation on your own numbers before you commit the list.

Set duration from the buying cycle and the attribution window, not from the day enough clicks have piled up. Leave room for delayed purchases, rep follow-up, returns, and cancellations. And do not extend the test because the interim numbers look good, or stop it early for the same reason, because peeking and stopping on a favorable reading inflates your false positive rate in a way no amount of downstream analysis repairs.

Randomize the test at the customer level and keep the assignment

Assign contacts using a reproducible random process, and assign at the customer or account level rather than by phone number, record, or opportunity. This is the failure that ruins more sales holdouts than anything statistical. What happens when one buyer has a mobile and a desk line? Or when two contacts at the same account sit twelve feet apart? They end up split across both groups. Your holdout has been treated. Your effect is diluted toward zero, and nothing in the output will tell you it happened.

When the audience has segments that genuinely behave differently, stratify. Randomize separately inside customer status, geography, account tier, or historical value, so those characteristics stay balanced without anyone hand-picking who gets the message.

Persist each assignment for the life of the test, in a field the campaign tool and the CRM both read. Before launch, compare group counts and baseline characteristics. A visible imbalance at that point is a broken assignment query, and it is far cheaper to find it now than in the readout.

Isolate the SMS treatment while reps keep calling

Send the defined campaign to the treatment group only, and suppress the holdout from equivalent SMS during the test. On a sales team, that suppression has to reach further than the campaign tool. Reps text from the call screen. They text after a voicemail. They text to confirm a meeting, and they text because the prospect asked them to. Where does that text show up in your suppression logic? Usually nowhere. If your reps send business texts from the same system they dial from, the holdout flag has to be visible there too, or the suppression only covers the automated sends and misses every manual one.

What about calls and email? It depends on the question. If you want the added value of SMS on top of an existing calling and email program, keep those other activities consistent across both groups and let the reps work normally. If the thing you are testing is a coordinated sequence, then the sequence is the treatment and the calls belong inside it. Those are two different experiments and they answer two different questions, so choose before launch instead of discovering afterward that you ran a hybrid.

Perfect isolation is not available. A rep will contact someone independently. A customer will forward an offer to a colleague. Paid media will reach one group harder than the other. Log accidental exposure rather than deleting the records, count it, and put the number in the readout. A test with 2% known contamination and an honest footnote is worth more than a clean-looking test nobody audited.

What the opt-out rules do to your holdout

This is the part marketing-side holdout guides skip, and it is the part that can turn a measurement problem into a compliance problem.

Under 47 CFR 64.1200, for the automated calls and texts the rule covers, a called party may revoke consent using any reasonable method that clearly expresses a desire not to receive further calls or texts. Replying to a text with stop, quit, end, revoke, opt out, cancel, or unsubscribe is reasonable per se. If someone replies with different wording, you still have to honor it when a reasonable person would understand the words as a revocation. You may not designate an exclusive means of revoking. And every revocation made by any reasonable means has to be honored within a reasonable time not to exceed ten business days from receipt.

What does that mean for a test you are about to launch? Three operating requirements, not legal trivia.

First, the opt-out channel is not just the SMS reply. A prospect who tells a rep on a call to stop texting has revoked consent, and if that never leaves the call notes, your suppression list is wrong and so is your test. Route verbal opt-outs into the same suppression the campaign tool reads.

Second, ten business days is a contamination window. Someone opts out on day three of a fourteen-day test. If your process takes eight days to propagate, that person sits in your treatment group receiving messages you already know they rejected. Shrink the propagation, then report opt-out counts and timing by group as a diagnostic, because a treatment arm with a rising opt-out rate is telling you something the conversion number is not.

Third, send time is part of the treatment. The same rule restricts telephone solicitation to residential subscribers to the hours between 8 a.m. and 9 p.m. in the local time at the called party’s location. Whether and how that reaches a given wireless number, message type, and relationship is a question for your counsel. The experimental point stands either way. If you fire one nationwide send at 8:30 a.m. Eastern, you delivered a breakfast message on the East Coast and a 5:30 a.m. message on the West Coast, which means you ran two different treatments and averaged them together. Schedule by recipient local time, or stratify by time zone and say so.

Run prelaunch QA before the first SMS goes out

Before the campaign activates, verify:

  • Eligibility, exclusions, audience counts, and assignment logic
  • Persistent customer-level treatment and holdout labels, visible in both the CRM and the texting tool
  • Suppression rules covering automated sends and manual rep texts
  • Links, timestamps, conversion events, and revenue fields
  • Offer terms, sender registration, opt-out handling, and internal approvals
  • Test start, end, attribution window, and decision rule
  • A named owner for monitoring delivery failures and accidental exposure

Carrier filtering deserves its own line on that list. Undelivered messages are not random, which is why 10DLC registration and delivery practices affect the test and not just the campaign. Copy matters too, and a message that reads like a blast gets filtered like one, so the same instincts behind SMS templates that get replies are doing measurement work as well as conversion work.

Freeze the primary analysis plan before you look at outcomes. Can you explore afterward? Of course. Just label it exploratory instead of presenting it as the hypothesis you started with.

Measure SMS lift by original assignment

Analyze everyone according to the group they were assigned to, including the people whose messages never arrived. ICH E9 states the principle plainly, that subjects allocated to a treatment group should be followed up, assessed, and analyzed as members of that group irrespective of their compliance with the planned course of treatment. Why not just drop the people whose messages never arrived? Because delivery failure is not random. It correlates with carrier, device, number age, and list quality, and randomization is the only thing making the two groups comparable in the first place.

Absolute conversion lift is the treatment conversion rate minus the holdout conversion rate.

Relative lift is absolute lift divided by the holdout conversion rate.

Incremental outcomes come from multiplying absolute lift by the treated population.

Report all of it. Both group sizes, both rates, the absolute and relative difference, and an uncertainty interval. A relative lift quoted without the base rate is the most misleading number in campaign reporting. Twenty percent sounds enormous. Then you see it means five conversions.

A delivered-message analysis can help diagnose execution problems. Run it, label it secondary, and never let it replace the assignment-based result.

Turn SMS lift into revenue and contribution profit

Compare average customer-level revenue between the two groups. Do not credit every purchase that followed a click. Incremental revenue per treated customer is the difference between the treatment average and the holdout average, and that subtraction is the whole point of having built a holdout in the first place.

For the economics, subtract what the campaign actually consumed. Message costs, discounts given, returns, cancellations, payment fees, fulfillment, and any variable cost the campaign triggered. Define profit before you report it. Is contribution profit the same as accounting profit? No, and the gap between them is where a positive test quietly becomes a negative one.

A worked sales SMS holdout test example

These figures are invented to show the arithmetic. They are not benchmarks and they are not Kixie results.

Two frosted glass bars of slightly different heights stand on a glass base on a pale violet background, with an I-shaped glass whisker marker on the taller bar reaching down past a thin glass rule set at the shorter bar's top height.

Take two hypothetical groups of 4,000 eligible customers each. Treatment records 240 purchases, a 6% conversion rate. Holdout records 200 purchases, a 5% rate.

  • Absolute lift is 6% minus 5%, or 1 percentage point
  • Relative lift is 1 divided by 5, or 20%
  • Estimated incremental purchases across 4,000 treated customers is 1% of 4,000, or 40

An approximate 95% interval around that 1-point difference runs from roughly 0 to 2 percentage points. Report the interval next to the estimate. The data are consistent with a 2-point lift and they are also consistent with nothing, which is a very different sentence from “SMS drove a 20% lift.”

Now the money. Say average revenue per customer is $84 in treatment and $72 in holdout. Incremental revenue is $12 per treated customer, or $48,000 across the group. If message costs ran $320 and incremental discounts plus variable fulfillment ran $18,000, the illustrative incremental contribution is about $29,680. Same caveat. Invented numbers, shown for the arithmetic.

Validity checks before anyone believes the test

Work this list before the decision, not after someone has already put the result in a board deck:

  • Sample-ratio mismatch, meaning the realized split differs from the intended split by more than chance explains. A chi-square goodness-of-fit test against the intended allocation is the standard check, and a failure here usually means the assignment or the delivery pipeline dropped records unevenly
  • Duplicate customers across groups
  • Delivery failures, and whether they cluster by carrier or segment
  • Cross-channel contamination, including manual rep texts
  • Tracking gaps and revenue fields that stopped populating
  • Seasonality, promotions, and anything else that hit one group harder
  • Early stopping, mid-test offer changes, and audience edits
  • Returns and delayed conversions that land after the window

One test, on one audience, in one season, with one offer, is one data point. It does not establish that the same lift shows up next quarter at three times the volume. Replication is what tells you whether the effect is stable enough to build a plan on, which is the thing to check before you scale into bulk SMS campaigns.

Decide what to do next, then write the test down

Use the decision rule you wrote before launch and the campaign economics you calculated. Scale, revise, rerun at a larger size, or stop.

What if the interval crosses zero? That is not proof of no effect. It usually means the test was underpowered, and the plausible range still holds both a profitable outcome and an unprofitable one. Saying “we could not tell” is a legitimate finding. Reporting it as “SMS does not work” is not.

Archive the hypothesis, the audience query, the assignment method, the message assets, the dates, the costs, the results with uncertainty, the execution problems, and the decision. The next test gets faster and more credible because this one was written down.

Sales SMS holdout test checklist

  • Define one primary business outcome and the attribution window
  • Document eligibility, exclusions, and required compliance review
  • Size the test from your baseline rate and the smallest lift worth acting on
  • Randomize at the customer level and persist the assignment where both systems can read it
  • Suppress the holdout from the tested SMS, including manual rep texts
  • Route verbal opt-outs into the same suppression list as replies
  • Schedule sends by recipient local time or stratify by time zone
  • QA tracking, links, offer details, group balance, and reporting fields
  • Analyze by original assignment and disclose contamination
  • Report both rates, absolute lift, relative lift, uncertainty, revenue, costs, and contribution
  • Check sample-ratio mismatch and duplicates before believing anything
  • Write down the decision and the limitations before starting the next test

Sales SMS holdout test FAQs

How large should the SMS holdout group be

Run a sample-size calculation rather than picking a percentage. The answer moves with your baseline conversion rate, the minimum effect you would act on, the power and significance you set, the treatment-to-holdout split, and how much eligible audience you actually have. A 10% holdout on a small list frequently cannot detect anything worth detecting.

Should holdout contacts still get calls and emails

Yes, if the question is what SMS adds on top of an otherwise consistent program, and that is usually the right question for a sales team. Keep calls and email running normally for both groups. If calls and email are part of the coordinated sequence you are testing, they follow the treatment definition instead, and the holdout gets none of it.

How do opt-outs and failed deliveries get handled

Honor every opt-out that applies, from any channel, within the time the rules require. For analysis, keep people in their originally assigned groups for the primary result, and report opt-out rates and delivery failures separately as diagnostics. An opt-out spike in the treatment arm is a finding about the message even when conversion looks fine.

Can clicks be the primary metric

Only when clicks are genuinely the business objective. If the objective is incremental sales, use purchases, qualified opportunities, revenue, or contribution profit. Clicks are cheap to move and easy to move in a direction that never reaches a rep.

How long should a sales SMS holdout test run

Long enough to cover the buying cycle plus the attribution window, set before launch and left alone. A test that ends when the clicks slow down measures the click curve, not the sales cycle. If returns or cancellations are material in your motion, the window has to outlast them.

Sources

How this article was built: the consent revocation, opt-out timing, and solicitation-hour requirements are quoted from the Code of Federal Regulations text of the FCC’s telephone solicitation rules, the intention-to-treat definition is quoted from the ICH E9 guideline, and the proportion comparison, sample-size behavior, and goodness-of-fit check come from the NIST/SEMATECH Engineering Statistics Handbook, each read directly on the review date.

  • 47 CFR 64.1200, Delivery restrictions, Federal Communications Commission, primary regulatory text via GovInfo, for paragraph (a)(10) providing that a called party may revoke prior express consent to receive calls or text messages by any reasonable method clearly expressing a desire not to receive further calls or text messages, that the words stop, quit, end, revoke, opt out, cancel, or unsubscribe sent in reply to an incoming text message constitute a reasonable means per se, that a reply using other words must be treated as a valid revocation request if a reasonable person would understand those words to convey a request to revoke consent, that all such requests must be honored within a reasonable time not to exceed ten business days from receipt, and that callers may not designate an exclusive means to request revocation; and for paragraph (c)(1) prohibiting any telephone solicitation to a residential telephone subscriber before the hour of 8 a.m. or after 9 p.m. local time at the called party’s location.
  • E9 Statistical Principles for Clinical Trials, International Council for Harmonisation, primary methodology guideline, for the glossary definition of the intention-to-treat principle as the principle asserting that the effect of a treatment policy can be best assessed by evaluating on the basis of the intention to treat a subject rather than the actual treatment given, with the consequence that subjects allocated to a treatment group should be followed up, assessed, and analysed as members of that group irrespective of their compliance to the planned course of treatment.
  • How can we determine whether two processes produce the same proportion of defectives?, NIST/SEMATECH e-Handbook of Statistical Methods, primary methodology reference, for the two-sample proportion z-test that pools both observed proportions into a single estimate, for the normal approximation to the binomial being the basis of that test when the samples are reasonably large, and for the Fisher exact probability test as the recommended technique when the two independent samples are small.
  • Sample sizes required, NIST/SEMATECH e-Handbook of Statistical Methods, primary methodology reference, for the derivation of required sample size when testing proportions in which the difference to be detected appears in the denominator of a squared term, so that smaller detectable differences require disproportionately larger samples, and for the dependence of the required sample size on the baseline proportion, the significance level, and the power.
  • Chi-Square Goodness-of-Fit Test, NIST/SEMATECH e-Handbook of Statistical Methods, primary methodology reference, for the chi-square goodness-of-fit test as the standard test of whether a sample of data came from a population with a specified distribution, which is the check applied to a realized treatment and holdout split against the intended allocation.

Sources verified and content reviewed by the Kixie Research Team on September 27, 2026. All source links checked on September 27, 2026.

Ready to close more deals with Kixie?

See how Kixie's AI-powered tools can transform your sales and support operations.

Start Free Trial