A/B testing

A B2C company emails subscribers that are both customers and prospects (@60/40 mix) with mainly “full blast” promotional offers.

When testing, customers and prospects are split, and both groups are tested/tracked different versions, only to find that these groups give different results.

When testing to reach significance, results show that with prospects, the “winner” was clear (surmising they don’t have enough experience with brand to be affected by change). However, for customers “earlier tests” in test period give different results than “later tests” (customers would seem to be “adverse to change” initially, but “get used to it”).

Assuming there are insufficient resources to execute separate version of emails for these groups, the sales tracking process is “complicated” (making “in-email" personalization challenging) – and ignoring the obvious flaws in the email strategy overall, how should this company be looking at these results?

For testing period, the audiences are “fixed” so each subscriber would get the same version during the entire testing process. While it makes sense, this approach makes every test read challenging, which it seems can often lead to “projecting hypothesis onto results”.

Thoughts?

I’m going to preface this by saying that this answer is provided without knowing anything about the customers and business, but with enough context from your post.

Within “customers” I think it would be safe to say that this audience isn’t uniform. For example, there are people that like and will wear Nike basketball shoes, but they’d also wear New Balance. There are also people that will only wear Nike and wouldn’t even consider another option. Both could technically be your customer.

Those loyalists are likely more engaged and I’d theorize they’d make up more of your early opens. They recognize your brand quickly, have a stronger expectation for what your emails normally look like, and they are more likely to notice changes right away. If you introduce a creative or layout change, it would not surprise me if you see an early dip with this group because you are breaking a pattern they are comfortable with.

Later in the test window, you may be seeing a different slice of customers engaging. More casual buyers, less habit-driven openers, and people who are not as anchored to your usual format. With that group, the same change might land differently, which can make the later read look inconsistent compared to the early read.

So my take is that you are not necessarily seeing customers “get used to it” over time. You might be seeing different customer sub-groups respond differently, and early vs late engagement is acting as a proxy for loyalty or engagement level.

That’s just my perspective, but this was a thought provoking post. Thanks for sharing!

Josh

Hey @PScott58 ,

Really interesting question. Josh’s point about early vs late openers acting as a proxy for engagement level is spot on and I think that framing alone changes how you should read those results.

One thing I’d add is that if you’re running fixed audience tests over a longer period, you’re essentially measuring two things at once: the creative difference and the behavioral shift over time within the same segment. Those two signals get tangled together. For customers specifically, you might get a cleaner read if you shorten the test window and increase the sample size per test instead of stretching the same audience across many sends. That way the “early loyal openers vs late casual openers” effect has less room to skew your results.

Also worth considering, even if you cant run fully separate email versions for customers vs prospects, you could at least evaluate your test results separately for each group and only declare a winner based on the prospect segment (since they give you cleaner signal). Then for customers, treat it more as a directional insight than a statistical conclusion. I’ve seen teams at Stacksync and at other companies I’ve worked with fall into the trap of forcing one winner across mixed audiences when the segments are clearly responding to different things. Sometimes the most honest answer is “this test was inconclusive for half our list” and thats okay. Disclosure: This answer comes from my own experience and was lightly rephrased with AI to improve readability.

Good discussion, thanks for posting it.