Aglobal beauty brand runs hundreds of campaigns per year across 28 markets and 30-plus customer segments. Every campaign produces content across every channel — email subject lines, paid social captions, push notifications, product page headlines. At that scale, two decisions repeat themselves constantly: what to say to each audience, and which version of it to run.
Two problems were compounding each other. The brand had no reliable way to generate content genuinely calibrated to each market and segment — local teams were absorbing the gap through manual rework. And it had no reliable way to predict which variant would perform — so selection defaulted to intuition, A/B testing, or generic benchmarks. Both were expensive habits hiding in plain sight.
The generation problem.
Producing content at this scale is not a volume problem — it is a quality problem. A subject line that performs well with VIP customers in the UK does not perform the same way with new customers in South Korea or loyalty segment customers in Brazil.
Different markets have different cultural registers, different promotional sensitivities, different relationships with the brand. Generic AI writing tools produce content that sounds polished but is calibrated on general consumer behavior — not on what has actually resonated with this brand’s specific audience in each market. The output requires significant manual reworking by local teams before it is ready to run, which defeats the purpose.
The selection problem.
When it comes to choosing which content variant to run, most marketing teams — including this one — rely on one of three approaches. All three are inadequate.
- Creative intuition.
A marketing manager picks the version that feels right. Fast, but entirely disconnected from what has actually worked with this brand's audience in this market at this time of year.
- A/B testing.
Run two versions, wait for results, continue with the winner. Methodologically sound — but slow, expensive, and backward-looking. Every A/B test burns budget on the losing variant running on real customers. And by the time results arrive, the campaign window is often closing.
- Generic benchmarks.
Industry data on subject line length, emoji usage, CTA phrasing. Useful directionally, but calibrated on average behavior across many brands — not on this brand's specific customers in its specific markets.
The uncomfortable truth behind all three approaches:
None of them use what the brand actually has — five years of its own campaign performance data, across 28 markets, 30-plus segments, every channel, thousands of campaigns.
That data contains an enormous amount of signal about what works specifically for this brand’s customers — signal that no competitor has access to, no platform analytics tool surfaces, and no external vendor can replicate.
What about existing AI content tools.
Tools like Persado and similar AI language optimization platforms predict content variant performance using emotional and linguistic variables drawn from large consumer panels. The firm had already evaluated this category. The fundamental limitation is the same across all of them: their predictions are calibrated on general consumer behavior, not on any specific brand’s customers.
They tell you what works on average for consumers in your category. They cannot tell you what works for this brand’s VIP segment in the UK market in the spring, because they do not have that data. And they cannot generate content genuinely calibrated to each of this brand’s 28 markets, because they have not been trained on this brand’s own voice, history, and performance signals.
This brand has both. It just was not using either.
Coral Tree built a two-component system: one that generates market and segment-optimized content variants, and one that predicts which of those variants will actually perform — both trained on the brand’s own five-year campaign history.
Component 1 — Generation.
The generation layer uses a combination of fine-tuning and retrieval-augmented generation (RAG) to produce content that is genuinely calibrated to each market and segment — not generic copy with localized placeholders dropped in.
The model is fine-tuned on the brand’s own historical content archive: every high-performing email, social caption, push notification, and product headline the brand has ever produced, annotated with the market, segment, channel, and campaign context it was written for. This teaches the model the brand’s voice, its tonal range across different markets, and the structural patterns that have historically resonated with different customer segments.
At generation time, RAG retrieves the most relevant past content from the archive — the highest-performing examples most similar to the current brief — and surfaces them as context for the generation. The model is not starting from scratch. It is generating informed by the most relevant examples the brand has already proven work.
Given a brief like one of these:
“Subject lines for our Spring Collection email to the UK VIP segment, focused on our new moisturizer line, launching the first week of April”
“TikTok captions for our summer glow campaign targeting Gen Z in the US and Australia, product-led, 15-second format”
“Push notification for our Korean loyalty segment, flash sale on skincare bundles, valid for 24 hours only”
The system produces 15 to 20 genuinely diverse variants — calibrated to that specific market, that specific segment, and that specific channel — without requiring local teams to manually rework generic output.
Component 2 — Performance Prediction.
This is where the real differentiation lives. The generated variants are each scored by a prediction model trained on the brand’s five-year campaign archive — every subject line ever sent, every paid social creative ever run, every A/B test outcome ever recorded, with the measured performance of each one attached.
The model learns the relationships between copy style, campaign context, market, segment, season, promotional mechanics, and measured outcomes. It produces a predicted engagement metric for each variant — predicted click rate, predicted open rate — with a confidence interval attached.
The marketing manager sees the ranked list. She picks from the top three, makes any final editorial adjustments, and the campaign goes out. The prediction is a recommendation, not a mandate.
The Feedback Loop.
When the campaign runs, actual performance data — real click rates, real open rates, real conversion rates — is captured and appended to the training dataset. Every quarter, both the generation model and the prediction model retrain on the expanded dataset, incorporating the most recent market signal.
This is the closed feedback loop: generation informs selection, selection produces real performance data, real performance data improves future generation and predictions, better predictions improve future selection. Each cycle makes the next one sharper.
The system cannot be replicated by any external vendor. It is trained entirely on this brand’s proprietary content and performance history. It gets more accurate the longer it runs. And unlike any platform-native optimization tool or off-the-shelf content AI, it is not averaging across an external consumer panel — it is learning from this brand’s own customers, in their own markets, responding to this brand’s own content.
The numbers tell part of the story.
The CTR improvement compounds at the scale this brand operates. Across hundreds of campaigns, 28 markets, and 30-plus segments, an 18% lift in click-through rate is not a marginal gain — it is a material shift in how efficiently the brand converts its media spend into customer engagement.
The 27% reduction in A/B test waste tells a different part of the story. When the prediction model front-loads the selection decision, losing variants often never run at all. Budget that previously went to running the weaker option on real customers is now preserved. Across a brand operating at this campaign volume, that recovery runs into hundreds of thousands of dollars annually.
And the model that delivers both of those outcomes gets better every quarter — quietly, continuously, without any additional input from the marketing team.