AI & Technology

Closing the Feedback Loop

By Nir Weingarten, Co-founder and CEO of Eikona

How Reinforcement Learning Is Replacing A/B Testing as The Gold Standard for Performance Marketing 

No One Knows What Works 

Marketers and brand teams have the best understanding of their audience. They cultivate and study it, engaging constantly through different channels, and gauging each interaction. 

When it comes to performance marketing, even the best marketers and creative teams struggle to know in advance which content will resonate with their audience – what would make people click, engage and purchase. In fact, it’s very common that a piece of content that was almost discarded or considered ugly turns out to be a winner. 

To overcome this, experienced marketers fight to resist their biases and intuition and turn to data driven marketing, the gold standard of which is A/B testing – the practice of producing several pieces of content and statistically comparing their performance, giving marketers insight into what resonated more. 

But A/B testing has some serious drawbacks – primarily that it doesn’t scale. Across the board, whenever I ask marketing leaders if they A/B test their content I constantly hear a version of ‘far less than we would have wanted to.’ 

The recent improvements in generative AI have opened the door for a new and far more scalable method of learning from performance called reinforcement learning from human feedback, or RLHF. 

Tools of Mass Persuasion 

Reinforcement Learning (RL) is intuitive. Formalize an environment with states, actions and rewards, and unleash a learning algorithm to find the actions that maximize the compounded reward. This approach has gained much attention since late 2013 when researchers from Deep Mind demonstrated how it could be deployed to autonomously learn to play Atari games, by using the games’ points as reward.  

Usage of RL in gaming reached its peak on 9 March 2016, when an RL driven AI model called AlphaGo beat world champion Lee Sedol in the game of Go, winning 4 out of a series of 5 games, in what was considered impossible for a machine to accomplish before that event. 

What if we could use human engagement as a reward for our RL models? 

More specifically, the type and intensity of human engagement with a model’s output. Then, given enough data, models could be trained to produce content that is extremely persuasive. 

Reinforcement Learning from Human Feedback does exactly that and also gained world attention when it was used in late 2022 to train the first ChatGPT models, by teaching a previously anonymous LLM model (GPT3) to generate answers that humans engaged with better. In fact, this type of training is one of the main reasons models are so often eager to please. 

RLHF can allow marketers to connect content performance data back to the models that generated it, in an ever-evolving feedback loop. This type of optimization solves all of the fallbacks of A/B testing – that is, it’s always on, scalable, and offers multivariant testing. 

Retention Channels: The Ground Zero for RLHF in Marketing 

Instead of testing two versions and picking a winner, a system can generate copy and analyze how people respond to it and use that response to retrain itself. 

There’s no test cycle. There’s no single winner. Just a system that keeps getting better at holding your attention.  

Retention marketing, which is data-rich yet under-monetized, might well be the birthplace of this new approach. 

Retention marketers are charged with servicing the most valuable asset of their organization, the clients. They command the most powerful channels, such as email, SMS, in-app, push, and direct channels that are all replete in first-party data such as purchase history and browser behavior. 

Yet most users still get the same message, at the same hour, on the same channel, like they did in 2001. 

Customer retention is underserved, and under-valued 

A now well-known Harvard Business Review paper from 2014 found that increasing customer retention rates by 5% increases profits by 25% to 95%. These numbers still hold true today and are especially compelling when considering that Customer Acquisition Cost (CAC) continues to rise, while marketing budgets remain almost stationary, having risen fractionally to 7.8% of company revenue in 2026.  

Meanwhile CMOs now direct 15.3% of marketing budget into AI and organizations are actively seeking marketing business processes that can be automated with new technology. A/B testing is one area which has great potential to be elevated by AI and in turn bolster retention rates.  

A/B testing exists because no one can say in advance which banner will convert better. So teams run the test and let the data talk. But as every marketer knows, A/B testing is limited. 

Each variant needs a full creative cycle, and in most cases, teams are struggling to deliver on deadline.   

Running the test itself takes real hours to set up, run and analyze and in most cases the results never reach statistical significance in any way.  

Even when a test does produce a clear winner, the reason why is rarely obvious: was it the headline, the color, the verb in the call to action, or some interaction between all three?  

And whatever gets learned doesn’t compound.  

A team learns something in March, the next brief opens blank in April, and the learning evaporates. 

Paid media stopped working this way years ago. Google’s PMax and Meta’s Advantage+ don’t ask a marketer which creative to run; they run all of them and let outcomes decide. Netflix hasn’t hand-picked a piece of key art since 2017, and instead it treats artwork as a multi-arm bandit and lets engagement choose, per member. 

Owned media, the channel with the best data and the lowest media price, is the last place still doing it by hand. And most teams still consider A/B testing best-in-class. To retain customers and start reaching customers more precisely, this needs to change. 

Closing the feedback loop 

The vast amounts of engagement data already in the wild, especially in places where it is directly accessible to marketers, such as in owned media, make a strong case that applying RLHF to these channels could produce a massive commercial benefit. 

This is where A/B testing can give way to a new technique, one that closes the feedback loop of user engagement by using it as an RLHF objective for the same AI models that generate the next campaign.  

In practice, that means adding another learning layer to the marketing stack. An adaptive layer that collects engagement data and uses it to fine-tune the models that generated it using RLHF techniques. 

This new type of media is already being deployed in production. Illustrated in recent research, Meta trained an LLM they call AdLlama using historical ad performance as the reward signal, then tested it across roughly 35,000 advertisers and 640,000 ad variations over ten weeks, resulting in substantial and statistically significant lifts. 

There are strong signals that RLHF training for generative models can dramatically affect user engagement. Given the recent improvements in models and workflows, a transition to a new type of media, generated to optimize engagement to its mathematical optimum, looks to be imminent and transformative for customer connection and retention. 

Nir Weingarten is the Co-Founder and CEO of Eikona, a start-up using GenAI and reinforcement learning to transform lifecycle marketing. A published AI researcher with a master’s degree in machine learning from Reichman University, he has spent over a decade leading multi-disciplinary technical and product teams across the fields of AI, data and performance.

Related Articles

Back to top button