
Getting a large language model to write marketing copy takes very little effort, but turning that copy into an email a business can safely send, one that renders across email clients, meets deliverability requirements and serves a real campaign goal, is a harder engineering problem. It is the one Alina Olefirenko has focused on as Co-Founder and CTO of AlpacaRelay, where she designed the architecture behind the company’s AI email system. The system generates emails as structured MJML components and evaluates them through a combination of deterministic checks, LLM-based judgment and visual rendering. It is now being extended to recommend campaigns from a business’s own customer data.
Olefirenko has spent more than a decade in software engineering and architecture. She joined pdfFiller when airSlate was still an idea, helped develop and launch airSlate, and later served as a Software Architect as the platform scaled, designing a schema-driven framework for its automation integrations. She is also a Senior Software Engineer at Prove.
We spoke with Olefirenko about her work addresses a challenge many engineering teams building with LLMs now face: getting model output to meet strict technical requirements in production. She discussed how she structured generation, why dark mode proved to be the hardest rendering case, how the system scores email quality, and what she is still testing as the company moves from generating emails to measuring business outcomes.
Walk us through how you arrived at treating AI email generation as a structured-output problem, and what that meant for the architecture you designed.
I came into this with many years in frontend engineering, including time as a frontend architect, so I already knew that email HTML is its own world. Clients support HTML and CSS differently, layouts still depend heavily on tables, and something that looks fine in one client can behave very differently in another.
When I started working on AI email generation, I tried the obvious approach first: describe the email in a prompt and let the model generate it. More detailed prompts and more instructions improved some things, but I still didn’t have the reliability or control I wanted. A model can produce HTML that looks reasonable without producing an email I would actually trust in production. That was the point where I started treating it as a structured-output problem.
MJML was an important part of that shift because it already gives you an email-specific abstraction over a lot of the rendering complexity. I designed the generation architecture around components, such as a hero, content sections, images, CTAs, columns and a footer, instead of one large block of HTML. Each piece can be generated under its own constraints and still be edited in our visual builder afterward. It also gave me a cleaner way to describe the domain to the model. For each component, I could define what it is for, what it is allowed to contain and how it relates to the rest of the email. That became the basis of the architecture we use now.
Can you take us inside the scoring layer you built, including the quality factors it evaluates and how you chose them?
I originally built scoring for a very practical reason: I needed a way to tell whether the emails we were generating were actually good enough.
At first I evaluated them against many of the same rules I was using during generation, which turned out to be useful in a way I hadn’t expected. If an email followed the generation rules but still looked wrong or scored badly, that meant the rules themselves were incomplete, so the scorer became part of the feedback loop for improving generation.
Reference material pushed it further. We studied publicly available templates and real marketing emails, although I didn’t assume an existing email was automatically a good example. Comparing more of them kept surfacing things the scorer didn’t cover yet, including visual balance, variation, styling, content structure and consistency. There is also the separate question of whether it’s a good marketing email at all, since something can be visually polished and still weak on persuasion. That’s where I brought in established concepts such as AIDA, PAS and Cialdini’s persuasion principles.
Under the hood, it’s a mix of different kinds of evaluation. Deterministic checks cover MJML validity, nesting, column rules, attributes, section order, duplicate content, contrast, expected content blocks and structural rules for each type of email. Semantic questions, such as whether the email matches its intended type or fits the business context and purpose of the campaign, go to an LLM. A visual layer renders the email and looks for issues that are hard to infer from markup alone, like empty space, awkward spacing or visual imbalance, and checks brand constraints where we have them. We roll all of this up into an internal Email Quality Score for a high-level view, but the evaluation underneath is much more granular. The main lesson for me was that no single metric tells you an email is good.
Where did cross-client rendering push back hardest on your architecture, and how did you work around it?
Dark mode was probably the hardest case, because it exposed gaps in the data model, the editor and the generation pipeline at the same time.
MJML solves a lot of email layout problems, but it doesn’t really have a component-level concept of light and dark themes, and Apple Mail, Gmail and Outlook each handle automatic color changes differently. I wanted dark-mode intent to remain part of the email model instead of becoming a CSS patch added at the very end. So I added paired dark-mode attributes to the blocks that need them. We keep those values in the MJML representation and turn them into a single dark-mode style block before compilation. The elements get deterministic classes so we can target them reliably, and we add the color-scheme metadata and CSS precedence needed to deal with client-side color inversion.
The editor had its own issue. It runs inside an iframe and can inherit the operating system’s theme, so the production media query made the preview unpredictable. We separated the editor preview from production behavior and added a manual dark-mode toggle.
Performance was another problem. Our first AI-based dark-mode annotation step added roughly 30 to 60 seconds to generation, and I didn’t want that on the critical path, so I moved it into a separate asynchronous procedure with concurrency protection.
We also have to account for a compiler difference. The editor uses a browser MJML compiler with softer validation, while server-side compilation and scoring are stricter, so we normalize and validate around that boundary instead of trusting a single render. Screenshot testing only gets us so far, too. Our rendering service uses Chromium, which is useful as a visual proxy but isn’t Outlook, so we still do real test sends. I make sure a test send uses the same HTML we plan to send in production, because a sanitizer can remove media queries or Outlook-specific markup, and then the test no longer represents the real email.
Tell us about the jump from writing emails to recommending which campaigns a business should run in the first place.
I started thinking beyond individual emails pretty early. Businesses rarely communicate in isolated messages, and even a simple use case often becomes a sequence. That affected the generation architecture, because context such as the business, brand and communication style should be established once and then reused across the individual emails.
Sending came early too. Initially I needed it for engineering reasons, since I had to see what a generated email actually looked like in Gmail, Outlook and other environments. As the product expanded, the next problem became deciding which campaign would be useful for a particular business at a particular moment, and that requires context about the business and its customers.
We explored a few ways to get that context, including our own tracking, existing analytics events, CRM exports and integrations with systems the business already uses. Building our own tracking layer is possible, but it adds setup, and many of the businesses we want to serve aren’t technical. I would rather remove setup than ask the customer to become an integration engineer. That’s why integrations with operational platforms became important. Square is one of the first, and that flow already works end to end. The business already has customer, appointment, service and transaction data there.
A simple example is a salon service with a recurring booking pattern. If a customer’s history shows they tend to come back for that service around every six months, and five and a half months have passed, the system can recognize that timing and recommend a reminder campaign. The owner gets an audience, a reason why now is a good time and a campaign ready to send.
Once a campaign has gone out, what happens to its results inside the system?
The goal is a closed loop, and I’m careful to separate what already works from what we’re still building.
What works now starts with what we internally call an “X-ray.” It scans the available business and customer data, builds a baseline view of the business and looks for problems or opportunities. Underneath that is a graph-based registry, which gives the analysis layer a normalized representation even when the raw data comes from different integrations. Every provider has a different schema, APIs change, and one business may eventually connect more than one system with overlapping customer data, so I didn’t want campaign logic tied directly to Square’s schema or anyone else’s. The registry represents customers, services, events, campaigns and the relationships between them, while preserving the attributes that matter to a particular business. After a campaign is sent, new customer and business events enter that same representation, so we have the data to compare what happens afterward with the state we saw before.
What we’re still building is the automatic connection between those later events and the reason the campaign was sent. If a campaign is meant to bring someone back for another appointment, an open or a click is useful context, but the result I really care about is whether the person booked or returned. We’re also building a scheduled reevaluation, with a weekly default cadence, that will look at what changed, what previous campaigns achieved and what new opportunities appeared, then use that for the next recommendations.
I also don’t want the analysis to look only for problems. Sometimes the useful signal is a pattern that’s already working. If customers who book one service often come back for a related one, that can become a follow-up campaign.
Which technical decisions do you keep for yourself as CTO, and which do you hand to your team?
I tend to keep the decisions that define the system: the initial technical direction, architectural boundaries, decisions that affect several parts of the product, and problems that are still too vague to hand off cleanly. Once the goal, constraints and expected outcome are clear, I’m very comfortable giving an engineer ownership of the implementation.
For me, the boundary is mostly about the cost of being wrong. If a decision will be expensive to reverse, materially change infrastructure cost, create a dependency on a partner or force several parts of the team to align around it, I want to be involved. If it’s local and we can change it later without redesigning half the system, I’d rather let the engineer who owns that area decide. I still want to understand the assumptions and trade-offs behind a proposed solution, but it doesn’t need to be the one I would have designed myself.
Our data architecture is a good example. I stayed very close to the decision about whether we needed a common registry, when to introduce it and what it needed to represent, because that affected data ingestion, analysis, campaign recommendations and future integrations. For sending and data ingestion, we have an engineer with years of experience in that area, and it makes sense for him to own that part of the system.
Did anything from your years designing workflow architecture and a schema-driven integration framework shape the way you built AlpacaRelay?
A lot of it did. At airSlate, I worked on a framework where very different automations, called Add-ons and later Bots, could be described through a common schema. That meant thinking carefully about the attributes a Bot needed, the data it could work with and the functions it could perform, and describing those in a general form instead of building every feature as a separate one-off experience.
That changed the way I approach architecture. I tend to start by asking what the vocabulary of the problem is: what the entities are, what defines them, how they relate to one another and what is universal versus specific to one case. That transferred directly to AlpacaRelay, first in the terminology and structure for email generation and later in the abstractions for businesses, customers, services, events and campaigns.
airSlate also shaped how I think about outcomes. We looked at user behavior through events and funnels, and a document workflow had a visible lifecycle. It had to be set up, people and roles assigned, actions completed and documents signed or delivered before it reached a successful state. I think about email campaigns the same way. Sending the email is one event in that chain, and I want to know whether the chain reached what the business actually wanted to happen.
I also learned from a weakness of very flexible systems. The more configurable a product becomes, the more complexity you can accidentally hand to the user, and at airSlate setup could be a difficult point in the funnel. With AlpacaRelay, I want to hide more of that complexity. A customer should be able to describe a problem in plain language, or the system should find an opportunity in their data, and the complicated analysis and setup happen underneath.
Looking ahead, which parts of the system are you still testing, and what would convince you they work?
We can generate a good email. The harder question I’m testing now is whether the campaigns we recommend consistently create value for the business, and I look at that in stages.
The first signal comes before sending: did the customer even want the campaign we recommended? If they repeatedly reject a recommendation, something is wrong in the analysis, the campaign idea or the way we generated it. Then there’s delivery health. That part is in a good place, but I don’t consider it permanently solved, because sender and domain reputation and list quality can change, so we keep an eye on hard and soft bounces and complaint rates. The metric I care about most is the business goal. If the purpose is to bring a customer back for another appointment, I want to know whether they actually booked again.
What would convince me is seeing that whole chain repeat reliably: we identify a real opportunity, the business accepts the recommendation, the campaign reaches the audience and the intended outcome happens. The next test is whether that accumulated outcome data actually improves the next recommendations once the attribution loop is complete.
I’m also still optimizing the system underneath. Different stages of the pipeline have different responsibilities, so we treated model selection as part of the generation architecture and evaluated which models made sense for each stage. Together with our ontology work, that reduced our generation cost by roughly three times.



