
I trained as a physician before I built a company, and medicine taught me to think of the prescribing decision as something that happens in the room. Write the script, and the decision is made.Â
Practice taught me otherwise. Plenty of prescriptions I wrote were never filled. The patient had already been forming a view before the visit, and would revisit it afterward, deciding whether the therapy was affordable and worth continuing. At the pharmacy counter a substitution could undo the choice entirely, and I would never learn that any of it had happened.Â
The decision was never a moment. It was a sequence, spread across days and three sets of hands, and the part of it I could see was the smallest part.Â
That is also the hardest problem in this industry, because following a decision across a sequence is exactly what an identity graph exists to do, and health data withdraws that option rather than restricting it.Â
The head start was imposed, not earnedÂ
For two decades, digital advertising has run on a premise it never had to examine: that coordinating anything across time requires knowing who you are speaking to. Identity graphs, cookie syncs and lookalike modeling all exist to serve it. Healthcare marketing never had permission to hold that premise, and I have come to believe it was the most useful constraint the industry was ever handed.Â
The prohibition on keeping a person-level record of someone’s condition arrived early enough to shape everything built afterward. The HIPAA Privacy Rule’s compliance date was April 2003. Maryland became the first US state to cap ordinary commercial collection at what is reasonably necessary, with consent unable to widen the limit, in October 2025. The same principle, twenty-two years apart.Â
That interval repeats at every layer that now matters for building AI. Prescriber identity has been resolvable against a federal registry since 2007, while the open web lost its cookie replacement last year with nothing shared to put in its place. Aggregate measurement was the only option available in healthcare from the outset, and open-web measurement reached modeled conversions in 2025. Provenance for every individual claim has been mandatory in pharmaceutical promotion since long before generative models existed.Â
I want to be careful about what I am claiming. None of this was foresight. This industry reached these positions by being penalized first, and tracking-pixel settlements across it now run well into nine figures. We were early to the problem, not wise about it.Â
Three people, three permission regimesÂ
Three people hold different parts of that arc, each governed differently. A licensed clinician acting professionally is identifiable through public federal enumeration. A patient carrying a diagnosis sits behind HIPAA and, increasingly, state consumer health statutes. A pharmacist at the point of dispense works inside a third system with obligations of its own.Â
Connecting the last step to the first ordinarily requires establishing that both involve the same person. The older sensitive-data regimes offer no precedent here, since credit reporting and financial privacy law govern how information about identified people may be shared and assume identification itself is permitted. This is the only case I know of where a commercial function lost identification of the end subject entirely and was still expected to influence individual decisions.Â
Identity was standing in for something elseÂ
Here is what the industry got wrong for years, and what took us too long to see. Nobody in pharmaceutical marketing ever needed to know who the patient was. What mattered was knowing that a treatment decision was underway, which is a materially different fact. Advertising built identity graphs because it could not observe the events it cared about, so it followed people around as an approximation of them. The profile was always a substitute for the event.Â
A clinical workflow produces those events directly. A diagnosis entered during an encounter, a procedure ordered, a prescription written at the point of decision. Each is timestamped, verified, generated by the workflow itself, and sufficient to establish that a decision is happening without anyone needing to know whose it is.Â
That makes the event usable as a coordination anchor across all three participants. Every activation keys off the same verified occurrence rather than a stitched-together patient profile, so the moment before the visit, the moment inside the workflow and the moment at the counter cohere because they reference one event. Activation runs against a clinical rule, and the prescriber it reaches is identified deterministically while the patient never is. That asymmetry is the whole design, and treating the two as one addressable population is where this industry has repeatedly gone wrong.Â
What the AI is actually doingÂ
Once events replace patient profiles, the modeling work changes shape, and four distinct jobs emerge.Â
The first is interpretation. A coded event on its own says very little, becoming meaningful only alongside the clinician’s specialty, the stage of the workflow, the treatment history attached to the encounter, and the coverage context that determines whether a therapy is even available. Weighing those inputs while a decision is still live is a modeling problem, and it constrains where the inputs can come from, because a coded event exists only inside the system that produced it. Tag-based or secondhand access sees an impression after the fact and never the coded occurrence, which is a data quality limit before it is a compliance one.Â
The training data is the part I find most interesting. The system learns from clinical events and their outcomes rather than from individuals and their accumulated histories, so there is no patient-level behavioral profile inside the corpus to protect, breach or unwind, and every label attaches to a verified prescriber rather than to an inferred likely one. Clinical history has real analytical value, but inferred intent decays quickly, and a probability is not a record of anything having happened.Â
The second is generation. Pharmaceutical promotion cannot say anything that has not cleared medical, legal and regulatory review, and an unapproved claim produces a regulatory record rather than an awkward brand moment. A general-purpose model writing free text with a review queue behind it cannot meet that standard. What works is constrained composition, where output is assembled from a library of pre-cleared claims with provenance attached to every assertion. Much of the enterprise software market arrived at a version of this in the past two years and calls it grounding.Â
The third is conversation. Clinicians now put questions to AI inside the workflow itself, and when a brand appears in an answer about dosing or comparative evidence, the failure mode is a clinician acting on poor information rather than a wasted impression. That sets a standard for retrieval accuracy and citation which advertising has never had to meet, and it is the standard every organization deploying a customer-facing assistant is now discovering for itself. It is also the only version of this I would want a physician treating my own family to be reading.Â
The fourth is agency. These systems increasingly act rather than advise, assembling plans and adjusting activation. An agent permitted to act inside a regulated workflow needs scoped permissions, a reviewable record of what it did and on whose authority, and a boundary it cannot cross without a human. That is close to what Europe’s AI Act will require of high-risk systems across every sector.Â
Why it resists being added afterwardÂ
A training corpus without lineage can assert data minimization but never demonstrate it, and lineage means the whole path from the clinical environment to the message, which every point of resale along the way degrades. Custody is precisely what regulators and courts now examine. A feature store built on identity-level joins requires rewriting rather than reconfiguring, and legal basis cannot be reconstructed after collection, so a record gathered without a documented basis for a given use has to be excluded and the model retrained.Â
Deletion exposes all of it at once, because a system unable to trace a record from source through feature to trained artifact can honor a deletion request against a database but not against a model.Â
This is what should concern anyone building AI outside healthcare. A model inherits the legal status of the data it learned from. Twenty states now have comprehensive privacy laws in force, and the strictest of them limit what may be collected rather than simply requiring disclosure, which means a system trained on person-level behavioral profiles carries that exposure into every prediction it makes and cannot be separated from it without being retrained.Â
My conviction is that the industries treating this as a compliance exercise have misread it. It is an architectural decision, it is made once, and it is made at the point where you choose what your systems are keyed to. Healthcare marketing had that choice made for it two decades ago and spent most of the time since learning what it cost. The rest of the market is choosing now, with far less time and no regulator obliging it to get the answer right.Â


