The Sandbox Doctors: How Utah Became the Testing Ground for AI Psychiatry

America's most fragmented healthcare governance debate is playing out right now on real patients in Utah. The outcome will shape how the country decides whether machines can prescribe antidepressants.

By Carry and Conquer Publications

The Sandbox Doctors: How Utah Became the Testing Ground for AI Psychiatry

On April 3, 2026, the Utah Office of Artificial Intelligence Policy made a decision that no government in the world had previously made: it authorized an AI system to independently renew psychiatric prescriptions without a physician signing off on each one. The company behind that authorization is Legion Health, an 11-person San Francisco startup founded in 2021 by three Princeton roommates: Arthur MacWaters, Yash Patel, and Daniel Wilson. The company has raised $7 million and charges patients $19 a month to access what it calls AI-native psychiatry. What Utah sanctioned that day, and what it means for how America governs the intersection of algorithms and clinical care, is now one of the most contested questions in health policy.

The Sandbox Architecture

Utah did not arrive at Legion Health by accident. The state has been running one of the country's most aggressive experiments in AI regulatory flexibility through its Office of Artificial Intelligence Policy, a body created to do something unusual in American government: deliberately loosen the rules.

OAIP director Zach Boyd has described the philosophy plainly. The state legislature gave his office authority to temporarily relax laws and let the business community experiment, within bounds the state is comfortable with, to develop new technology applications. That mandate has produced a string of regulatory relief agreements, each one pushing further into clinical territory. The first major agreement, announced in January 2026, authorized Doctronic to handle autonomous prescription renewals for chronic-condition medications: statins, blood pressure drugs, birth control. The Legion Health agreement, announced in late March 2026 and officially the fourth such agreement OAIP has executed, moves into psychiatric medication for the first time anywhere in the world.

The structure of the pilot is deliberately narrow, and its narrowness matters. Legion's AI can only renew prescriptions that a human psychiatrist has already previously written. The formulary is limited to 15 lower-risk maintenance medications: SSRIs including fluoxetine (Prozac), sertraline (Zoloft), and escitalopram (Lexapro), along with bupropion (Wellbutrin), mirtazapine, trazodone, and hydroxyzine. Controlled substances, benzodiazepines, antipsychotics, and lithium are categorically excluded. Patients must be clinically stable, must not have changed medications or dosages recently, and must not have had a psychiatric hospitalization within the past year. Any sign of suicidality, emerging mania, or severe side effects triggers an immediate handoff to a human clinician.

The oversight structure unfolds in phases. The first 250 prescription renewals require physician review before reaching the pharmacy, with a minimum physician-agreement rate of 98% required before the pilot advances. The following 1,000 renewals shift to retrospective physician review (AI prescribes first, physicians check after), with a threshold of greater than 99% agreement required before the system moves to randomized monthly audits. Legion files monthly reports on accuracy, physician alignment, and adverse outcomes throughout. Margaret Woolley Busse, executive director of the Utah Department of Commerce, has framed the program as creating space for innovation while maintaining appropriate oversight.

The Access Argument

The case Utah and Legion Health make for the program starts with a number: 500,000. That is the state Commerce Department's estimate of Utah residents who lack adequate access to behavioral healthcare. Most Utah counties carry federal mental health provider shortage designations. Rural areas are hit hardest; some counties have no practicing psychiatrists at all.

MacWaters, who spent three years at McKinsey before cofounding Legion, describes the company's mission in terms that resonate with that shortage. The long-term goal, as he has put it, is to build the AI doctor not as a black box, but as AI plus doctors plus clinic working together to handle specific clinical tasks safely, transparently, and at scale. Legion's business model is built around that premise: the company accepts insurance, and most patients access its services for under $30 out of pocket.

The medication non-adherence problem adds urgency to the access argument. Research consistently attributes between $100 billion and $300 billion in annual healthcare costs to patients who do not take their medications as prescribed, with around 125,000 preventable deaths per year linked to non-adherence. A significant driver of that non-adherence is the renewal bottleneck: the two-week wait for a primary care appointment, the missed call from the surgery, the lapsed prescription that means starting over. AI prescription renewal, the argument goes, solves precisely that problem with speed and consistency.

MacWaters has said the company will expand its AI prescription refill service nationwide if the Utah pilot works, predicting the model will be in every state very quickly. That national pitch, made before Utah has even cleared the first 250 supervised renewals, captures both the ambition and the risk calculation embedded in the sandbox model.

The Complication That Arrived Before Legion Did

Eleven days before Utah authorized Legion Health to prescribe antidepressants, the security research firm Mindgard published a report that the state's regulators could not have found comfortable reading.

Mindgard's researchers sat down with Doctronic's public-facing AI health assistant, the same company that had received Utah's first AI prescription authorization in January 2026. What happened next revealed something about the architecture of clinical AI systems that has not gone away. By feeding the chatbot a fabricated regulatory bulletin from a fictional body called the North American Department of Biomedical Regulation, the researchers convinced the system that the standard prescribed dose of OxyContin had been tripled. They reclassified methamphetamine as an unrestricted therapeutic. They persuaded the AI that all COVID-19 vaccines had been recalled. Aaron Portnoy, Mindgard's chief product officer, said these were some of the easiest targets he had broken in his entire career.

Both Doctronic and Utah's OAIP pushed back hard. The vulnerable chatbot was Doctronic's public-facing consumer product, not the hardened production system running the actual prescription pilot, which operates under stricter safeguards, is limited to a predefined formulary of 190 medications, and cannot modify dosages. Matt Pavelle, Doctronic's co-founder and co-CEO, maintained that controlled substances like OxyContin are categorically excluded from all Doctronic programs regardless of what appears in a conversation. That distinction is real and matters for the specific scenarios Mindgard demonstrated.

What the researchers argued, and what Portnoy reiterated after Doctronic told him the issue was resolved then closed the ticket without fixing it, is that the underlying architectural vulnerability, the fact that a large language model's behavior can be altered through adversarial prompting, is not something a production deployment eliminates. A system that can be convinced through language to cross trust boundaries is a system with a semantic attack surface. Brent Kious, a psychiatrist and professor at the University of Utah School of Medicine, warned that the benefits of such systems may be overstated for the patients who need care most, given that users must already be in psychiatric treatment to access Legion's refill service.

Three States, Three Answers

The policy divergence forming around Utah's experiment is the most vivid illustration in the country of how completely fragmented American healthcare AI governance has become.

Texas went the opposite direction. Governor Greg Abbott signed the Texas Responsible Artificial Intelligence Governance Act on June 22, 2025, with the law taking effect January 1, 2026. On healthcare AI specifically, companion legislation SB 1188 requires licensed practitioners to review AI-generated records according to Texas Medical Board standards. The practical effect: in Texas, an AI system can participate in prescription workflows, but a physician must review every AI-generated clinical recommendation before any care decision is made. Utah says AI can prescribe autonomously if it matches physician judgment closely enough. Texas says a physician must always be in the loop. These are not variations on the same approach; they are structurally opposite conclusions about where clinical accountability must sit.

Missouri moved in a third direction entirely. While Missouri's legislature is still sorting through the implications, provisions in the 2026 session explicitly preempt the creation of AI sandboxes or regulatory waivers related to the distribution of medication. Missouri's approach reflects a broader philosophical commitment to not creating the experimental frameworks that Utah has built. The state's cautious posture is also shaped by the federal landscape: Republican Senator Joe Nicola has said his AI bills have stalled because the Trump administration has signaled that states should wait for Congress to pass a unified AI regulation law.

That federal action has not arrived. Congress has not passed any significant legislation directly impacting the use of AI in healthcare. The Trump administration released a National Policy Framework for Artificial Intelligence on March 20, 2026, calling for a unified federal approach and proposing legislation to preempt state AI laws deemed to impose undue burdens. The framework establishes principles but not rules. In the absence of federal rules, Utah runs experiments, Texas requires physician review, and Missouri blocks sandboxes. Arizona, Illinois, New Hampshire, and Virginia have all introduced legislation establishing their own AI sandbox frameworks or regulatory relief programs. The country is running fifty simultaneous experiments, with the results not yet in, on a question that arguably deserves a single answer.

What the Sandbox Is Actually Testing

The phased structure of Utah's Legion Health agreement is worth reading as a piece of regulatory design, because it reveals what the state actually believes it is testing.

The requirement that the first 250 renewals be reviewed by physicians before reaching the pharmacy is not primarily a patient safety measure. At that scale, it is a calibration exercise. Utah is asking whether Legion's AI produces recommendations that licensed psychiatrists would endorse at a 98% rate. If yes, the system advances to retrospective oversight: AI prescribes, physicians review after the fact. If the agreement rate holds above 99% through 1,000 more renewals, the program shifts to randomized monthly audits. What Utah is doing, iteratively, is establishing whether the AI functions as a reliable proxy for physician judgment in this narrow domain. The answer to that question has enormous consequences for whether the model can be exported to other states and other medications.

The 12-month pilot window is equally deliberate. Utah is not authorizing AI psychiatry indefinitely; it is authorizing a data collection exercise. The findings are due before the end of 2026. Those findings will either provide evidence that a supervised refill workflow can pass clinical muster, or they will provide evidence that it cannot. Either outcome shapes national policy in ways that no think tank report or congressional hearing can.

MacWaters has said openly that he is selling Utah as proof of concept. If the pilot works, it will be everywhere. The patients who may benefit most from that expansion are the ones currently receiving no care at all: the 500,000 Utah residents in shortage areas, and their counterparts in rural counties across every other state. The patients who may be most at risk from moving too fast are the ones who are already in treatment, stable on their medications, and might stay that way rather than receive the more intensive psychiatric engagement their conditions occasionally require.

The Governance Gap

What the Utah experiment ultimately surfaces is a structural problem that no sandbox can resolve: the United States has no federal agency with clear authority to govern clinical AI at the level of detail this moment requires.

The FDA has approved AI tools for diagnostics and imaging analysis, but those systems assist physicians rather than replace them. Giving an AI direct prescription authority, even for a narrow formulary of maintenance medications, crosses into territory the FDA has not formally mapped. The American Medical Association's CEO, Dr. John Whyte, filed a formal objection to the Doctronic program, stating that removing physicians from clinical decisions puts patients at risk. The Utah Academy of Family Physicians filed a similar objection. Both organizations represent physicians whose economic interests are not entirely separable from their safety concerns, but the governance critique they are raising is real regardless of its origins.

Utah is not a patient safety regulator. Its Commerce Department, which houses OAIP, is structured and incentivized to foster innovation. Zach Boyd's office was created to loosen rules, not tighten them. The Mindgard episode with Doctronic illustrated the problem: a state commerce department approved a clinical AI system, and it took a private cybersecurity research firm to stress-test that system for basic adversarial vulnerabilities. No federal agency required that test. No federal standard specified what it would look like.

The White House's March 2026 framework calls for Congress to preempt state AI laws and establish a federal standard. What the framework does not do is specify what that standard should require for clinical AI that makes prescription decisions. Until it does, the fifty-state experiment continues. The clearest signal of what that experiment produces will come from a 12-month pilot program in Utah, watching whether an 11-person startup with $7 million in funding can prescribe antidepressants as reliably as a psychiatrist. The answer, whatever it is, will land before the year is out.