Concept PRD · Mental Health · AI/ML · Regulated Domain
Precision Psychiatry PM
A product vision for using neuroimaging data and machine learning (specifically graph neural networks trained on brain connectivity patterns) to match patients with depression and anxiety to targeted treatment pathways, moving mental healthcare from trial-and-error toward precision medicine.
AI/ML
Health Tech
FDA Regulated
Ethical AI
Concept PRD
Problem
Mental healthcare runs on trial and error.
50%
of depression patients don't respond to their first prescribed medication
11 yrs
average time between symptom onset and correct diagnosis for mood disorders
$225B
annual cost of depression and anxiety in the US, including lost productivity
When a patient presents with depression or anxiety, a psychiatrist selects a treatment based almost entirely on self-reported symptoms and clinical interviews. There is no biomarker test, no objective measurement. You prescribe, wait 6-8 weeks, and find out if it worked.
This isn't a failure of care. It's a data problem. We haven't had reliable ways to measure the biological underpinnings of mental illness. That is changing. Advances in neuroimaging, genomics, and machine learning now make it possible to identify patterns in brain activity that predict treatment response with meaningful accuracy.
The product opportunity: a clinical decision support tool that surfaces these insights at the moment of prescribing, not replacing clinical judgment, but giving it an objective foundation.
The Science
What makes this possible now.
Precision psychiatry is not a new idea. Three converging factors make it viable as a product in 2026 in ways that were not true even five years ago.
ML Prediction Accuracy: Treatment Response by Input Type
Not actual artifact, generated via UXPilot for privacy reasons
Neuroimaging at scale
Large-scale fMRI datasets (UK Biobank, HCP, ABCD Study) have made it possible to train models on brain connectivity patterns across tens of thousands of patients. What required a research lab in 2015 can now be done with cloud ML infrastructure.
ML model maturity
Graph neural networks and transformer architectures have shown strong performance on brain connectivity data, achieving 65-72% accuracy in predicting SSRI response in recent research, compared to ~50% for clinical judgment alone. Not a silver bullet, but a meaningful signal.
Regulatory momentum
The FDA has approved several AI-based clinical decision support tools under De Novo pathways in recent years, establishing precedent for ML-assisted psychiatric tools. The regulatory path, while complex, is navigable.
Discovery & Stakeholder Research
Three stakeholders, one system.
A clinical decision support tool sits at the intersection of patients, clinicians, and payers. Each has different incentives, constraints, and definitions of success. Getting the product right requires understanding all three.
🧠
Primary User
Psychiatrist / Prescriber
- Wants objective data to support treatment decisions without replacing clinical judgment
- Fears liability from over-relying on algorithmic recommendations
- Needs seamless EHR integration and will not adopt a separate tool
- Skeptical of "black box" AI; needs explainability to trust outputs
- Reimbursement for additional diagnostic steps must be justified
🙋
End Beneficiary
Patient
- Wants faster path to a treatment that works
- Highly sensitive to privacy, as neuroimaging data feels deeply personal
- Needs informed consent that's genuinely understandable
- Concerns about insurance discrimination from mental health data
- Wants to know how data will be used and stored
🏥
Economic Buyer
Health System / Payer
- Wants to reduce cost of failed treatment episodes
- Needs clinical validation data (RCT or real-world evidence)
- Requires HIPAA compliance and BAA before any data sharing
- Will only reimburse with CPT code coverage for the diagnostic step
- Values outcomes data: reduced hospitalizations, improved adherence
AI Product Design
Designing for high stakes.
This is not a chatbot. The AI design decisions here carry clinical and legal weight. Every choice, from model architecture to output format, needs to be made with that gravity in mind.
Model architecture
This product does not use a large language model. The core AI is a graph neural network (GNN) trained on fMRI brain connectivity data, classifying patients into treatment response profiles. A secondary LLM layer translates model outputs into plain-language clinical summaries for the prescriber. These are two separate systems with different validation requirements.
Training data requirements
The model requires labeled neuroimaging datasets with known treatment outcomes, ideally longitudinal data showing which patients responded to which treatments over 12 months or more. Minimum viable dataset: ~5,000 labeled patient scans across depression subtypes. Initial data partnerships with academic medical centers (e.g., Stanford, UCSF, Mass General) are required before any commercial training.
Bias & fairness: critical risk
Neuroimaging datasets are historically skewed toward white, highly educated, Western populations. A model trained on biased data will produce biased treatment recommendations, potentially at the expense of already-underserved mental health populations. Bias auditing across race, gender, age, and socioeconomic status is not optional. It is a launch blocker.
Explainability requirement
Psychiatrists will not act on a black-box recommendation. The product must surface which brain regions and connectivity patterns drove the model's output, using attention visualization or SHAP values, in a form a clinician can evaluate. "The model recommends X" is not sufficient. "The model identified reduced prefrontal-amygdala connectivity consistent with SSRI responder profiles" is.
Output design
The product outputs a ranked probability distribution across treatment options, not a single recommendation. It surfaces confidence intervals, flags low-confidence cases, and always presents outputs as "supporting evidence" rather than directives. The psychiatrist retains full clinical authority. This is a deliberate product and liability decision.
Privacy & data architecture
Neuroimaging data never leaves the health system's infrastructure. The model runs on-premise or in a HIPAA-compliant VPC. Federated learning architecture allows model improvement across sites without centralizing patient data. All data access requires explicit patient consent under a plain-language consent framework co-designed with bioethicists.
Regulatory Pathway
The FDA path to market.
A clinical decision support tool that influences treatment selection for a psychiatric condition is a Software as a Medical Device (SaMD) under FDA jurisdiction. The likely pathway is De Novo, which is used for novel, low-to-moderate risk devices without a predicate. Here's how I'd structure the approach.
Phase 1
Pre-Submission Meeting (Q-Sub)
Engage FDA before filing to align on intended use, risk classification, and required clinical evidence. This is a free FDA program and dramatically reduces rejection risk. Key questions: Is this a decision support tool or a diagnostic? What level of clinical validation is required?
Months 1-6
Phase 2
Clinical Validation Study
Prospective multi-site study across 3-5 academic medical centers. Primary endpoint: improvement in treatment response rate at 12 weeks vs standard of care. Secondary endpoints: time to remission, clinician confidence scores, patient satisfaction. Target: 500+ patients per arm.
Months 6-24
Phase 3
De Novo Submission
File De Novo request with clinical study data, algorithmic performance documentation, bias audit results, cybersecurity assessment, and labeling. De Novo review typically takes 12-18 months. If approved, establishes a predicate for future 510(k) submissions in the category.
Months 24-42
Phase 4
Commercial Launch + Post-Market Surveillance
Limited launch with 10-15 health system partners. Post-market surveillance is required, covering ongoing monitoring of real-world outcomes, adverse events, and model drift. FDA requires re-submission if the algorithm is retrained with new data that changes performance characteristics.
Month 42+
Not actual artifact, generated for portfolio purposes
Product Vision & PRD
What we're actually building.
The MVP is a clinical decision support module embedded in the psychiatrist's EHR workflow. It surfaces a treatment response probability profile at the moment of prescribing. No separate app, no separate login. Here's the core PRD framework.
North Star Metric
% of patients achieving remission at 12 weeks in practices using the tool vs matched controls. Secondary: time to first effective treatment.
MVP Scope
Depression (MDD) only. Three treatment dimensions: SSRI vs SNRI vs therapy-first. EHR integration with Epic via SMART on FHIR. Explainability view for clinician. Consent capture workflow for patient.
Customer & GTM
Initial customer: large academic medical centers with existing neuroimaging infrastructure and research appetite. GTM: partnership model with 3-5 founding health systems who co-fund the clinical study in exchange for early access and co-authorship.
Business Model
B2B SaaS to health systems, priced per-scan or per-seat. Reimbursement path via CPT code for "AI-assisted psychiatric assessment." Long-term: licensing to EHR vendors (Epic, Cerner) as a native module.
What's Out of Scope (V1)
Anxiety disorders, bipolar, schizophrenia. Direct-to-consumer. Autonomous prescribing. Genomic data integration. Real-time monitoring. Pediatric populations.
Success at 18 Months
5 health system pilots live, 1,000+ patients assessed, clinical validation data submitted with De Novo application, zero reportable adverse events attributable to tool recommendations.
The user journey for the prescriber:
- Patient consent captured during intake via a plain-language form co-designed with bioethicists and stored in the EHR
- Neuroimaging scan ordered as part of the standard psychiatric workup, integrated into the existing radiology workflow
- Model runs automatically on scan data within the health system's own infrastructure. No data leaves the facility.
- Treatment probability profile surfaces in the EHR at the moment of prescribing. The clinician sees ranked options with confidence intervals and brain region explainability.
- Prescriber makes the final decision. The tool is always advisory, never directive. The decision and any override rationale are logged automatically for post-market surveillance.
Risks & Mitigations
Hard problems with real answers.
Surfacing risks is not pessimism. It is the job. Every risk here has a mitigation path, because a PM who finds the problems is also the person responsible for solving them.
Bias & Equity
Model underperforms for underrepresented populations
Neuroimaging datasets historically skew toward white, educated, Western populations. A model trained on biased data produces biased recommendations, potentially widening the mental health disparities it is meant to address. Mitigation: mandatory demographic performance reporting as a launch condition, diversity requirements built into data partnership agreements with academic medical centers, and a published equity commitment reviewed by an independent ethics board before any commercial deployment.
Critical
Liability
Clinicians defer to the model when they shouldn't
Even with "advisory only" framing, a time-pressured psychiatrist may follow the recommendation without critically evaluating it. Mitigation: mandatory explainability (the model must show its reasoning, not just its output), required override documentation in the EHR, and liability language in health system contracts clarifying that the prescriber retains full clinical responsibility. Legal framework developed with malpractice counsel before any clinical deployment.
Critical
Adoption
Psychiatrists don't trust the tool and don't use it
Clinical culture change is slow, and trust in algorithmic systems in psychiatry is low. Mitigation: co-design the output format and explainability interface with psychiatrists from day one, treating them as co-developers rather than end users. Run a structured pilot with 10-15 clinicians before any broad rollout. Track clinician override rates as a leading indicator of trust. If override rates are high, that's a signal to improve explainability, not to push adoption harder.
High
Reimbursement
No CPT code coverage means health systems can't justify the cost
CMS reimbursement for AI-assisted psychiatric assessment is not yet established. Without it, health systems absorb the cost entirely. Mitigation: pursue a parallel reimbursement advocacy track from year one, working with a clinical economics team to build the ROI case. Use the clinical validation study to generate real-world outcomes data that supports a CPT code application. The founding health system partners fund the pilot in exchange for early access and co-authorship on the outcomes research.
High
Model Drift
Performance degrades as treatment protocols evolve
A static model trained in 2026 will drift from its validated performance as clinical practice and patient populations change. Mitigation: a continuous monitoring eval running against a held-out validation set post-launch, with a defined performance threshold that triggers re-validation. FDA notification is required for performance-altering updates under the approved De Novo, so the drift detection system needs to run before FDA thresholds are crossed, not after.
Medium
User Journey
What the experience actually looks like.
The product has to work for three different people: a patient consenting to something they may not fully understand, a psychiatrist under time pressure, and a health system trying to contain costs. Here's how the journey maps across all three.
| Actor |
Step 1 |
Step 2 |
Step 3 |
Step 4 |
Step 5 |
| Patient |
Referral Referred by GP for psychiatric evaluation. Receives pre-visit materials explaining the neuroimaging component. |
Consent Reviews plain-language consent form. Asks questions. Explicitly opts in to neuroimaging data use. |
Scan Completes fMRI scan as part of standard intake. ~45 minutes. No additional appointment required. |
Consultation Meets psychiatrist. Learns treatment recommendation and the reasoning behind it. |
Treatment Begins recommended treatment. Outcome tracked at 4, 8, and 12 weeks for post-market surveillance. |
| Psychiatrist |
Patient intake Reviews patient history and symptom profile in EHR ahead of consultation. |
Consent confirmed EHR flags that patient has consented to AI-assisted assessment. No additional action required. |
Model runs Model processes scan automatically. No separate login or tool needed. Output appears in the EHR at the moment of prescribing. |
Reviews probability profile Sees ranked treatment options with confidence intervals and brain region explainability. Makes final decision. |
Documents rationale Decision and override (if any) logged automatically. Feeds post-market surveillance data. |
| Health System |
Tool deployed Integrated into Epic via SMART on FHIR. Activated for consenting psychiatry patients. |
Consent workflow live Consent capture built into standard intake flow. Compliance team signs off. |
Model runs on-prem Data never leaves the health system's infrastructure. HIPAA compliance maintained. |
Outcomes tracked Dashboard shows treatment response rates vs baseline. Data supports reimbursement case. |
ROI measured Fewer failed treatment episodes, reduced hospitalizations, improved patient retention metrics. |
Not actual artifact, generated for portfolio purposes
Evals Framework
Measuring what actually matters.
Evals are to AI products what unit tests are to traditional software, except harder, because LLM outputs are non-deterministic and "correct" is rarely binary. For a high-stakes clinical tool, defining and running the right evals isn't optional. It's the foundation of trust.
I'd structure the eval program around the three-step lifecycle: Analyze failure modes qualitatively, Measure them quantitatively, Improve based on evidence. That cycle runs continuously, not just at launch.
Step 1
Analyze
Inspect real model outputs on representative patient scans. Qualitatively identify where the model fails, and not just when it is wrong but how it is wrong. Does it fail on certain demographic groups? Certain symptom profiles? Does the plain-language summary misrepresent the model's confidence? You can't measure what you haven't observed.
Step 2
Measure
Build specific evaluators for the failure modes you found. Run them at scale across your validation dataset. This is what turns qualitative observations into prioritized problems. Without measurement, you are guessing which fixes matter most. In a clinical context, that is not acceptable.
Step 3
Improve
Make targeted interventions based on measurement data: prompt changes for the LLM summary layer, retraining for the GNN on underrepresented populations, architecture changes for specific failure modes. Then cycle back to Analyze. The loop never stops. It just gets faster as your eval infrastructure matures.
The specific metrics I'd build evaluators for:
Calibration accuracy
Does the model's 70% confidence score actually correlate with 70% treatment response in real-world outcomes? Miscalibrated confidence is more dangerous than low accuracy.
Code-based
Demographic parity
Is accuracy consistent across race, gender, age, and socioeconomic status? Any gap here is a launch blocker, not a post-launch fix.
Code-based
Explainability faithfulness
Does the brain region attribution shown to the psychiatrist actually match what drove the model's prediction? Misleading explainability is worse than no explainability.
Human
Summary accuracy
Does the LLM-generated plain-language summary accurately represent the model's probability output? Evaluate for hallucination, overconfidence, and clinically dangerous misrepresentation.
LLM-as-judge
Clinician override rate
What % of cases does the psychiatrist override the top recommendation? High override rate signals low trust or poor calibration. Zero override rate signals over-reliance, which is also a problem.
Human
Model drift detection
Does performance degrade over time as patient populations or treatment protocols evolve? Requires a continuous monitoring eval running against a held-out validation set post-launch.
Code-based
Why This Matters
The PM's role in high-stakes AI.
I did not build this because I expect to ship a psychiatric AI tool. I built it because the thinking required to do it right is exactly what an AI PM needs in any high-stakes domain.
The questions here show up everywhere AI touches consequential decisions: how do you validate a system that influences outcomes you can't easily measure? How do you design for explainability when users need to trust, not just use, the output? How do you move fast without moving recklessly?
These are not engineering questions. They're product questions. The PM's job is to own them explicitly, not hand them off to ethicists or lawyers after the code is written.