The Mental Health AI Stack Needs a Triage Layer, Not Another Chatbot
Founders and clinicians need AI that sees gradients of risk, not just messages. That requires a triage layer in the stack
In early 2026, Chat & Ask AI exposed roughly 300 million private messages tied to 25 million users through a Firebase misconfiguration, including conversations from people asking how to end their lives. In practice, the database was left with security rules so open that anyone who knew the URL could read or potentially write to the data. This was an architectural failure that made the security lapse possible in the form it took.
What matters is not only that the data was exposed, but how the system understood it before it was exposed. Mental health AI is full of products that look sophisticated at the surface and remain clinically primitive underneath. The interface may sound warm. The prompts may feel thoughtful. The memory may even appear helpful. But in system after system, the underlying stack still collapses radically different human disclosures into one undifferentiated stream of data and then treats the whole stream as ordinary app telemetry.
A clinician who treated those disclosures as interchangeable would be negligent. Mental health AI should not be built that way either. A request for sleep hygiene advice, a disclosure of substance use, a trauma narrative, a suicidal plan, a therapy note, and a fleeting product preference are not versions of the same thing. They are different classes of event with different meanings, different legal protections, different retention implications, and different duties attached to them.
Mental health does not need less innovation. It needs innovation that starts with the reality that this category of data is different, and that systems built to hold it must be different too.
This is not an indictment of AI. It is a call to the field to stop applying generic AI product assumptions to a domain whose data is unusually intimate, unusually consequential, and unusually regulated.
The Argument Founders Need to Hear
A common mistake in mental health AI strategy is to assume the next breakthrough is better conversation. It is not. Better phrasing, more natural tone, and longer context windows can improve user experience, but they do not solve the central safety problem in this domain. One structural product gap is that most mental health AI systems have no meaningful triage layer between user disclosure and system action.
A triage layer is not just a classifier. It is the part of the architecture that decides what kind of event this is, what kind of data this creates, which legal regime applies, who may access it, how long it should be retained, whether it may be used for model improvement, whether the user should be warned, whether a human must review it, and whether the system should stop behaving like an open-ended chatbot altogether.
Without that layer, other improvements are downstream refinements that do not touch the core clinical risk. A system may sound safer while remaining structurally unable to distinguish ordinary support from acute risk. It may add memory while quietly increasing legal exposure. It may market personalization while expanding the amount of sensitive material stored in ways the user neither expects nor meaningfully controls.
Why Clinicians Should Care
Clinicians are being asked, implicitly and explicitly, to lend credibility to systems that operate on assumptions most clinicians would never tolerate in practice. The profession is grounded in distinctions: presenting problem versus acute risk, psychotherapy notes versus the designated record set, substance use records versus ordinary scheduling data, therapeutic rapport versus consumer engagement. Many AI systems collapse those distinctions at the data layer long before a clinician ever sees the interface.
That matters because licensed clinicians do not merely hold confidential information in the abstract. They work inside professional, legal, and ethical duties regarding what is documented, how it is stored, when it is shared, and when action is required. Information related to mental and behavioral health and the HIPAA mental health information sharing guidance make clear that these records sit inside specific duties and exceptions.
Substance use disorder records remain subject to the 42 CFR Part 2 final rule, which HHS updated in 2024 to allow a single consent for treatment, payment, and health care operations while preserving special handling requirements for this category of data. That does not mean the same obligations automatically transfer, intact, to every AI company handling mental health-related conversation. It does mean the opposite claim is dangerous.
A B2C chatbot is not a therapist because it uses reflective language. A foundation model provider is not suddenly outside consequence because it presents itself as infrastructure rather than care delivery. And a B2B company selling workflow automation into clinical settings does not escape mental health-specific risk simply because its product category is framed as software rather than treatment.
What the Law Already Says
The law already rejects the idea that mental health data is one thing. At minimum, systems in this space can touch four overlapping modes: covered-entity care under HIPAA, psychotherapy notes, substance use disorder records under the 42 CFR Part 2 final rule, and direct-to-consumer wellness or mental health data that may sit outside HIPAA but remain exposed to FTC enforcement, state consumer protection law, and in some jurisdictions stricter state mental health privacy rules.
That is the architecture problem stated in legal form. The same person can move through all four modes in one product journey: self-guided onboarding, support chat, clinician contact, relapse discussion, scheduling workflow, between-session journaling, AI-generated summary, product analytics. Technically, it is easy to route all of that into one logging and retrieval pipeline. Legally and clinically, that is the move that should raises the alarm.
Psychotherapy notes are a particularly pointed example. HIPAA defines them narrowly as notes recorded by a mental health professional that document or analyze the contents of a therapy session and are kept separate from the rest of the medical record. They specifically exclude information that belongs in progress notes, such as diagnosis summaries, treatment plans, session times, test results, prognosis, and documented progress.
Because psychotherapy notes receive heightened protection and usually require specific patient authorization before disclosure, a clinician cannot treat psychotherapy‑note‑equivalent material sent into an AI system as if it were ordinary documentation convenience. The legal category is different, and the obligations attached to it are different.
Psychotherapy‑note‑equivalent material is not “just more data.” It lives under a different rulebook. When you route it into an AI stack, you inherit that rulebook whether or not your product was designed for it.
In many circumstances, it may be an impermissible disclosure absent specific authorization, even where a business associate agreement exists. For clinicians, that raises the question: when an AI note assistant, ambient scribe, supervision tool, or recall system handles session‑level material, what exactly is it handling?
Does AI Inherit the Clinician’s Obligation?
The difficulty is that mental health AI rarely sits neatly inside a single role, a single duty, or a single legal relationship. That is the problem.
A licensed clinician’s duties arise from licensure, professional standards, privacy law, evidentiary doctrine, contractual arrangements, employer policy, and the specifics of the treatment relationship. A foundation model provider, by contrast, may occupy the role of processor, subprocessor, infrastructure vendor, or general-purpose model provider. A B2C company may operate outside provider status entirely. A B2B company may be a business associate in one deployment and a consumer data company in another.
The key distinction is not whether AI simply inherits the clinician’s obligations wholesale. It is which obligations attach at which layer, and what happens when the architecture blurs the layers so badly that nobody can tell where the responsibility actually sits.
This ambiguity is one reason mental health AI cannot be governed solely through interface disclaimers. The foundation model provider has obligations that may relate to training, retention, security, incident response, and contractual limits on sensitive use. The application company has obligations tied to collection, notice, consent, routing, storage, deletion, and product claims. The clinical organization has obligations tied to recordkeeping, supervision, scope of practice, and lawful disclosure. None of those layers can responsibly point to the others and say the problem lives elsewhere.
The Devil’s Advocate Case
There is a serious counterargument to tackle here. One could say that all of this caution risks freezing useful progress. Mental health systems are overburdened. Access is poor. Administrative load is crushing. People want continuity, responsiveness, and support between appointments. AI systems can clearly help with intake, scheduling, psychoeducation, structured check-ins, documentation support, signal detection, and some forms of guided self-management.
That argument is not wrong. In fact, it is one reason the stakes are so high. The case for better architecture is strongest precisely because AI can be useful in mental health.
If the field reduces every critique to anti-AI fear, it will miss the actual point: the more these systems matter, the less acceptable it is to run them on infrastructure that treats crisis disclosures, therapy-adjacent material, and product analytics as one undifferentiated stream.
A balanced position is this. AI can absolutely support mental health care and mental health-adjacent workflows. But support is not exemption. Utility does not cancel duty. And personalization is not a moral free pass for indefinite storage.
De-Identification Is Not a Shield Against Risk
A familiar move in health technology is to ask whether the problem becomes easier once the data is de-identified. For some purposes, it does. De-identification can reduce privacy risk, support analytics, and narrow the chance that sensitive material can be linked back to a named person. But in mental health, that answer is incomplete. De-identification is a privacy technique. It is not a clinical theory of responsibility.
A clinician cannot dissolve duty by mentally converting a suicidal patient into an abstract record. In a therapeutic relationship, obligations attach to the fact of risk inside a duty-bearing relationship, not merely to whether the note contains a direct identifier. Conducting research with participants at elevated risk for suicide, the HIPAA Privacy Rule and Sharing Information Related to Mental Health, and discussion of mandatory reporting when a patient may be violent all point toward the same principle: action follows risk, not merely record format.
If a patient discloses suicidality or a credible threat toward another person, the clinician’s responsibilities arise from the foreseeability and seriousness of that disclosure. Assessment, documentation, escalation, protective action, and in some contexts disclosure are not optional simply because one imagines the data could later be scrubbed, coded, or pseudonymized.
That is where the contrast with AI vendors becomes apparent. A foundation model provider, a B2C mental health app, and a B2B workflow company may each argue that de-identified or aggregated data is no longer meaningfully clinical and therefore can be treated as telemetry, safety tuning input, or model improvement material. Sometimes the law may indeed treat those actors differently from a licensed professional in an active therapeutic relationship. But that looser legal posture is not an architectural solution.
If the system can detect a high-risk pattern in human disclosure, the field still has to answer the question of what duty follows from that detection and at which layer of the stack it sits. De-identification cannot be the “escape hatch” in a mental health AI architecture. It may change what can be shared, studied, sold, or reused. It does not change the fact that some disclosures are action-relevant in real time.
A triage layer must distinguish two separate questions that product teams often collapse into one: can this data be retained or reused in a lower-risk form, and does this disclosure trigger a duty to act now?
Memory, Context, and Privacy
This is where my argument becomes more technical (and hopefully interesting). Memory and context are not decorative features in mental health support. They are central to usefulness. A system that remembers a person’s stressors, preferred coping strategies, medication concerns, relationship themes, or recent deterioration can feel far more coherent and more supportive than a stateless chatbot that begins every exchange from zero.
But personalization creates a privacy fork in the road. A naïve implementation is bulk accumulation: store everything, retrieve opportunistically, and call the result continuity. A more disciplined implementation is selective, user‑governed memory: only retain what is necessary, separate preference memory from acute‑risk memory, make memory legible to the user, allow editing and deletion, and bind retrieval rules to both clinical acuity and legal category.
Personalization does not have to mean maximal storage. It can mean better design. A privacy-respecting architecture might separate at least four layers of memory:
transient session context that expires quickly,
user-authored preference memory that is visible and revocable,
clinically sensitive state markers with stricter controls, and
acute-risk traces that trigger safety workflows but are excluded from ordinary retrieval and model training.
That is also where user agency either exists or does not. If a company says the system is personalized but cannot show the user what is remembered, why it is remembered, how long it is kept, whether it informs model training, whether it is shared with downstream vendors, and how to revoke it, then the personalization is not under user control. It is monitoring dressed up as support.
This is also a good place to acknowledge an earlier promise in digital health: that people would be in charge of their own health data. The strongest expression of that idea is data sovereignty and, in some frameworks, the self‑sovereign patient—a model in which individuals have meaningful control over access, consent, portability, and revocation for their health information. Those of us who worked on early NLP/NLG systems in health saw this tension directly. Much of the implementation discourse around this became inflated and technically overclaimed. The ambition often outran the actual plumbing, so to speak. But the core intuition was sound. Mental health AI should move toward architectures that increase intelligibility and control for the person disclosing, not architectures that quietly deepen asymmetry.
The Deeper Technical Problem
The safety weakness of many current systems is not simply that they miss dangerous phrases. It is that they are stateless in the wrong places and sticky in the wrong places. They forget what should be interpreted over time and remember what should have been tightly constrained.
A clinically credible mental health stack should be stateful about risk trajectory, because suicidality, eating disorders, trauma escalation, coercive control, psychosis, and relapse often reveal themselves across sequences rather than isolated utterances. At the same time, it should be deliberately non-sticky about broad retention, because indefinite preservation of intimate disclosures increases discovery exposure, breach harm, internal misuse risk, and user vulnerability if those records are later repurposed or compelled.
The key technical question is not whether a model has memory. It is what kind of memory, at which layer, under whose control, with what retention logic, and under which legal theory.
What a Triage Layer Actually Does
A triage layer would sit between raw conversation and everything downstream. It would classify both the type of data and the level of acuity at the point of entry. It would determine whether the content belongs in ordinary workflow context, protected clinical documentation, a substance use-protected segment, transient personalization memory, or a high-acuity safety channel.
A triage layer would also separate the questions that are too often collapsed into one:
What should the model see right now?
What should the application store?
What should a clinician review?
What can be used for quality improvement?
What must never enter training?
What can the user delete?
What must be retained because law or care obligations require it?
What should trigger immediate human intervention?
The mental health AI stack really does need a triage layer, not another chatbot, because the missing capability is not chat generation. It is disciplined sorting under conditions of intimacy, risk, and legal complexity.
Practical Recommendations
For founders
If you are building mental health AI, these are the minimum architectural moves that turn safety from a claim into a design.
Build data classification at ingress, not as a downstream tagging exercise after everything has already been captured.
Separate preference memory, workflow memory, clinical documentation, and acute-risk material into distinct stores with distinct permissions and retention schedules.
Exclude high-acuity and deeply sensitive content from training by default, and make the policy legible rather than burying it in generalized privacy language.
Design user-facing memory controls that allow inspection, correction, expiration, and deletion of personalized context wherever legally possible.
Do not market the system as therapy or confidentiality-equivalent support if the legal and operational protections of therapy are not actually present.
For clinicians and provider organizations
Ask vendors exactly what categories of mental health data they ingest, how they distinguish psychotherapy notes from other records, whether Part 2 segmentation is supported, and whether any session-level material reaches model training or third-party subprocessors.
Require clear answers on retention, deletion, discoverability, audit logs, human review pathways, and what happens when a conversational system detects escalating risk.
Treat memory features as a recordkeeping and privacy issue, not as a harmless convenience layer.
Distinguish documentation support from clinical judgment. A system can assist workflow without inheriting the authority to make legally or clinically loaded determinations on its own.
For the field as a whole
Start defining the architecture patterns that should be considered baseline for systems handling psychologically sensitive data: segmentation, tiered retention, stateful risk monitoring, memory transparency, training exclusion, and user agency over personalization.
Where This Leaves Our Field
Mental health data has unique properties and unique consequences. It can expose state of mind, risk of harm, trauma history, substance use, family conflict, social stigma, legal vulnerability, and the private disclosures on which treatment sometimes depends. Those properties demand deliberate architecture, not generic data handling.
The next real advance in mental health AI will not be another chatbot that sounds more human. It will be infrastructure that knows, from the first moment of disclosure, what kind of data it is holding, what kind of risk it may represent, what kind of memory it is allowed to keep, and when the system must stop acting like a chatbot at all.
In other words, the mental health AI stack does not need another layer of conversation; it needs a layer of triage. Until that exists, safety will be something systems talk about, not something they are built to provide.
Scott Wallace, PhD, is a clinical and neuropsychologist turned mental health technology strategist with 35+ years at the intersection of care and code. He has engineered some of North America’s earliest digital mental health platforms (C, C++, JavaScript, early iOS), developed several mental health mobile apps, and led NLP/NLG‑based conversational system design well before large language models went mainstream.
His workplace mental health programmes have been adopted by major employers across Canada and the United States, most notably reaching every employee at 3M worldwide, and his work helped TELUS earn Excellence Canada’s Gold awards for Healthy Workplace and Mental Health at Work under the National Standard for Psychological Health and Safety. He has produced award‑winning mental health video, led the digital division of a major EAP provider through a successful exit, and keynoted for multinationals including Scott Paper.
Clinically, Scott has worked with all age groups in hospital settings, the criminal justice system, and neuropsychological assessment (including temporal lobectomy candidates for refractory epilepsy), and spent two years as a regular national morning‑show psychologist in Canada. He now advises founders, health systems, and investors on AI‑enabled mental health, with a focus on AI safety, governance, clinical risk, and the unit economics of care.



Triage as a design constraint is the right call, because the failure mode is not a bad reply, it is an unowned decision. A model that grades severity and routes it has to know who owns the escalation, what evidence it hands over, and what happens when nobody picks it up. Otherwise the gradient just produces a score nobody acts on, and the liability lands on whoever shipped the chatbot.
My piece argues the same thing from the other end. Governance is not a policy document sitting above the stack, it is the layer where the agent actually decides and hands off, and if that layer is ungoverned nothing above it holds. A triage layer without a named owner per risk band is exactly that gap.
https://cyrilsimonnet.substack.com/p/you-are-exactly-as-sovereign-as-your