2 Comments
User's avatar
Cyril Simonnet's avatar

Triage as a design constraint is the right call, because the failure mode is not a bad reply, it is an unowned decision. A model that grades severity and routes it has to know who owns the escalation, what evidence it hands over, and what happens when nobody picks it up. Otherwise the gradient just produces a score nobody acts on, and the liability lands on whoever shipped the chatbot.

My piece argues the same thing from the other end. Governance is not a policy document sitting above the stack, it is the layer where the agent actually decides and hands off, and if that layer is ungoverned nothing above it holds. A triage layer without a named owner per risk band is exactly that gap.

https://cyrilsimonnet.substack.com/p/you-are-exactly-as-sovereign-as-your

Scott Wallace, PHD's avatar

I appreciate such a considered reply. Reading your piece, I think we’re making the same argument from different domains: you’re naming the sovereignty gap at the agency layer; I’m watching that same gap show up as unowned triage in mental health AI.

I believe that a model that grades severity and routes only counts as triage if each risk band has a named owner, clear evidence handoff, time‑bounds, and a defined “what happens if nobody picks this up.” Otherwise it’s the care equivalent of the GuardFall pattern you describe: an agent that can act on what it reads, whose limits and duties nobody actually set.

My argument is basically your last line, turned inward on clinical workflows: governance is not the policy sitting above the stack, it’s the layer where the agent decides and hands off. If that layer is ungoverned, you get exactly the gap I worry about most in mental health AI – risk detection that looks like safety from the outside, but lands on no one in particular when it matters.