Hacker News

MentalHealthBench

32 points by gmays ago | 14 comments

grafmax |next [-]

Mental health plus ad profiles, what could go wrong? I can't help but wonder about the moral fiber of the mental health professionals they say they consulted for this.

KellyCriterion |next |previous [-]

Wondering when the first insurance/etc. company is announcing a "partnership with OpenAI", i.e. shoveling all the data of their members to OAI :-D

rsfern |next |previous [-]

I find this troubling, there are so many potential ethical issues with this application of language models. Maybe it would be more appropriate for models to issue safety refusals and help the user learn how to get help from a licensed mental health professional

fnordpiglet |root |parent |next [-]

In the United States at least there’s a huge lack of mental health professionals, they’re expensive, insurance goes out of their way to make it inaccessible, and wait times for appointments can be months out. For psychiatrists in some regions there are zero that accept insurance and there can be a six or longer month wait. Many regions have no inpatient beds or any sort.

My parents are clinical psychologists with 40 years experience, and they both agree LLMs could be a great boon to many many unserved people if made well. They think this is a wonderful avenue of research. But agree in the moment it’s not ready. I think the ethics of allowing people to suffer due to economics while ignoring a technology that can help them is worse, no? Telling them to seek a health professional when none are available is worse ?

rsfern |root |parent |next [-]

I agree that accessibility is a big problem, and take your point about ignoring technology that could help, but if the technology if not there yet then I think more research and discourse is warranted before deploying it widely. As a society we need to decide what kind of requirements we expect of these systems, then agree on standards of performance

Is it ok for models to offer specific mental health interventions, or just provide neutral information? To what degree is that technically achievable? Does the answer depend on the scenario?

Health care workers have mandated reporting responsibilities (in the US, I assume other countries are similar) for some situations. Should models be bound by this responsibility as well? (I think yes but it is a thorny question). How to achieve that while preserving patient privacy for issues that don’t fall under mandated reporting rules?

We also have standards of care for health workers. Benchmarks are good, but that’s similar to a licensing requirement, and we have accountability mechanisms when people don’t adhere to their ethical and professional duties. What should be the liability situation when a model responds out of standard resulting in harm (to the standard of evidence we would hold a human health worker)?

Maybe I am not plugged in to this space enough, but my impression is these questions aren’t really in the foreground

fnordpiglet |root |parent [-]

I think the expectation that ChatGPT will be your therapist is wrong, however the expectation that you can build a therapeutic system using LLMs is not necessarily wrong. You are right that there are regulatory aspects, and liability, as well as guard rails built into our system (but less for mental health than medical as it’s largely a neglected aspect of our care delivery).

Also, it’s a different mode of delivery. Some people work well remotely, others better in person. Likewise some people open up more easily to a machine than a person. One doesn’t replace the other.

Further, LLMs can be very useful for diagnostic work in addition to a psychologist to help accurately classify mental and personality disorders. This is my fathers speciality and he finds modern LLMs to be fairly adept at classifying.

A lot of psychologist work in a clinical setting is case notes and an LLM can be quite skilled at producing standard expected case note structures from clinical notes, improving the time spent by the professional in care vs paperwork.

All of these are opportunities for use of LLMs as specialist models, agentic frameworks around general models, guard rails, delivery mechanisms, etc. I wouldn’t look at the world being OpenAI and Anthropic, but as model providers as providing screws and bolts and it’s up to us to build the products. LLMS are trained with an enormous corpus of witting on mental illness, but also writings by the mentally ill, care professionals, and all sorts of other materials. By eliciting all these dimensions the applications to mental health are enormous - it is a giant learned model of human thought and experience, not just of factual information, but the entire complex.

Here’s an idea I’ve talked to my dad about - how about using a model set trained on the writing of many people of specific mental illnesses. Then you can interact with them and they will produce responses likely of someone with that illness. This can be used to train clinicians, but also to build better inventories and diagnostics since you know the classification of the responses to some reasonable degree. This should accelerate research into mental illnesses.

The possibilities are vast, limited by our imagination and biases.

blueone |root |parent [-]

> Likewise some people open up more easily to a machine than a person. One doesn’t replace the other.

This is a fundamental personal failure. Not being able to open up to a professional whose purpose is to help you. The first step is to be honest with yourself, then honest with the person whose entire professional purpose is to help you.

AI psychologists will only ever be able to do so much, but I imagine it will be a whole lot of “you’re right” with very few harsh truths, criticisms, or personal development.

fnordpiglet |root |parent [-]

Dude, in that lens, all mental health is a personal failure.

The fact is a lot of people have trust issues, anxieties, social issues. Some are paranoid, suffer extreme neuroses, or panic in a social situation. Some have felt betrayed by professionals seeking to help them, putting them into inpatient care against their wishes, or medicating them with debilitating medications to sedate them, which is common for people experiencing mania. That’s not a personal failure, that’s just personal issue. And someone seeking mental help is having personal issues more often than not, unless they’re just seeking life coaching.

Ascribing these as fundamental personal failures, and therefore they deserve their personal hells with no relief unless they comply with your view of how people must be, is cruel.

The sycophantic behavior you’re outlining is a result of reinforcement learning and alignment, not a fundamental property of generative language models.

csnover |root |parent |previous [-]

> I think the ethics of allowing people to suffer due to economics while ignoring a technology that can help them is worse, no? Telling them to seek a health professional when none are available is worse ?

Not necessarily. Half-measures—especially those that try to offer technological solutions to social problems—allow the actual root problems to remain unaddressed, fester, and grow. The ACA was a hack for a failing system, and then when that quit working, more hacks were added on top, and on and on. LLMs are just another hack that let the can get kicked down the road once more.

If and until actual meaningful health care system reforms happen, the rich will continue getting their $63,000 14-day concierge service from Mass General Brigham whenever they feel blue, and the rest of us will end up getting Doctor Amazon telling us to practice deep breathing and buy some supplements.

fnordpiglet |root |parent [-]

I can’t argue this, but I can build an LLM based care provider today with likely decent results in patient outcomes for those adapted to it and good cost dynamics, and across a variety of critical use cases in the care vertical from providers to end delivery. I can not however change American society, our economic structure, or political dysfunction single handedly, can you? Should I not do the former and let people suffer in protest because the latter is not possible, and hold my breath hoping it’ll change? How many lives must be destroyed in protest hoping everything changes first?

Fine, let’s have socialized delivery of mental health care. Let’s go back in time 20 years and establish a pipeline of capable providers that doesn’t exist. Let’s do those things if we can. But I don’t see it happening in my life time.

I do however see an opportunity to take the technologies presented to use in the immediate present and providing a mechanism to bring help to people who need it today with good efficacy where it’s appropriate. Is it OpenAI that brings that? I doubt it - they’re selling shovels, not digging for gold. Is it a general model that provides direct care? Doubtful IMO. Is it a chatbot interface ? I doubt that too. But I think if we are creative and thoughtful we can use these technologies to address huge swaths of the care providing problem, which is more than talk therapy. And we can do it today, and don’t have to wait for MAGA/MAHA to accept reality on realities terms.

j45 |root |parent |previous [-]

The transcripts of mental health professionals could likely have an improvement in the baseline below-average experiences of most folks with such supports.

In other words, mental health professionals will resist this technology to replace them until they realize they can put it in front of people like a new kind of web form, at which time the same will become amazing and wonderful.

At the very least, a tool like this could serve as an anti-virus and firewall for poor and harmful experiences from mental health professionals towards clients.

cimi_ |next |previous [-]

How about putting a fucking red banner at the top of that page saying 'WE DON'T RECOMMEND ANYONE USE OUR STOCHASTIC PARROTS FOR THERAPY, WE DO THIS FOR DAMAGE CONTROL'.

I can't understand how anyone can do this in good conscience.

mzhaase |previous [-]

This seems to miss the point for me, therapy is a series of 20+ conversations, steered by the therapist.

Current LLMs would instead be steered by previous tokens, meaning the patient does. So someone that for example has schizophrenia may actually convince the LLM it's true.

appstorelottery |root |parent [-]

This is a multi-session problem, or at least this is how I've solved it in the past (i.e. the user chat/report thread isn't the one driving the conversation, the sub-sessions/workflows are guiding the analysis).

p.s. there's some interesting literature on how in some circumstances (i.e. religious delusions), where actually validating the patient - basically confirming their experience can lead to rapid recovery (however this approach is not found in western culture). I do agree with you in terms of the risks of LLM's driving folks deeper into delusional beliefs & behaviors - it's very real.