“It’s Probably Nothing, But…” Why We’re Asking AI About Our Health, And Why That Should Give Us Pause

AI health questions

“It’s Probably Nothing, But…” Why We’re Asking AI About Our Health, And Why That Should Give Us Pause

  • Post comments:0 Comments

AI health questions are becoming a normal part of modern life. When symptoms appear late at night and medical help feels slow or hard to access, asking a chatbot can seem like the quickest and easiest next step. But while AI can be useful for understanding information and preparing for a conversation with a clinician, it is still far less reliable when people start treating it like a diagnosis.

It’s 11pm. You’ve noticed a weird lump, a rash that wasn’t there yesterday, or a headache that just won’t shift. Your GP surgery is shut, the NHS 111 queue feels endless, and A&E seems like overkill. So you do what millions of us are quietly doing every week. You open ChatGPT and start typing.

AI health questions

You’re not alone. OpenAI reckons more than 200 million people ask ChatGPT health and wellness questions every single week, and around 70% of those conversations happen outside clinic hours. A recent BBC report, drawing on research from Oxford University and others, paints a picture that anyone with a smartphone will recognise: AI has quietly become one of the world’s biggest “doctors”, whether the medical profession likes it or not.

The trouble is, new research suggests our digital second opinion might not be as reliable as we’d like to believe.

Why we’re turning to the bots

If you’ve ever tried to book a same-day GP appointment, you already know half the answer. Access is hard. Waiting lists are long. Private care is expensive. And even when you do get through, you’ve got roughly ten minutes to explain something you’ve been worrying about for weeks.

Compare that to an AI chatbot:

  • It answers immediately, at any hour
  • It never sighs, never rushes you, never makes you feel silly for asking
  • You can describe embarrassing symptoms without blushing
  • It pulls from a vast library of medical information in seconds
  • It’s free (or close to it)

There’s also something comforting about the way AI talks to us. It sounds calm, authoritative, and organised, which, when you’re anxious at midnight, feels genuinely reassuring. One doctor quoted in recent coverage put it well: the advice you get from AI is often better than nothing, and probably better than what your cousin on WhatsApp would offer.

So the appeal is completely understandable. The question is whether “better than nothing” is actually good enough when your health is on the line.

What the research is telling us

This is where things get uncomfortable. A study led by the Oxford Internet Institute and the Nuffield Department of Primary Care Health Sciences, the one the BBC and others have been reporting on, ran a large experiment with nearly 1,300 participants using AI chatbots to work through realistic medical scenarios written by doctors. The headline finding? People using AI didn’t make better decisions than people who simply Googled their symptoms or trusted their own judgement.

In some cases, they did worse.

A separate study published in BMJ Open went further. Researchers tested five of the biggest chatbots (ChatGPT, Gemini, Meta AI, Grok and DeepSeek) and found that roughly half the medical responses were problematic, with nearly one in five described as highly problematic. Another piece of research found that in 52% of emergency scenarios, chatbots “under-triaged”, treating genuinely dangerous situations as less serious than they actually were.

One example from the Oxford research really stuck with me. Two participants were describing the same underlying condition, a subarachnoid haemorrhage, which is a life-threatening type of stroke. One person said they’d “suddenly developed the worst headache ever”. The AI correctly told them to get to A&E immediately. The other person described a “terrible headache” with the same other symptoms. The AI suggested it was probably a migraine and recommended lying down in a dark room.

Same condition. Same symptoms. Two tiny words of difference. One answer could save a life; the other could end one.

The communication problem nobody warned us about

Here’s the bit I think is most interesting, and the one most people don’t see coming. The problem isn’t really that AI “doesn’t know” medicine. On paper, these models pass medical exams with flying colours. The problem is what happens when a real, worried human tries to talk to one.

Doctors are trained to ask the questions you didn’t know you should answer. “Does the pain radiate anywhere?” “When did it start, exactly?” “Have you had this before?” A chatbot, by contrast, mostly works with whatever you happen to type, and most of us don’t know what matters and what doesn’t.

The Oxford researchers called this a “two-way communication breakdown”. We don’t know what to tell the model, and the model doesn’t know what to ask. The result is a response that often mixes good advice with bad, wrapped in a tone of calm confidence that makes it very hard to tell which bit is which.

And there’s a sneakier risk that’s been getting more attention recently: anchoring bias. If AI tells you your chest pain is “probably anxiety”, you might describe it that way to your GP the next day, leaving out details that would have pointed them towards something more serious. The AI hasn’t just given you a wrong answer. It’s subtly reshaped the conversation you’ll have with the human who can actually help.

So… should we stop?

Honestly? Probably not, and probably we won’t anyway. The genie is well out of the bottle. AI genuinely can be useful for understanding medical jargon, thinking through questions to ask your GP, learning about a diagnosis you’ve already been given, or getting a rough sense of whether something sounds urgent. Used well, it can make you a more informed, more confident patient.

But there’s a meaningful difference between a tool that helps you prepare for a medical conversation and a tool that replaces one. Today’s chatbots are much better at the first than the second.

A few sensible ground rules, if you’re going to use AI for health questions:

  • Treat it like a starting point, not a verdict
  • Be specific and generous with detail; vague questions get vague answers
  • If the AI uses words like “urgent”, “emergency”, “severe”, or “immediately”, take that seriously, but don’t treat the absence of those words as reassurance
  • For anything sudden, severe, or getting worse quickly, skip the chatbot entirely and ring 111 or 999
  • Never let an AI’s suggestion stop you booking an appointment you were already worried enough to consider

The diagnosis…

AI is extraordinary technology, and it’s going to keep getting better at this. But right now, what the research keeps showing us is that medical knowledge and medical judgement aren’t the same thing, and a chatbot that sounds certain isn’t the same as one that’s correct.

If there’s one thing worth taking away from all this, it’s that the goal shouldn’t be choosing between AI and your doctor. It should be using AI in a way that helps you get to the right human, faster, and knowing, honestly, where its limits are.

Because a confident answer at midnight is not the same thing as a right one.

Author: Deborah Holmwood, Client Change & Transformation Partner.

Follow our LinkedIn company page to stay up to date with all our new blogs!

Leave a Reply