Dr. KI: The Double-Edged Diagnosis

One in two people uses AI for medical advice, but according to *Nature Medicine*, the accuracy rate among laypeople is only 35% (compared to 95% for professionals). The diagnostic value therefore depends primarily on the quality of the question asked.

Photo: Generated using AI
Hanna Sachse
April 30, 2026

AI in Medical Counseling: Key Findings

Discrepancy in accuracy: While about half of insured individuals use AI for health-related questions, laypeople achieve an accuracy rate of only 35% when making preliminary diagnoses. Medical professionals, on the other hand, achieve a 95% accuracy rate through optimized queries (Sources: Bitkom / Nature Medicine).

Limitations in patient triage: Recent studies show that general models offer no diagnostic added value compared to conventional screening methods. Accuracy remains stagnant at around 74%, with 70% of errors attributable to unfounded recommendations to see a doctor. Furthermore, newer models (such as GPT-5) provide inconsistent answers to identical queries in 42% of cases (Sources: University of Oxford / TU Berlin 2026).

Specialization and the Role of the Pharmaceutical Industry: Specialized bots based on verified data (e.g., Uro-Bert, Lupus-GPT) are considered an alternative. For the pharmaceutical industry, such systems offer the potential to convey information in an understandable way and ensure adherence; however, diagnostic liability remains legally excluded (Sources: BMG / Industry Analysis).

Nearly one in two patients now seeks medical advice from an AI. But while professionals get spot-on medical results with precise queries, for many laypeople the journey leads into the unknown. In the end, it’s often not the computational power of the code that matters, but the quality of our questions.

Although generative AI models like ChatGPT possess an enormous amount of medical knowledge, laypeople hardly benefit from it in their daily lives. According to a study published in the journal *Nature Medicine*, ordinary users receive the correct preliminary diagnosis in only 35% of cases. By comparison, experts achieve a 95% accuracy rate when using precise input.

According to experts, the problem lies less in the AI’s computing power and more in the interaction: Laypeople often omit important information in their prompts (search queries) or unconsciously steer the model in a direction that leads to an incorrect diagnosis. In addition, users tend to anthropomorphize the AI, which reduces their critical distance from the results.

Despite this error rate, demand remains extremely high, as chatbots are available around the clock and allow for anonymous counseling on sensitive topics. Experts are therefore urgently calling for greater AI literacy and the use of specialized bots to prevent dangerous misdiagnoses.

The Trend Toward Digital Second Opinions

The reluctance to have medical symptoms assessed by AI is rapidly declining. An estimated half of those with public health insurance in Germany have already used ChatGPT or similar models to ask health-related questions (“Digital Health 2025” by Bitkom Research). The main reasons for this are the constant availability (24/7) and the difficulty in securing appointments with primary care physicians, especially in rural areas. AI models such as ChatGPT, Claude, or Gemini also offer anonymity, which is seen as an advantage, particularly when it comes to sensitive topics such as mental health, addiction, or sexual health.

Data Set Factor

A key risk associated with general-purpose chatbots is their data source. Because they draw on the entire Internet, reliable medical knowledge is mixed with inaccurate information. This can lead to dangerous self-diagnoses that delay necessary visits to the doctor.

Experts see the solution in specialized systems based on verified expert data. Examples such as “Uro-Bert” (urology) or “Lupus-GPT” demonstrate how AI can serve as a precise interface for expert knowledge. However, data protection remains an unresolved issue: To receive an accurate, personalized response, users must disclose highly sensitive data. This presents a dilemma between medical accuracy and the protection of privacy.

The Role of the Pharmaceutical Industry

For the pharmaceutical industry, integrating AI into patient counseling can open up new possibilities along the patient journey. Specialized bots could be used to explain complex package inserts in an understandable way or to support the management of side effects in chronic conditions. AI thus serves as a bridge between scientific evidence—which is often difficult to understand—and patients’ need for information. Nevertheless, the issue of liability remains central: As long as AI models are not permitted to make binding diagnoses, their use serves primarily to provide information and promote adherence, but not to replace medical expertise.

If you'd like to know more:

Specialized Medical Bots

  • Uro-Bert: Developed by Uro-GmbH Nordrhein (a network of private urologists), this chatbot offers anonymous counseling on sensitive topics such as erectile dysfunction, incontinence, and prostate cancer. It serves as a low-threshold information resource that, when necessary, refers users to specialized clinics or emergency care.
  • Sucht-GPT: A project funded by the Federal Ministry of Health (BMG) that provides anonymous information on addiction disorders (e.g., gambling or drug addiction) to those affected and their families. The bot helps users identify symptoms and connects them with professional support services.
  • Lupus-GPT: This specialized bot was developed in collaboration with lupus patients and medical professionals. Studies show that specialized AI models can often answer questions about systemic lupus erythematosus just as accurately and empathetically as medical specialists.
  • Cancer Information Service (DKFZ): The service is experimenting with chatbot solutions to present complex information on cancer prevention and early detection in an accessible way. However, experts expressly warn against using general-purpose AI systems such as ChatGPT for treatment decisions, as their databases are often outdated or do not disclose their sources.

‍

Oxford Study: Why Dr. ChatGPT Can't Replace a Visit to the Doctor

A new study by the Oxford Internet Institute and the Nuffield Department of Primary Care Health Sciences at the University of Oxford reveals a significant discrepancy between the theoretical capabilities of AI models (LLMs) and their practical utility for patients. While these models now achieve excellent results in standardized medical knowledge tests, they pose a significant risk to real-world users seeking help with symptoms.

The key findings of the study:

  • No added diagnostic value: Participants who consulted an AI did not make better medical decisions than those who used traditional methods such as online searches or their own judgment.
  • Dangerous Misdiagnoses: The study warns that chatbots can make incorrect diagnoses and often fail to recognize when immediate emergency medical care is needed.
  • Poor Communication: A two-way communication problem has been identified: Users often do not know what information the AI needs to provide accurate advice, while the AI's responses often contain a mix of good and bad recommendations.
  • Inconsistency: The AI often provided completely different answers when there were minor variations in the wording of the question.

‍

Background: ChatGPT in Health Counseling

A recent study by the Technical University of Berlin (published in 2026 in *Communications Medicine*) examined the reliability of ChatGPT in providing an initial assessment of health complaints. The team led by Dr. Marvin Kopka analyzed 22 model versions (up to GPT-5) using 45 real-world patient case studies.

Key findings:

  • Systematic overcaution (“conservative triage”): The AI tends to classify symptoms as more urgent than they actually are from a medical standpoint. While the need for treatment is usually correctly identified, the biggest weakness lies in harmless cases: 70% of all errors occurred because the AI recommended a medical evaluation even though self-care would have been sufficient.
  • Stagnating Accuracy: Since the introduction of GPT-4, accuracy has plateaued at around 74%. According to the researchers, better results on medical knowledge tests (state exams, etc.) do not automatically translate to better practical patient care.
  • Lack of Consistency: Identical queries often result in different recommendations. This is particularly noticeable with GPT-5, where responses were inconsistent in 42% of cases when the same scenario was entered multiple times.
  • Limited value for healthcare delivery: Since the models almost always recommend a doctor’s visit “just to be safe,” they lack a guiding effect. Instead of relieving pressure on the healthcare system, such recommendations could actually increase the number of unnecessary doctor’s visits.

The researchers' conclusion: The standard version of ChatGPT is not currently suitable as a standalone tool for patient management. The technology's potential lies more in its integration into quality-assured symptom-checker apps, where medical oversight is ensured behind the scenes.

(Study: Kopka, M., He, L., & Feufel, M.A., “Evaluating the accuracy of ChatGPT model versions for giving care-seeking advice.” *Commun Medicine* (2026). https://www.nature.com/articles/s43856-026-01466-0)

‍

More Articles
Deepfake Medical Professionals: “You want real truth in advertising”
July 30, 2026
Agent-Based AI in Medicine: Potential, Practice, and Guidelines
July 23, 2026
Pharmaceutical Communication: The Celebrity Factor vs. Patient Trust
July 9, 2026

Kicking Off the Industry Dialogue: The New Forward Pharma Podcast

Kicking Off the Industry Dialogue: The New Forward Pharma Podcast
August 10, 2026
Listen now
Listen to the podcast
"FORWARD PHARMA"
analyzes the key trends in the healthcare and pharmaceutical industries. We provide context for current topics and speak with experts.
Click here for the
FORWARD PHARMA
Podcast