Decades of male-focused medical research could bias healthcare AI

MarcoPace68/Shutterstock

Many people will learn CPR using a flat-chested manikin. A 2024 study of 20 models of CPR manikins sold worldwide found that three quarters were described as male or had no sex specified. Of the 20, only one offered a breast overlay.

The manikins reflect a wider tendency in medical teaching and research to treat the male body as standard. The terms “male and female” and “men and women” in this article reflect the sources, which often fail to distinguish sex from gender or say whether gender-diverse people were included.

The lack of female representation can have consequences. In a US study of 19,331 out-of-hospital cardiac arrests, 39% of women who collapsed in public received bystander CPR, compared with 45% of men.

For CPR, even a short delay can affect survival. Another US study found that people who received bystander CPR four to five minutes after a witnessed cardiac arrest had 27% lower odds of surviving to hospital discharge than those who received it within one minute. In England, fewer than one in 12 patients for whom ambulance services attempted resuscitation survived 30 days, according to figures reported in 2024.

If bystanders hesitate to perform CPR on women because of the presence of breasts, training on more representative manikins could help. Historical gaps in representation now have another consequence. Healthcare increasingly uses artificial intelligence trained on medical records and research data produced when male bodies were often treated as the norm.

When old data trains new systems

Healthcare AI learns associations from health data. But medical records reflect clinical decisions as well as biology. If women and men with similar symptoms have historically been assessed differently, an algorithm may learn that pattern without knowing why it exists.

For example, an experimental study of medical students and residents found that when coronary heart disease symptoms were presented alongside psychological stress, women received fewer coronary heart disease diagnoses and cardiology referrals than men, and their symptoms were more likely to be interpreted as psychogenic. An AI system trained on records shaped by such decisions could treat that pattern as though it reflected disease itself.

Reviews of AI in medicine warn that unbalanced datasets can produce uneven performance. It can also be difficult to assess performance: a 2024 review of 692 AI-enabled medical devices approved by the US Food and Drug Administration found that demographic information and details of performance studies were often missing from public documents.

Understanding that risk requires looking at how the medical evidence AI may draw on was created.

How men became the medical default

Females are consistently underrepresented in medical studies. Hormonal changes across the menstrual cycle have historically been viewed as a source of unwanted variation. After the thalidomide tragedy of the late 1950s and early 1960s, the US Food and Drug Administration in 1977 recommended excluding women who could become pregnant from early drug trials. The policy was reversed in 1993 amid concern that it was limiting knowledge about how medicines affected women.

Women of childbearing potential were routinely excluded from industry-sponsored trials. Male animals have also been favoured in laboratory research, despite evidence that females are no more biologically variable, and many studies fail to report the sex of cells used in experiments.

These gaps can affect treatment. Women clear the sleeping drug zolpidem more slowly than men, and in 2013 regulators recommended halving the starting dose for women. A 2020 analysis found that women experienced adverse drug reactions nearly twice as often as men, with differences in how drugs move through the body explaining many of them.

Representation has improved. The 2016 Sex and Gender Equity in Research (SAGER) guidelines encourage researchers to consider and report sex and gender. But equal participation does not guarantee equal understanding. A 2026 review of 574 papers found that 61% included both sexes, yet only 44% of those analysed results by sex.

Biological sex and gender can both affect health, and separating their effects is difficult. Sex-related biology may influence drug metabolism, while gendered experiences may affect access to healthcare and how symptoms are interpreted. Yet many health databases collapse both into one binary category, limiting what researchers can conclude.

Biology or bias?

Heart-attack symptoms show why precision is needed. An analysis of more than one million people with acute coronary syndromes found that 74% of women and 79% of men experienced chest pain, but women had higher odds of reporting other symptoms including neck or jaw pain, fatigue and shortness of breath.




Read more:
Women are at a higher risk of dying from heart disease − in part because doctors don’t take major sex and gender differences into account


Stroke care presents a similar problem. An Australian study of more than 202,000 people admitted with stroke found that among those under 70 arriving by ambulance, women were less likely than men to be assessed as having a stroke by paramedics and less likely to be managed under the pre-hospital stroke protocol.

Finding a difference is only the starting point. A study of 6.9 million people in Denmark found that women were older at their first hospital diagnosis for most conditions examined. That does not necessarily mean doctors took longer to diagnose them: disease onset and use of healthcare may also vary.

Medical research is becoming more representative, but much of today’s evidence was collected when male bodies were more often treated as standard. Careful design could help AI identify patterns that older research missed. As tomorrow’s healthcare is built from yesterday’s records, researchers must ask whether an apparent difference reflects biology or the way patients encountered healthcare. Otherwise old assumptions could become embedded in new technology.

The Conversation

Suzannah Williams research group is funded by Children with Cancer, Rosetrees, Revive and Restore, Oxford Medical Research Fund, and John Fell OUP.

Brittany Vining does not work for, consult, own shares in or receive funding from any company or organisation that would benefit from this article, and has disclosed no relevant affiliations beyond their academic appointment.

Scroll to Top