The Female Health Data Gap Is an Engineering Problem
Medicine treats the female body as an edge case because the datasets do. Missing data is a systems failure, and systems failures can be fixed.
In 1977, the US Food and Drug Administration published a guideline recommending that women of childbearing potential be excluded from early drug trials. It defined that category generously: any premenopausal woman capable of becoming pregnant. Single women. Women on the pill. Women whose husbands had had a vasectomy.
It also listed the exceptions, the cases where a woman could be given an experimental drug before anyone knew what it did to a foetus. One of them was women who had been institutionalised long enough to establish that they were not pregnant.
That line makes the logic legible. The policy was not really about protecting women. It was about protecting a hypothetical foetus, and a confined woman was the one place that protection could be guaranteed without anybody having to ask her.
Most people know roughly that story. Women were shut out of medical research, and it took until the nineties to fix. It is mostly true. It is also, I think, the wrong thing to still be angry about, because the problem moved.
The part that got fixed
First, a correction I had to make to my own thinking. The 1977 guideline was not a blanket ban. It covered Phase 1 and early Phase 2 trials, the first small studies that check whether a drug is safe in humans at all, and for later trials it only recommended exclusion until animal reproduction studies were done. What happened next is a very recognisable engineering failure: sponsors read the restriction, could not be bothered to work out its boundaries, and applied it to everything. The FDA said as much in 1993. The rule was narrow and the implementation was broad, which is what always happens when the safe default is to exclude.
The fix came in two pieces. In June 1993 the NIH Revitalization Act made inclusion a matter of federal law, and in July 1993 the FDA revised its guideline. Then, mostly, it worked. In the trials supporting the FDA's 53 novel drug approvals in 2020, 56% of participants were women.
So if your argument is "women are excluded from clinical trials," that number is a problem for you. Aggregate enrolment is not the scandal, and has not been for a while.
The part that did not
The 1993 statute did not just require that women be enrolled. It required that trials be "designed and carried out in a manner sufficient to provide for a valid analysis of whether the variables being studied in the trial affect women... differently than other subjects." It even specified that cost was not a permissible reason to skip it.
Enrol them, and analyse the results separately. Two requirements. We did the first one.
In 2020 a team re-ran a survey of sex inclusion across nine biological disciplines. Studies including both sexes had become much more common since 2009, but among those that included both, the share actually analysing the data by sex had fallen, to 42%. A 2026 evaluation of publications from NIH R01 grants found 44% ran a sex-based analysis. Two independent measurements, six years apart, landing on 42% and 44%.
And then the number I cannot get out of my head. Somebody went through 224 heart failure randomised controlled trials published in high-impact journals between 2000 and 2020: 228,801 participants, 28% of them women. Not one of those 224 trials reported its adverse events, the side effects and harms people actually suffered, broken down by sex. Zero.
That is the whole ballgame. Side effects are where a drug that behaves differently in women shows up. Every one of those trials recorded who had which side effect, and knew the sex of every participant. The join was never performed.
I write software for a living, so this is where it stops feeling like a medical story to me. We have a pipeline. It ingests correctly: the field is populated, right type, present on every record. Then at the aggregation step somebody takes the mean across the population, writes it to the output, and the column is silently dropped. No error. Nothing fails. The report is generated, peer reviewed, published, and used to set the dose of a drug given to millions of people.
Nobody in that chain did anything wrong, exactly. Every step passed its own checks. The information was destroyed at aggregation, the one place nobody thinks to look, because aggregation feels like summarising rather than deleting. The 1993 law was a schema requirement with no validation and no consequence for violating it. We have all shipped one of those.
Here is what falls through. Across 86 drugs with documented sex differences in how the body processes them, 76 showed higher exposure in women at the same dose. Same pill, more drug in the bloodstream.
You will often see this summarised as "women experience adverse drug reactions twice as often as men." I could not stand that number up. The literature says more like 1.4 to 1.7 times for recorded reactions, and one large Dutch population study found identical hospitalisation rates in both sexes. The inflation is what lets people dismiss the whole thing.
Same with the most famous example in the field. In 2013 the FDA halved the recommended dose of zolpidem for women, its first ever sex-specific dosing recommendation, after finding that about 15% of women versus 3% of men still had impairing blood levels eight hours after a 10mg dose. It gets told as a clean parable. It is not: a 2019 analysis argued the reduction was not supported by the evidence and risked underdosing women into untreated insomnia.
I find the contested version more damning. A blockbuster sleeping pill had been on the market for a quarter of a century, and we were still having a first-of-its-kind argument about dosing women differently, on evidence thin enough to go either way. The scandal is not whether the FDA got it right. It is that nobody had the data to settle it.
The cycle argument, done properly
The version of this argument that gets passed around goes: men run on a 24-hour cycle, women run on a 28-day cycle, and medicine only accounts for the first. I believed it. Then I checked, and most of it does not survive.
The 28 days is not real in the way people think. In an analysis of 612,613 cycles, the mean length was 29.3 days and only 13% of cycles were exactly 28 days long. The implied contrast is wrong too: women have circadian rhythms as well, measurably different from men's, and men have hormone pulses every couple of hours, a daily testosterone swing, a seasonal swing, a decades-long decline. The rhythms are nested, not competing.
But there is a version underneath the tidy one that holds up. It is about how big the swing is, and what medicine chose to institutionalise.
A man's testosterone falls by roughly 20 to 40% between morning and late afternoon, and clinical guidelines respond to that by telling you to draw the blood before 10am. A woman's ovarian hormones swing by an order of magnitude more over a few weeks, and there is no equivalent instruction anywhere.
And when studies do try to record where in her cycle a participant was, they usually count days from the last period. Tested against hormone measurement, counting forward 10 to 14 days from the start of menstruation misclassifies ovulation 82% of the time. The field imported the 28-day model into its measurement code. The myth is not just something people believe. It is a bug that shipped, and it corrupts the data of the studies conscientious enough to try.
The studies that asked the wrong question
Everything above is about drugs, where at least somebody is counting. For conditions that only happen to women, the problem is not a skipped analysis. It is that the study asked the wrong question in the first place.
Take polycystic ovary syndrome. The Cochrane review of lifestyle intervention in PCOS pools somewhere between nine and twelve small trials, a few hundred women in total, all rated low quality. It found a change of 1.68 kilograms. And the line that stopped me: no included study reported live birth, miscarriage, or menstrual regularity.
Those are the outcomes women with PCOS actually walk into the appointment about. In a survey of 1,385 women with PCOS, a third waited more than two years for a diagnosis and nearly half saw three or more health professionals to get one. What waits at the end of that is a literature that never measured the thing they came in about. Not neglect at the analysis stage. A gap designed into the protocol.
Menopause is the same failure in a different costume. The SWAN study followed women through the menopause transition with repeated body scans. Fat around the middle went from rising 1.2% a year to rising 5.5% a year, a more than fourfold acceleration. Body weight over the same period showed no acceleration whatsoever: the test for a change in the weight trajectory came back at p equals 0.98. At exactly the moment women most want an answer, the one measurement everybody relies on is the one that shows nothing. The scales are not lying. They are just not sensitive to what is happening.
The gap that measures itself
There is a statistic you see everywhere on this subject: only 1% of healthcare research funding goes to female-specific conditions outside cancer. The Gates Foundation used it in 2025 when announcing $2.5 billion for women's health. I went to find the source. It is the title of a chart, Exhibit 2 in a McKinsey article from 2022, and what the chart shows is the share of the commercial drug development pipeline, counted in assets rather than in pounds. The article's own text says under 2%. Somewhere between the chart and the press release, a pipeline number became a funding number.
You do not need the drifted version. The actual budgets are public. In 2024 the US National Institutes of Health spent $28 million on endometriosis and $9 million on polycystic ovary syndrome. It spent $1,151 million on diabetes. Adenomyosis, closely related to endometriosis and affecting an enormous number of women, has received two NIH grants in total. My favourite entry in this genre: a trial found that sildenafil relieved period pain, and the research stopped for lack of funding. Sildenafil is Viagra, which did $400 million of US sales in its first three months on the market.
Upstream of the money, though, there is a schema problem. The Global Burden of Disease dataset is what the world uses to decide which conditions deserve attention. It says 1 to 2% of women have endometriosis. The World Health Organisation says 10%. Nobody can resolve an eightfold disagreement about how many women have a disease, because nobody has properly counted. And menopause is not in that dataset as a category at all. It is folded into "other gynaecological diseases." Every woman lives through it, and in the file that governs global health priorities there is no field for it.
So when analysts rank conditions by measured burden to decide where the money goes, two of the largest women's health conditions turn up already shrunk by an order of magnitude, or absent entirely. Underfunding produces bad data, and the bad data justifies the underfunding. A feedback loop with no error signal, which is the worst kind.
You can watch it play out in diagnosis times. One charity's surveys put the wait for an endometriosis diagnosis in the UK at 8 years in 2020 and 9 years 4 months in 2025, and around 11 years for women from ethnically diverse backgrounds. In the most recent survey, 83% said a clinician had told them they were making a fuss about nothing, up from 69% two years earlier. Not a system slowly improving. A system going backwards, fastest for the women already waiting longest.
What I actually think should happen
Almost none of this requires new science. It requires the analysis step to be non-optional: results and side effects broken down by sex as a submission requirement at journals and regulators, the way a missing conflict-of-interest statement gets a paper bounced today. Not a recommendation. A field that fails validation.
Where a condition mostly affects women, the people who have it should get a say in what the trial measures. If everyone in the waiting room is there about periods and fertility, a literature that reports neither is misaimed rather than underfunded, and no amount of money fixes an outcome nobody chose to record.
And where cycle phase matters, measure it rather than count it, and where it was not measured, say so. "We did not record this" is a perfectly respectable sentence, and far more useful than silence, because it makes the gap visible to the next person.
One last thing. While researching this I went looking for a number: what proportion of drug trials record where in her cycle a female participant was. There are audits like that for sport science and for preclinical biology. For clinical drug trials, as far as I can find, there is none. Nobody has counted.
I was annoyed about that for a while before realising it is the cleanest possible statement of the problem. The measurement of the gap has the same shape as the gap. You cannot report on what was never recorded, and the absence does not announce itself. It just quietly comes out in the average.