Artificial intelligence has moved into medicine fast — drafting clinical notes, suggesting diagnoses, triaging symptoms. But according to Flinders University, whose findings were surfaced this week in a Google News roundup on large language models, the latest generation of AI systems still reproduces racial and gender stereotypes when applied to medical scenarios.
The headline claim is blunt: newer models, despite the safety work and scale increases that separate them from earlier releases, have not shed the biased patterns researchers have flagged in clinical AI for years. Flinders University frames this as an ongoing problem rather than a solved one.
Detail beyond that is limited in the material available here. The item circulating is a summary of the university's announcement, and it does not spell out which models were tested, what clinical scenarios were used, or how large the measured gaps were. Those specifics matter, and readers should treat the finding as a signal to check the underlying study rather than a final scorecard on any particular product.
Why this is worth paying attention to anyway: bias in a medical AI system is not an abstract fairness complaint. If a model is more likely to downplay pain, suggest a different workup, or reach for a different diagnosis depending on a patient's race or gender, the output is simply wrong medicine — and it arrives wrapped in the authority of a computer, which makes it harder for a clinician or patient to push back.
The finding also complicates a common assumption in the industry: that bias is a growing pain newer, bigger models will outgrow. Flinders University's work suggests that assumption has not held.
As hospitals and pharmaceutical companies weave these tools into real decisions about real patients, evidence that the newest models still encode old stereotypes is a reason to keep humans — and independent testing — firmly in the loop.