
About Osseus
We are building
medical superintelligence.
A system reliably better than the median clinician on a widening set of tasks, deployed under supervision, and available in places that have no specialist at all.
The goal
Why we think it is reachable.
Almost everything medicine knows is written down somewhere, and almost none of it is written down in a form a model can learn from. The trial literature is public and thin. The practice, the billions of encounters a year in which clinicians decide under uncertainty and then find out what happened, is private and mostly unrecorded. That is why medical AI has stalled at the level of an impressively well-read student: it can recite the guideline, but nobody has shown it a few million instances of deciding, with the answer attached.
The second problem is reward. Reinforcement learning carried coding and mathematics a long way because both have cheap, honest verifiers, and you cannot unit-test a differential diagnosis. So we manufacture the verifier: real patient trajectories turned into gradeable environments, with specialist-written rubrics where judgement is required and the recorded outcome as the anchor wherever the world settled who was right.
Both problems need a position inside the institutions where care happens, which is why we build the record system too. It will not arrive as a single model release. It arrives as a loop that keeps closing faster: care producing data, data producing models, models returning to care.

Who it is for
Most clinical AI has never met most patients.
Most clinical AI is trained on a narrow slice of the world: a handful of high-income countries, and mostly patients of European descent within them.
A model fitted to that slice does not generalise off it. Accuracy falls on the patients the data left out, which makes it both less fair and, simply, less good.
Better medicine for everyone needs more of the world in the training data. That is the gap our partnerships fill.
More than half of the datasets behind published clinical AI come from two countries, the United States and China. Almost every one of the most-used databases is from a high-income country.
PLOS Digital Health, 2022The imbalance is in who is enrolled as much as where. Of the participants behind genome-wide association studies, 86.5% are of European descent, while those recorded as African, not counting African American or Afro-Caribbean samples, are 0.47%.
Cell Genomics, 2024Of 106,950 publicly available skin cancer images, ethnicity was recorded for 1.3% of them. Among those, not one patient was of African, Afro-Caribbean or South Asian background.
The Lancet Digital Health, 2022The models inherit it. Chest X-ray classifiers trained on three of the largest public datasets consistently underdiagnosed Black, Hispanic and female patients, telling them they were healthy when they were not, and were worst of all at the intersections.
Nature Medicine, 2021It costs accuracy, not only equity. Risk scores built from European cohorts keep 51% of their predictive power when applied to South Asian patients, 47% for East Asian and 39% for African.
Human Genomics, 2024
Principles
Four commitments we can be held to.
- Built for the patients it serves
- We train on data from the populations a model will treat, and report performance by group rather than in aggregate.
- Consent is the strategy
- The only durable supply of clinical data is one patients and institutions would renew if asked again tomorrow.
- Outcome over eloquence
- A recommendation that reads beautifully and was wrong is worse than useless.
- Autonomy is earned slowly
- One task, one site, in public, with a clinician holding the override.
Careers
We are hiring people who would do this anyway.
- Research engineers
- Post-training, RL infrastructure and evaluation. Your work reaches a real patient pathway within months.
- Clinician-builders
- Practising specialists who want to encode their judgement into environments rather than advise from a board seat. Part-time works.
- Deployment engineers
- Hospital networks, on-premise inference, and integration with systems designed in 2004.
- Product engineers
- The EHR and the agent layer on top of it. If you have opinions about why clinical software is so bad, we want to hear them.
