Skip to content
A wide marble colonnade opening onto a misty river valley at dawn

About Osseus

We are building
medical superintelligence.

A system reliably better than the median clinician on a widening set of tasks, deployed under supervision, and available in places that have no specialist at all.

The goal

Why we think it is reachable.

Almost everything medicine knows is written down somewhere, and almost none of it is written down in a form a model can learn from. The trial literature is public and thin. The practice, the billions of encounters a year in which clinicians decide under uncertainty and then find out what happened, is private and mostly unrecorded. That is why medical AI has stalled at the level of an impressively well-read student: it can recite the guideline, but nobody has shown it a few million instances of deciding, with the answer attached.

The second problem is reward. Reinforcement learning carried coding and mathematics a long way because both have cheap, honest verifiers, and you cannot unit-test a differential diagnosis. So we manufacture the verifier: real patient trajectories turned into gradeable environments, with specialist-written rubrics where judgement is required and the recorded outcome as the anchor wherever the world settled who was right.

Both problems need a position inside the institutions where care happens, which is why we build the record system too. It will not arrive as a single model release. It arrives as a loop that keeps closing faster: care producing data, data producing models, models returning to care.

A misty river valley at dawn seen from a marble terrace

Who it is for

Most clinical AI has never met most patients.

Most clinical AI is trained on a narrow slice of the world: a handful of high-income countries, and mostly patients of European descent within them.

A model fitted to that slice does not generalise off it. Accuracy falls on the patients the data left out, which makes it both less fair and, simply, less good.

Better medicine for everyone needs more of the world in the training data. That is the gap our partnerships fill.

  • More than half of the datasets behind published clinical AI come from two countries, the United States and China. Almost every one of the most-used databases is from a high-income country.

    PLOS Digital Health, 2022
  • The imbalance is in who is enrolled as much as where. Of the participants behind genome-wide association studies, 86.5% are of European descent, while those recorded as African, not counting African American or Afro-Caribbean samples, are 0.47%.

    Cell Genomics, 2024
  • Of 106,950 publicly available skin cancer images, ethnicity was recorded for 1.3% of them. Among those, not one patient was of African, Afro-Caribbean or South Asian background.

    The Lancet Digital Health, 2022
  • The models inherit it. Chest X-ray classifiers trained on three of the largest public datasets consistently underdiagnosed Black, Hispanic and female patients, telling them they were healthy when they were not, and were worst of all at the intersections.

    Nature Medicine, 2021
  • It costs accuracy, not only equity. Risk scores built from European cohorts keep 51% of their predictive power when applied to South Asian patients, 47% for East Asian and 39% for African.

    Human Genomics, 2024

Principles

Four commitments we can be held to.

Built for the patients it serves
We train on data from the populations a model will treat, and report performance by group rather than in aggregate.
Consent is the strategy
The only durable supply of clinical data is one patients and institutions would renew if asked again tomorrow.
Outcome over eloquence
A recommendation that reads beautifully and was wrong is worse than useless.
Autonomy is earned slowly
One task, one site, in public, with a clinician holding the override.

Careers

We are hiring people who would do this anyway.

Founding team: We chose to build Osseus over a health AI PhD at the University of Oxford and ML engineering at Apple. Between us we have published AI safety research at Imperial College London, health AI research at Harvard-HKU, and built state-of-the-art agentic systems at Apple.
Research engineers
Post-training, RL infrastructure and evaluation. Your work reaches a real patient pathway within months.
Clinician-builders
Practising specialists who want to encode their judgement into environments rather than advise from a board seat. Part-time works.
Deployment engineers
Hospital networks, on-premise inference, and integration with systems designed in 2004.
Product engineers
The EHR and the agent layer on top of it. If you have opinions about why clinical software is so bad, we want to hear them.
An ancient aqueduct striding across a green valley at dawn

Begin

If this sounds like your life's work, say so.