Skip to content
Back

Science

One molecule in. One honest number out.

The engine reads a molecule’s structure and returns one call: will this hurt a liver, and how sure are we. This page is what goes into that call, how it is tested, and what we will not claim.

The question

Drug-induced liver injury is the single biggest safety reason a candidate dies, and it is usually found late. It is the first endpoint the engine learns. The others follow, in the order the data and the toxicologists allow.

  1. DILILiver injuryDrug-induced liver injury: the first endpoint, and the one the benchmark result is about.First
  2. hERGCardiotoxicityBlock of the hERG channel, the classic route to QT prolongation.Next
  3. AmesMutagenicityThe bacterial reverse-mutation test, the standard genotoxicity screen.Next
  4. CytoCytotoxicityGeneral toxicity to cells, the broadest early signal.Next
  5. CYPMetabolismInteraction with cytochrome P450 enzymes, where many clinical surprises start.Later

The input

The engine takes the molecule as drawn: atoms, bonds, stereochemistry. Nothing has to be synthesised, nothing has to be dosed. That is the whole point, and it is also the limit.

What structure cannot see

  • Dose and exposure.
  • Formulation and route.
  • Most of metabolism, which is why CYP is its own endpoint.

So every prediction is a prior, not a verdict: the best call available before a single milligram exists.

The data

Safety data is public, which is what makes the first result provable on compute alone. It is also messy: different assays, different thresholds, duplicates, salts, label noise. Most of the work is here.

  • Tox21
  • ToxCast
  • ChEMBL
  • DILIrank
  • Therapeutics Data Commons
  • MoleculeNet
  1. 01StandardiseOne canonical form per molecule: salts stripped, tautomers and charges resolved, stereochemistry kept where the assay kept it.
  2. 02DeduplicateSame molecule, same assay, one record. Disagreements stay visible instead of being averaged away.
  3. 03Judge the labelA toxicologist decides what counts as positive for each endpoint, assay by assay.
  4. 04Split without leakageMurcko scaffold and time splits, so the test set holds chemistry the model never saw.
  5. 05Keep provenanceEvery record carries its source and version. Every result re-runs in one command.

The benchmark

The test set is locked before any modeling starts and never touched again. It is public, so anyone can run it. A win means beating the best freely available tool on that set, on data the model never saw, with the splits and the metric published alongside.

  1. 01Locked firstChosen and frozen before any modeling. Never tuned against.
  2. 02PublicThe set, the splits and the metric are published. Anyone can rerun it.
  3. 03Beats the free baselineThe bar is the best freely available predictor, on chemistry the model never saw.

The model

Architecture is the least interesting part. Several families are in play, kept or dropped by what survives the benchmark. All of them are grounded in physics and real assay data.

  1. 01Graph neural networksThe molecule as a graph of atoms and bonds.
  2. 02TransformersOn the graph, or on the SMILES string.
  3. 03Pretrained encodersRepresentations learned on millions of unlabeled molecules, tuned on small assay sets.
  4. 04Multi-taskEndpoints learned together, so sparse assays borrow strength from each other.

Uncertainty

A toxicity prediction without a confidence is a guess with a decimal point. Each call comes with a calibrated probability, and an honest “we do not know” when the molecule sits outside what the data covers.

  1. 01CalibratedA predicted 80% is right about 80% of the time, checked on held-out data.
  2. 02ConformalPrediction sets with a guaranteed coverage rate, not a bare point estimate.
  3. 03Out of distributionChemistry far from the training data is flagged, not scored with false precision.

The scientists

Endpoints are defined and labels judged by people who know the biology, with the toxicology and pharmacology faculty of the Medical University of Warsaw. Metrics are chosen to match real safety decisions, not leaderboard habits. Once the engine is proven, synthesis and assays through partners feed measured results back into it.

Results

No result is published yet, and nothing on this page claims one. When the benchmark run is done it goes here, with the split, the metric and the baseline beside it.

DILI, held-out scaffold splitPending
All news