Careers
Four teams, one proof.
Everyone we hire in year one converges on a single result: a drug-toxicity model proven on a public benchmark, on data it never saw. The model, the data that fuels it, and the science that proves it means something.
7 roles open
Modelingthe engine
Mission
Turn ideas into benchmarked results fast, reproducibly, and without breakage.
What you will do
- Implement model architectures and training loops.
- Build reproducible training/eval pipelines; run and track experiments.
- Optimize GPU training; keep the benchmark harness clean.
Must have
- Python + PyTorch; solid software engineering (git, tests, clean modular code).
- Reproducible ML + experiment tracking (W&B / MLflow).
- Comfortable with RDKit basics and a GNN library (PyG / DGL).
Experience
2+ yrs building real ML systems (not just notebooks). Some exposure to molecules/graphs, or clear evidence you ramp fast on a new domain.
Nice to have
Cheminformatics; distributed / mixed-precision GPU training; MLOps tooling.
Mission
Make the model's confidence trustworthy, so we know which predictions to believe, and investors trust the numbers.
Must have
- Uncertainty quantification, conformal prediction, deep ensembles, calibration (temperature / Platt scaling).
- Statistics; PyTorch; active-learning fundamentals.
Experience
PhD or equivalent in ML / statistics; publications in UQ or calibration. Chemistry exposure is a plus, not required.
Data & Cheminformaticsthe fuel
Mission
Turn messy public toxicity data into clean, honest, benchmark-ready datasets, the single biggest lever on whether the model wins.
What you will do
- Aggregate & standardize datasets (Tox21, ToxCast, ChEMBL, DILIrank, hERG, Ames, Therapeutics Data Commons).
- SMILES standardization, salt-stripping, tautomer/charge handling, de-duplication.
- Design leakage-free splits (Murcko scaffold, time) so results are real.
- Featurization + matched-molecular-pair analysis; document data provenance and QC.
Must have
- RDKit (expert) + Python / pandas.
- Deep understanding of molecular-data pitfalls, activity cliffs, duplicates, assay variability, label noise.
- Familiarity with public ADMET/tox datasets and benchmark suites (TDC, MoleculeNet).
Experience
3+ yrs cheminformatics / computational-chemistry data work. Has personally built a clean dataset others trusted.
Nice to have
QSAR modeling; medicinal-chemistry intuition; some ML; toxicology familiarity.
Mission
Make every dataset and every result reproducible, versioned, and traceable to its source.
What you will do
- Build ingestion / ETL pipelines and the experiment data store.
- Schema design (Postgres); data versioning (DVC / LakeFS); object storage (S3/GCS).
- Provenance & lineage on every record; automate dataset builds.
Must have
- Python + SQL / Postgres.
- Pipeline tooling (Airflow / Prefect / Dagster); DVC or LakeFS.
- Cloud storage; reproducible, tested data builds.
Experience
3+ yrs data engineering. Bonus for scientific / research data.
Nice to have
MLOps overlap; comfort with molecular data formats.
Science & Validationthe credibility
Mission
Make sure we predict the right toxicity, measured the right way, and that a benchmark win is scientifically and clinically meaningful, not a data artifact.
What you will do
- Select & define endpoints, hERG (cardiotoxicity), DILI (liver), Ames (mutagenicity), cytotoxicity, CYP metabolism.
- Source and vet ground-truth data; judge label quality biologically.
- Define metrics that map to real safety decisions and regulatory relevance.
- Interpret model outputs, sanity-check against known chemistry, and represent the science to investors/partners.
Must have
- Toxicology / pharmacology: ADMET, tox endpoints, in-vitro/in-vivo assays.
- QSAR & read-across concepts; judging data quality biologically.
- Comfort collaborating closely with an ML team.
Experience
PhD or industry background in computational toxicology / pharmacology / safety pharmacology; exposure to in-silico tox strongly preferred.
Nice to have
Coding (Python / RDKit); regulatory tox (ICH, OECD); DILI/hERG domain depth; translational or pharma experience.
Mission
Credibility + guidance for near-zero cost, validate endpoints and benchmark meaningfulness, open doors, and lend names that answer "do these software people understand biology?"
What you will do
- Faculty / clinicians in toxicology, pharmacology, medicinal chemistry.
- Hepatology & cardiology for DILI / hERG relevance.
- Bonus: anyone with translational or pharma-industry ties.
Experience
A few hours a month. Small advisory equity (0.1–0.5%, 1–2 yr vesting) and/or honorarium. We aim for 2–4 engaged advisors, not a wall of logos.
Enablingkeeps all three teams fast
Mission
Keep training fast, reproducible, and cheap; make results one-command re-runnable.
What you will do
- Environments/containers (Docker); CI/CD (GitHub Actions).
- GPU/cloud orchestration and cost control; model registry; serving (FastAPI / BentoML).
- Monitoring & reproducibility tooling across teams.
Must have
- Docker, cloud (AWS/GCP), CI/CD, Python.
- MLOps tooling (MLflow / model registries); GPU infra.
Experience
3+ yrs infra / DevOps / MLOps.
Nice to have
Terraform / IaC; experience keeping GPU bills down.
Don't see an exact match? We still want to hear from strong scientists and engineers.
General application