Treatment response from H&E biopsy scans

We built the AI and ML pipeline that trains a model on H&E-stained cancer biopsy scans to predict a patient’s chance of responding to treatment.
- Context
- A European CRO working on cancer biopsies wanted to know whether the tissue itself carries enough signal to say, before treatment starts, which patients are likely to respond. The input is the routine artefact of any pathology lab: H&E-stained biopsy slides, digitised at high resolution. Nothing new is collected, which is what makes the question worth asking. If the answer is in material that already exists, the path from research to clinical use is much shorter than it is for a new assay.
- Challenge
- Three properties make this harder than a standard image task. A whole-slide scan is far too large to put in front of a model directly, so each slide has to be broken into tiles and the tiles reassembled into a single prediction per patient. The label does not live in the image: response is a clinical outcome recorded later, so every slide is weakly labelled and the model has to find the signal without being told where it is. And the number of patients is small even when the number of tiles is very large, which means a model can memorise the cohort long before it learns anything that transfers.
- Approach
- We built the pipeline end to end. Ingestion and quality control of the scans, tissue segmentation to drop background and staining artefacts, tiling with stain normalisation so slides from different scanners are comparable, feature extraction over the tiles, and an aggregation step that turns thousands of tile representations into one patient-level prediction. Splits are by patient and never by tile, so the same patient cannot appear on both sides of a split. The validation scheme and the reporting metric were agreed before training started. Data and code are pinned and every training run is recorded, so any figure in a report can be traced back to the run that produced it. The shape of the pipeline is the one this class of problem calls for: whole-slide image to tiles, a pathology feature encoder over the tiles, attention-based multiple-instance learning to aggregate them, and a biomarker prediction per patient at the end.
- Outcome
- The pipeline trains and evaluates a model that predicts response to treatment from the biopsy scan alone, and it does so reproducibly: the same data and the same code give the same numbers again. Validation is by patient-level cross-validation against a metric agreed before training, and the held-out split stayed untouched until the end. The work was scoped as a technical feasibility study, and we publish no performance figures from it.
- Stack
- Python throughout, with versioned data, pinned dependencies and recorded experiment runs. The partner holds the code, the model weights and the full training record, and can retrain without us.






