By Léon Van Wouwe, Clinical Innovation Director, Volv Global
Why the next useful thing AI does in solid tumour care is about timing, not detection.
In most solid tumours, meaning cancers that form a mass of tissue rather than blood cancers, the patient is not hidden. They have a name on the list for the multidisciplinary team meeting, where surgeons, oncologists, radiologists and pathologists agree each patient’s plan. They have a stage, a tumour type confirmed under the microscope, a clinic slot and a named consultant. Oncology is among the most protocolised parts of medicine. And yet sit in on any of those meetings and you notice how many decisions have to be taken before the information that would fully inform them is at hand.
Talking to clinicians, you will hear that the molecular test is still pending. The staging scan cannot take into account the microscopic spread that will only manifest itself nine months from now. The consultation to check on whether treatment is working is scheduled for week twelve, but the disease advanced in week five. The prognosis under discussion is a trial population average, while the actual patient involved has a combination of other conditions that no trial really accounted for by design.
I have spent much of my recent efforts on the other side of this problem. Finding cancer earlier, identifying the patient nobody has identified. That work matters enormously, catching cancer early remains one of the strongest outcome predictors we have, and it is a central part of what we do at Volv Global. But it answers one question, and there is a second sitting underneath it that attracts far less attention. In the common solid tumours the patient has usually already been found. What remains uncertain is sequence and timing. Which therapy, in what order, at what intensity, for how long, and when to stop, switch or refer to a speciality cancer care facility.
The field already accepted that prediction changes decisions
Predictive analytics is sometimes discussed as though it were a new proposition for cancer care. It is not. Consider a blood test now used after bowel cancer surgery. It looks for circulating tumour DNA, or ctDNA: fragments of genetic material shed by cancer cells into the bloodstream. These fragments survive in the blood for under two hours, so once the tumour producing them has been removed, the trace it left behind disappears within a day. That is what makes a test taken four to seven weeks later informative. If tumour DNA is still detectable then, something is still shedding it: cancer cells left at the margin, or deposits too small for a CT scan to see. The test cannot say where they are, and it is not a diagnosis of relapse. It says the probability of relapse is high enough to justify chemotherapy after surgery, or low enough to withhold it.
The DYNAMIC trial, led by Jeanne Tie and colleagues, put this to a randomised test in stage II colon cancer, where the tumour has grown through the bowel wall but has not reached the lymph nodes. Guiding treatment by ctDNA cut the use of chemotherapy after surgery from around 28 per cent of patients to 15 per cent, and at five years the proportion of patients alive with no return of their cancer was the same in both groups, 88 against 87 per cent. Fewer patients treated, but similar disease outcomes were achieved across both groups. (1)
DYNAMIC-III tested the same idea one stage later, where the cancer has reached the lymph nodes. Patients with no detectable ctDNA after surgery had their treatment reduced, and the reduction was substantial: oxaliplatin, the harsher component of the standard regimen, was used in 34.8 per cent of them against 88.6 per cent under standard care, with fewer severe side effects and fewer hospital admissions.
But the trial did not clear the bar it had set itself. Three-year recurrence-free survival was 85.3 against 88.1 per cent, and non-inferiority was not met. (2)
Those results need to be considered together, because the pair says more than either alone. The prediction held. Patients with no detectable ctDNA did well, and far better than the patients in whom it was found.
What the trial did not prove was that the reduction in treatment given on the strength of that prediction was safe. And the same paper shows why that is unsurprising. Among the patients who did have ctDNA detected, the amount of it mattered. Recurrence-free survival fell from 77 per cent in those with the lowest levels to 23 per cent in those with the highest. The signal was graded. Yet the decision it was asked to drive was binary.
Also, the trial did not find an answer for those higher-risk patients. It tested that question separately, and intensifying their chemotherapy did not improve outcomes.
So the principle that a prediction can replace a stage-based rule as the basis for a treatment decision is established. We we cannot decide yet, how far to act on any one prediction, in which patients, and what to offer those it flags as high risk. That is an argument for better prediction rather than less of it, and it narrows the question usefully: what else can be predicted well enough to act on, from which data, and with what margin for being wrong.
What the models can now do
These models can do quite a lot now, and mostly without asking for new assessments. These models run on the information routine care already generates.
Two results show how this helps answer different questions, helping with clinical decision making. The first question is around what is likely to happen to this patient. The second probes into whether a particular treatment will work for them.
The first is prognosis from the slide already in the file. CHIEF, the Clinical Histopathology Imaging Evaluation Foundation model published in Nature in 2024, was tested on nearly 20,000 digitised tissue slides from two dozen hospitals and cohorts internationally. It separated patients into groups with markedly different survival using nothing but the routine haematoxylin and eosin stain, the pink-and-purple dye every pathology laboratory already applies to every specimen. No extra test, no extra sample, no extra cost. (3)
The second is response prediction that beats the biomarkers in current use. COMPASS reads the pattern of gene activity in a tumour sample and was assessed across seven cancer types. In a trial held back from its training, patients it flagged as likely to respond to the immunotherapy atezolizumab had one-year survival of 86 per cent, against 40 per cent for those it did not, separating them better than either biomarker used to make this call today. (4) The authors are careful to note that without a comparator arm of untreated patients, they cannot fully separate the two questions, so some of that signal may be prognosis rather than drug-specific prediction.
These are research results, and I will come to their limitations. But together they point to something structural: at the moment a clinician has to make a decision and choose between options, there is usually more predictive signal already in the record than is being used for making the decision.
The prognostic problem we prefer not to discuss
Prediction in oncology is not only about picking drugs. It is also about how long we think a patient has, and here human judgement has a specific and instructive weakness.
It is easy to assume that clinicians are optimistic, but the evidence tells a more complicated story. Consider a question widely used in clinical practice: “Would you be surprised if this person died within the next year?” An analysis of 56 patient groups, involving nearly 70,000 people, found that clinicians were usually right when they answered “yes”: 89 per cent of those patients were still alive a year later. But when clinicians answered “no”, identifying the patient as at risk of dying, only 40 per cent died within the year. In other words, three in five patients flagged as approaching the end of life were still alive a year later. (5)
This calls for two considerations. The first is that the judgement is more dependable in one direction than the other. Part of that gap simply reflects how uncommon death within a year is in most of these populations: a base rate the clinician cannot observe from a single consultation. The second is more awkward than error in either direction: in one prospective cohort, how confident an oncologist felt about a prediction bore no relationship to whether it proved right. (6)
The problem is not that clinicians cannot predict what will happen. Their judgements provide useful information, and studies comparing them with models rarely show either one making better predictions by a wide margin. What a well-built model offers here is not sharper foresight. It is a reliability that can be measured, published and monitored. A clinician cannot know, in the room, whether this particular judgement falls in the dependable direction or the unreliable one, and nothing about the experience of making it will indicate that. A model’s calibration can at least be checked before it is deployed and watched while it is in use. That is the difference between a competence problem and a calibration problem, and it is the argument for building one carefully rather than not at all.
So what is the consequence of miscalibration? Rarely, a wrong decision. But what it does often lead to, is the right decision taken late. The referral that would have helped in month four happens in month eleven. The conversation that would have changed how someone spent their remaining time happens once there is little time left to shape. And where nobody recognises the moment at all, the decision is not so much made as settled by default. The pattern is visible in cancers where the prognosis is not even ambiguous.
Take platinum-resistant ovarian cancer, a milestone after which median survival was around 15 months and the prognosis is about as legible as oncology gets. The median time from that point to a palliative care referral was nine months, and 43 per cent of referrals came within three months of death. Referral was reacting to decline rather than anticipating it. (7)
And that timing matters in a way that has been tested rather than observed. Patients with metastatic lung cancer randomised to palliative care alongside standard treatment from diagnosis had better quality of life, less than half the rate of depressive symptoms, and less aggressive care at the end of life, 33 against 54 per cent. (8) A secondary analysis found the mechanism you would hope for: those patients were more likely to hold an accurate understanding of their own prognosis, and among patients who did, chemotherapy near the end of life was far less common. (9)
If late recognition is the problem, the obvious fix is to make the moment visible. One trial has tested exactly that, and it is the result I find most instructive in this whole field. At the University of Pennsylvania, a machine learning model predicting six-month mortality from routine electronic health record data was paired with simple prompts to oncology clinicians: a weekly list of high-risk patients, and a nudge before the relevant appointment. Across 20,506 patients and 41,021 encounters, conversations about goals and priorities in serious illness rose from 3.4 to 13.5 per cent of encounters with patients at high risk of death, and systemic therapy at the end of life fell from 10.4 to 7.5 per cent. (10)
The lesson is not that the algorithm was clever. The lesson is that the algorithm on its own would have changed nothing. What changed care was a prediction attached to a specific decision, delivered to a specific person, at the specific moment the decision was live. The same lesson has been learned in reverse outside oncology: an implementation study of a validated prognostic model in advanced dementia found that making physicians aware of a patient’s predicted one-year mortality risk was, on its own, insufficient to change referral behaviour. (11) Almost every disappointing AI deployment in healthcare I have seen failed on that last clause rather than on accuracy.
The constraint is not model capability
This piece would be incomplete without examining the quality of the evidence behind these models. Here, “bias” has two different meanings. Algorithmic bias occurs when skewed or incomplete training data causes a model’s predictions to reproduce existing inequalities. Methodological bias comes from weaknesses in a study’s design, sample size or validation. Many problems in the current evidence base fall into this second category. A systematic review of machine-learning models for predicting outcomes in oncology assessed 152 models across 62 publications. It judged 84 per cent of the models developed to be at high risk of bias, mainly because of small sample sizes and split-sample internal validation, where a single dataset is divided into development and testing groups. The review also found that 82 per cent of development analyses and 57 per cent of validation analyses did not assess calibration, whether predicted risks matched observed outcomes. (12)
Calibration deserves a plain explanation because it is crucial to understanding whether a model’s predictions are useful. Discrimination asks whether a model correctly ranks patients from lower to higher risk. Calibration asks whether its numbers mean what they say. For example, among patients given a 20 per cent risk of an event, do roughly 20 per cent actually experience it? Discrimination is what gets published. Calibration determines whether a clinician can safely act on a risk estimate. A model can rank patients very well while consistently overstating or understating their risk. This is the same problem described earlier in human prognostication, but expressed through more sophisticated mathematics.
Strong results on one dataset do not guarantee that a model will perform well on others. An independent evaluation posted this year as a preprint, and not yet peer reviewed, found that several models using gene activity to predict responses to immunotherapy, including COMPASS, performed considerably less well on external datasets than their original reports suggested. This is not a reason to dismiss the work. It shows why independent validation is needed.
There is also a practical issue: what the available data actually contain. Cancer stage, tumour subtype and assessments of treatment response are often recorded in the text of pathology and radiology reports rather than in structured fields. Claims and administrative records reliably show treatment sequences, admissions and events, but provide much less information about how well patients responded to treatment. Before its scope is finalised, any serious programme should state which clinical decisions its data can support and which it cannot. This is not a limitation to be worked around behind the scenes. It defines the limits of what the programme can credibly offer.
Where prediction models have the most use
A prediction in clinical care only adds value where room for action exists and the window for making a decision is still open. On that basis, four typical situations stand out.
- Recurrence risk after treatment given with the intention to cure, used to intensify for the patient who needs more and, just as importantly, to de-escalate for the patient who does not.
- Early identification of a response that will not last, ahead of the scheduled scan, rather than discovering the failure a cycle or two late.
- The referral window: predicting when a patient will cease to be eligible for cell therapy, transplant or a clinical trial, rather than discovering it in hindsight.
- Relapse risk conditional on stopping planned therapy. The two-year stopping point for immunotherapy was never established by comparing durations; it entered practice as a trial design choice and hardened into convention. Most clinicians do not follow it, and the observational evidence suggests stopping is safe on average without identifying who the exceptions are. Predicting which individual patient relapses on stopping is the missing piece. (13)
Each has an identifiable decision owner, a defined moment and a real alternative course of action. Prognostication with no attached action does not earn a place in medicine and patient care, however good the model.
Prediction cuts both ways
Notice that two of those four point towards giving less treatment, not more. Prediction cuts both ways, and de-escalation may prove the larger prize: less toxicity, fewer hospital days, less of what has come to be called time toxicity, the share of a patient’s remaining life spent in waiting rooms and infusion chairs. Treatment intensity would then be matched to individual risk rather than to a stage category, which is itself a crude probability wearing a confident label.
That is the cultural obstacle rather than the technical one. We are comfortable acting on a measurement, a hard result, the observed disease stage. We are much less comfortable acting on a probability, even when the probability is better calibrated than the current disease stage. And a probability raises a question a guideline conveniently does not ask: if the model said 12 per cent and the patient relapsed, who takes ownership for that?
Clinical decisions remain sovereign
Tools of this kind are now reaching the field rather than existing solely in the literature. Volv Global’s inFlow is one example, predicting progression, response, relapse and event risk from routine longitudinal records. There are others, and the more the better. What matters far more than which tool is the discipline around it: prospective validation in the setting where it will be used, calibration reported and not only discrimination, an explicit action attached to every prediction, and a clear boundary of use. These are decision-support tools. They do not diagnose, prescribe or replace clinical judgement.
What this is actually for
Strip away the modelling and focussing on what this means clinically and the value becomes more clear: It is the patient who reaches the right line of therapy while still well enough to benefit from it. It is the patient spared six months of chemotherapy they were never going to need. It is the conversation that happens in month four rather than month eleven, while it can still shape how someone spends the time they have.
There is growing evidence these things can be predicted. We now need to think carefully if we are willing to reorganise a decision around a number, and to be accountable for doing so. That, rather than model architecture, is where the next few years of progress in oncology AI will be decided. It is technology, for the use of humans.
Which leaves me with a question I would rather put to you than answer myself. Of those four situations, which would you build first, and what would have to be true of the model before you would act on its number in front of a patient? If you sit in those meetings, your answer is probably better than mine. I would like to hear it.
About the author
Léon van Wouwe is Clinical Innovation Director at Volv Global SA, working with pharmaceutical partners on patient identification, care-gap analytics and real-world evidence in rare disease and oncology.
Links:
- Volv Global inFlow – Predicting patient outcomes
References
- Tie J, Cohen JD, Lahouel K, et al. Circulating tumor DNA analysis guiding adjuvant therapy in stage II colon cancer. N Engl J Med 2022;386:2261-2272. Five-year outcomes: Tie J, et al. Nat Med 2025;31:1509-1518. doi:10.1038/s41591-025-03579-w
- Tie J, Wang Y, Loree JM, et al. Circulating tumor DNA-guided adjuvant therapy in locally advanced colon cancer: the randomized phase 2/3 DYNAMIC-III trial. Nat Med 2025;31:4291-4300. doi:10.1038/s41591-025-04030-w
- Wang X, Zhao J, Marostica E, et al. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature 2024;634:970-978. doi:10.1038/s41586-024-07894-z
- Shen W, Moon I, Nguyen TH, et al. Generalizable AI predicts immunotherapy outcomes across cancers and treatments. Nat Med 2026;32:3010-3022. doi:10.1038/s41591-026-04502-7
- Gupta A, Burgess R, Drozd M, Gierula J, Witte K, Straw S. The Surprise Question and clinician-predicted prognosis: systematic review and meta-analysis. BMJ Support Palliat Care 2024;15:12-35. doi:10.1136/spcare-2024-004879
- Kim YJ, Yoon SJ, Suh SY, Hiratsuka Y, Kang B, Lee SW, et al. Performance of clinician prediction of survival in oncology outpatients with advanced cancer. PLoS One 2022;17(4):e0267467. doi:10.1371/journal.pone.0267467
- Haag JG, Adler AD, Sheeder J, Brubaker LW, Lefkowits C. Patterns of palliative care referral in platinum resistant ovarian cancer demonstrate reactive rather than proactive approach. Gynecol Oncol Rep 2022;43:101053. doi:10.1016/j.gore.2022.101053
- Temel JS, Greer JA, Muzikansky A, et al. Early palliative care for patients with metastatic non-small-cell lung cancer. N Engl J Med 2010;363:733-742. doi:10.1056/NEJMoa1000678
- Temel JS, Greer JA, Admane S, et al. Longitudinal perceptions of prognosis and goals of therapy in patients with metastatic non-small-cell lung cancer: results of a randomized study of early palliative care. J Clin Oncol 2011;29:2319-2326. doi:10.1200/JCO.2010.32.4459
- Manz CR, Zhang Y, Chen K, et al. Long-term effect of machine learning-triggered behavioral nudges on serious illness conversations and end-of-life outcomes among patients with cancer: a randomized clinical trial. JAMA Oncol 2023;9:414-418. doi:10.1001/jamaoncol.2022.6303
- Subramaniam A, Tan WS, Tan HTR, et al. “Triggering the palliative intent”?: a qualitative implementation evaluation of a prognostication model for advanced dementia (PRO-MADE) in a geriatric tertiary care setting for the integration of early palliative care. BMC Palliat Care 2026;25:156. doi:10.1186/s12904-026-02138-5
- Dhiman P, Ma J, Andaur Navarro CL, et al. Risk of bias of prognostic models developed using machine learning: a systematic review in oncology. Diagn Progn Res 2022;6:13. doi:10.1186/s41512-022-00126-w. Companion analysis of the same review: Methodological conduct of prognostic prediction models developed using machine learning in oncology: a systematic review. BMC Med Res Methodol 2022;22:101. doi:10.1186/s12874-022-01577-x
- Sun L, Bleiberg B, Hwang WT, et al. Association between duration of immunotherapy and overall survival in advanced non-small cell lung cancer. JAMA Oncol 2023;9:1075-1082. doi:10.1001/jamaoncol.2023.1891
Photo by Gustavo Sánchez on Unsplash