Machine learning models using radiomics achieved 80% accuracy (pooled AUC 0.798) for predicting high PD-L1 expression in non-small cell lung cancer, with performance varying significantly by imaging type and algorithm, though clinical utility remains to be established .
Researchers conducted a systematic review and meta-analysis of 35 studies examining whether artificial intelligence models trained on imaging features (radiomics) could predict PD-L1 expression status in non-small cell lung cancer (NSCLC) patients. PD-L1 is a checkpoint protein that tumors use to evade immune detection, and its expression level directly influences whether patients will respond to immunotherapy drugs called immune checkpoint inhibitors (ICIs). Currently, determining PD-L1 status requires invasive tissue biopsies. An accurate non-invasive imaging-based alternative could streamline treatment planning and reduce procedural burden.
The analysis revealed that prediction performance depends heavily on which PD-L1 threshold is being targeted. When predicting higher PD-L1 expression (Tumor Proportion Score 50, or TPS50), AI models achieved a pooled area under the curve (AUC) of 0.798 in validation datasets. This represents reasonable discriminatory ability: an AUC of 0.80 means the model correctly ranks a randomly selected high-PD-L1 case as higher risk 80% of the time compared to a low-PD-L1 case. For lower thresholds (TPS1), the pooled AUC was 0.782, suggesting modestly lower performance at detecting minimal PD-L1 expression.
The choice of imaging modality significantly influenced results. Models built using contrast-enhanced CT or 18F-fluorodeoxyglucose (FDG) PET/CT outperformed those using non-contrast CT when predicting TPS50 status. This makes mechanistic sense: contrast enhancement and metabolic uptake provide additional tissue characterization beyond plain structural imaging. However, the pattern reversed when predicting any TPS (the lowest threshold): non-contrast CT models actually performed better, suggesting that different biological phenomena at different PD-L1 expression levels may be captured by different imaging features.
Algorithm selection also mattered. Machine learning approaches (traditional statistical algorithms like random forests or support vector machines) showed superior performance for TPS50 prediction, while deep learning models (neural networks) performed better for TPS1 prediction. The authors note this likely reflects differences in how these algorithm classes extract patterns from high-dimensional imaging data, though the mechanistic explanation remains unclear. Across all comparisons, performance metrics ranged broadly, indicating substantial heterogeneity in study quality and methodology that tempers confidence in the absolute numbers.
If you or a family member faces NSCLC and needs immunotherapy planning, this research suggests that radiomics-based AI models are not yet ready to replace tissue biopsies for PD-L1 assessment. While the 80% accuracy sounds reasonable, clinical decision-making for cancer treatment typically requires higher confidence thresholds. Current results show that roughly 1 in 5 patients would receive an incorrect PD-L1 classification using these models alone.
That said, this work validates the general principle that non-invasive imaging contains meaningful information about PD-L1 expression. The heterogeneity in the results highlights that imaging modality and algorithmic approach matter significantly. Future development should focus on standardizing which imaging protocols and analysis methods produce the most reliable predictions, and on validating whether these models add value when used alongside clinical and pathologic information rather than as standalone tools.
If you are undergoing lung cancer evaluation at a center using radiomics-based decision support, ask specifically whether the model has been validated in your demographic group and imaging setting, and whether it informs rather than replaces tissue-based assessment.
| Parameter | Details |
|---|---|
| Study type | Systematic review and meta-analysis |
| Number of included studies | 35 |
| Sample size | Not reported (aggregated across included studies) |
| Primary outcome | Pooled AUC for PD-L1 expression prediction |
| TPS50 validation AUC | 0.798 (95% CI not specified in abstract) |
| TPS1 validation AUC | 0.782 |
| Best performing imaging modality | Contrast-enhanced CT or FDG-PET/CT for TPS50 prediction |
| Best performing algorithm | Machine learning for TPS50; deep learning for TPS1 |
| Journal | European Respiratory Review |
| PubMed ID | 42419776 |
| Publication year | 2025 (accepted) |
Pan X, et al. Prediction of PD-L1 expression in nonsmall cell lung cancer using artificial intelligence models based on radiomics: a systematic review and meta-analysis. *European Respiratory Review*. PubMed ID: 42419776
ProtocolEngine provides general health information based on published research. This is not medical advice. Consult a healthcare professional before starting any supplement or health protocol.