Skip to content
All projects

FIG. 02.4 — Project notes

  • Independent
  • Finished

OncoLens: Explainable Tumor Classification Dashboard

A recall-first, explainable tumor classifier with an interactive Streamlit dashboard, built on the public Wisconsin Diagnostic Breast Cancer dataset.

Educational only — not a clinical tool

Built for learning on public data. It is not a medical device and must not be used for diagnosis.

Test recall at t = 0.2
0.976 (41/42)
Test ROC AUC
0.996
Nested-CV ROC AUC
0.9949 ± 0.0051
Tests
48 pytest
Screenshot of the OncoLens Streamlit app: cell-measurement sliders, a malignant prediction, and a SHAP waterfall chart explaining it.
FIG. 02.4 — OncoLens Streamlit app · Figure from the project repository

01Problem

In cancer screening, a missed malignant case costs more than a false alarm.

02Question

Can a transparent model be tuned to favor catching malignant cases — and explain each prediction?

03Data

UCI Wisconsin Diagnostic Breast Cancer dataset, loaded via scikit-learn: 569 samples, 30 features describing cell nuclei from digitized fine-needle aspirate images (212 malignant / 357 benign).

04Methods

  • Stratified 80/20 split (455 train / 114 test); the test set was used once.
  • Five models — logistic regression, random forest, gradient boosting, RBF SVM and XGBoost — tuned with GridSearchCV and compared by nested cross-validation.
  • Scaling inside every cross-validation pipeline, so nothing leaks from validation folds.
  • Calibration check (Brier score, expected calibration error).
  • Recall-first decision threshold chosen from out-of-fold training predictions: the highest threshold with out-of-fold recall ≥ 0.98, which gave 0.2.
  • Exact SHAP explanations (LinearExplainer), both global and per patient.
  • Error analysis of the remaining mistakes.
  • A Streamlit app for interactive predictions and explanations.
  • 48 pytest tests with continuous integration.

05Tools

  • Python
  • scikit-learn
  • XGBoost
  • SHAP
  • Streamlit
  • pytest
  • Docker

06Visualizations

Line chart of recall, precision and specificity against the decision threshold, with the chosen threshold of 0.2 marked (recall 0.982, precision 0.923) and the default 0.5 shown for comparison.
FIG. 02.4.1Figure from the project repository ·Threshold trade-off on out-of-fold training predictions: lowering the threshold to 0.2 catches more malignant cases.
Chart comparing nested cross-validation performance of logistic regression, random forest, gradient boosting, RBF SVM and XGBoost.
FIG. 02.4.2Figure from the project repository ·Model comparison by nested cross-validation.
SHAP beeswarm plot showing the distribution of per-sample feature contributions for the logistic regression model.
FIG. 02.4.3Figure from the project repository ·SHAP beeswarm: how each feature pushes the model’s output across samples.

07Findings

  • Logistic regression was chosen (nested-CV ROC AUC 0.9949 ± 0.0051). All five models were close (0.986–0.995); the differences were within fold-to-fold variation.
  • Held-out test set (114 samples, 42 malignant), threshold 0.5: recall 0.929 (39/42), accuracy 0.965, ROC AUC 0.996.
  • At the chosen threshold of 0.2: recall 0.976 (41/42), precision 0.911, specificity 0.944, accuracy 0.956.
  • Missed malignant cases fell from 3 to 1, at the cost of 3 extra false alarms.

08Limitations

  • Not a clinical tool — educational only.
  • Small, old, single-institution dataset (early 1990s).
  • Uses pre-computed features, not raw images.
  • Small test set: one case equals 2.4 points of recall.
  • The recall target behind the threshold is a value judgment, not one set by clinicians.
  • SHAP explains the model, not the biology.