IDWSDS 2026, International Day of Women in Statistics and Data Science. Session S115, October 6, 13:30 to 14:00 UTC. Reliable Variable Selection for Biomedical Data Science: From Shrinkage Estimation to Interpretable Learning. Organizer: Mina Norouzirad. Chair: Chiara Orsini. Sponsor: Bernoulli Society for Mathematical Statistics and Probability.
Reliable Variable Selection
for Biomedical Data Science
From shrinkage estimation to interpretable learning
Mina Norouzirad
Center for Mathematics and Applications (NOVA Math)
NOVA School of Science and Technology, NOVA University Lisbon
Joint work with Marta Belchior Lopes and Tomás da Rosa Bandeira
IDWSDS 2026  |  Plenary session S115
This work is funded by national funds through the FCT – Fundação para a Ciência e a Tecnologia, I.P., under the scope of the projects UID/00297/2025 (https://doi.org/10.54499/UID/00297/2025), UID/PRR/00297/2025 (https://doi.org/10.54499/UID/PRR/00297/2025), and UID/PRR2/00297/2025 (https://doi.org/10.54499/UID/PRR2/00297/2025) (Center for Mathematics and Applications – NOVA Math).

A question that sounds simple

619
patients with a brain tumour
20,504
genes measured per patient
3
tumour subtypes to tell apart
Which genes really tell the subtypes apart?
And would we find the same genes if we collected the data again?

Gliomas: why the subtype matters

  • Gliomas arise from the glial (support) cells of the brain and make up most malignant adult brain tumours.
  • The subtype drives prognosis and the choice of therapy.
  • Subtyping increasingly relies on molecular profiles, not only imaging and histology.
Class shares in the TCGA cohort (n = 619).

From a biopsy to a very wide table

p = 20,504 genes (columns) n = 619patients entry xij = expression of gene j in patient i (RNA-seq, TCGA)
  • Response: subtype ∈ {astro, GBM, oligo} ⇒ multinomial logistic regression.
  • There are about 33 times more unknowns per class than patients.

Why biomedical data is hard for classical regression

  1. High dimensionality: p ≫ n, so the maximum-likelihood fit is not unique.
  2. Multicollinearity: genes in the same pathway move together.
  3. Noise: measurement and biological variability.
  4. Class imbalance: 41 / 32 / 27%.
  5. Heterogeneity between patients.

What goes wrong

  • Overfitting: perfect on training data, poor on new patients.
  • Unstable selection: a small change in the data gives a different gene list.
  • Conclusions that domain experts cannot interpret or trust.

The correlation trap

Two genes carry the same signal. Which one is “the” biomarker?
ρ(A,B) = 0.95
Simulated: n = 60. A and B are two noisy readouts of the same signal; C–H are noise. Each bar counts how often a gene was selected.
  • The LASSO tends to keep one gene from a correlated group, almost arbitrarily.
  • Prediction barely changes, but the gene list does.
  • The elastic net keeps the group together.
Reliable selection must use the correlation structure, not fight it.
Part 1

The shrinkage toolkit

The idea of shrinkage

β = arg minβ  { − ℓ(β)  +  λ P(β) }
fit to datapenalty
  • Accept a little bias to buy a large reduction in variance.
  • λ controls how hard coefficients are pulled towards zero; chosen by cross-validation.
  • The shape of P decides the behaviour: shrink only, select, or group.
  • Stein (1956) and Hoerl & Kennard (1970) planted the idea; modern omics made it indispensable.

Penalty geometry: why some penalties select

Drag the slider from ridge to LASSO. The solution is where the loss contours first touch the region.
RidgeElastic netLASSO
mixing parameter α = 0.00
β1 = –
β2 = –
Corners in the constraint region are what produce exact zeros. Drag the black OLS point to try other data.

Four estimators, four behaviours

Penalty P(β)ShrinksSelectsGroups correlated
RidgeΣ βj²✓partly
LASSOΣ |βj|✓✓
Elastic netα Σ |βj| + (1−α)/2 Σ βj²✓✓✓
Adaptive LASSOΣ wj |βj|,  wj = 1/|βj|γ✓✓via weights
  • Adaptive LASSO penalises unimportant genes more and important genes less; it has the oracle property (Zou, 2006).
  • Its quality depends entirely on the pilot estimate β that produces the weights.
Part 2

Letting correlation guide the penalty

A correlation-based penalty

Tutz & Ulbricht (2009): encode gene–gene correlation directly in the penalty
Pcor(β) = Σi<j [ (βi − βj)² / (1 − ρij)  +  (βi + βj)² / (1 + ρij) ]  =  β⊤Wβ
  • ρ → 1: a large cost for different coefficients, so correlated genes share the signal.
  • ρ → −1: pushes coefficients towards opposite signs.
  • ρ = 0: reduces to a ridge-type penalty.

Interpretation

A generalised ridge: with θ = W1/2β and X* = XW−1/2, it is ordinary ridge on “whitened” genes. Redundant contributions are attenuated instead of arbitrarily dropped.

How correlation shapes the penalty

Two correlated genes whose true effects are equal. Draw new samples and watch where each estimator lands.
correlation ρ between the genes0.90

The computational wall

3 GB
for one 19,072 × 19,072 correlation matrix
days
forming W−1/2 on a high-memory server, without finishing
  • After removing near-zero-variance genes, p = 19,072 remain.
  • The penalty is the right idea, applied at the wrong scale.
  • A method is only reliable if it is also computable and reproducible by others.

Screen, score, sharpen: a two-phase pipeline

1. ScreenElastic net keeps aninformative gene subset 2. ScoreCorrelation-based penaltyon the subset gives scores 3. SharpenMultinomial adaptive LASSOweight = 1 / |score|, per class 19,072 genes in≈ 18 genes out
  • Weights are class-specific: a gene can matter for glioblastoma but not for oligodendroglioma.
  • glmnet accepts only one weight per gene, so we built pampam, a fork that accepts a p × K penalty matrix.
Part 3

What we found

Study design

  • TCGA RNA-seq, n = 619, three subtypes.
  • 100 stratified train/test splits.
  • Tuning by 10-fold cross-validation inside each training set.
  • Compared: ridge, LASSO, adaptive LASSO, elastic net, elastic net + correlation penalty, and the full pipeline.

What we measure

  • Test accuracy, weighted F1, multinomial AUC
  • Number of genes kept
  • How often each gene is selected across the 100 runs (stability)
  • Runtime

Almost everything is accurate. Very little is interpretable.

Mean test accuracy over 100 splits. Press the button to spread the methods by how many genes they keep.

The numbers behind the picture

MethodAcc (train)Acc (test)AUCwF1# genes
Ridge0.9960.9820.9800.98219,074
LASSO0.9920.9850.9820.98537
Adaptive LASSO (ridge pilot)0.9930.9850.9820.98522
Elastic net0.9950.9860.9820.986494
Elastic net + correlation penalty0.9720.9630.9420.962494
Screen–score–sharpen pipeline0.9910.9850.9820.98518
Means over 100 stratified splits. AUC: Hand–Till multinomial AUC; wF1: weighted F1.
  • Same accuracy as elastic net to within 0.1 point, with 27 times fewer genes.
  • Errors are almost all astrocytoma ↔ oligodendroglioma; glioblastoma is separated almost perfectly.

Beyond accuracy: stability, generalisation, cost

0.6
points train–test gap
vs 1.4 for ridge: the sparse model generalises better
27 s
per fit
vs 275 s for adaptive LASSO with a ridge pilot on all genes
100/100
runs select ARSD, FBXO42, PTBP2
a core that does not depend on the split

Honest reading

The pipeline does not beat elastic net on accuracy, and it is close to a well-tuned adaptive LASSO. Its gains are a smaller, more stable signature, class-specific weights, and a 10-fold faster fit. In biomedical data science those are often the gains that matter.

A compact, stable, interpretable gene panel

Times each gene was selected across 100 runs. Filter by the subtype a gene leans towards.
  • ARSD, KIAA0495 chiefly separate glioblastoma.
  • PTBP2, DRG2 lean towards astrocytoma; FBXO42, UBA2 towards oligodendroglioma.
  • The astro–oligo signal is spread over many genes: this is the hard pair.
  • ATRX is a known glioma marker, a sanity check from biology.
Closing

What makes variable selection reliable?

Five properties, and how to get them

AccurateShrinkage controls variance; tune λ by cross-validation, report test performance only.
StableUse the correlation structure; report selection frequencies over resamples, not a single list.
TransparentPrefer penalties whose behaviour you can explain: geometry, weights, grouping.
ReproducibleMethods must be computable at real scale; share code and packages (pampam).
InterpretableA small panel that a clinician or biologist can check against what is already known.

Take-home messages

  1. In biomedical data, accuracy is rarely the bottleneck: many methods predict well. The difference lies in which genes they keep, and how stably.
  2. Correlation is information, not a nuisance. Penalties that use it give more reliable selections.
  3. Shrinkage ideas from the 1970s, combined with adaptive weights and good software, give models that are accurate and understandable.
pampam on GitHub
github.com/tomasb21/pampam

Thank you

Questions and discussion welcome
Mina Norouzirad  |  NOVA Math  |  mnrzrad.github.io
With Marta Belchior Lopes and Tomás da Rosa Bandeira
This work is funded by national funds through the FCT – Fundação para a Ciência e a Tecnologia, I.P., under the scope of the projects UID/00297/2025 (https://doi.org/10.54499/UID/00297/2025), UID/PRR/00297/2025 (https://doi.org/10.54499/UID/PRR/00297/2025), and UID/PRR2/00297/2025 (https://doi.org/10.54499/UID/PRR2/00297/2025) (Center for Mathematics and Applications – NOVA Math).

References

  • Hoerl, A.E., Kennard, R.W. (1970). Ridge regression: biased estimation for nonorthogonal problems. Technometrics, 12, 55–67.
  • Tibshirani, R. (1996). Regression shrinkage and selection via the lasso. JRSS B, 58, 267–288.
  • Zou, H., Hastie, T. (2005). Regularization and variable selection via the elastic net. JRSS B, 67, 301–320.
  • Zou, H. (2006). The adaptive lasso and its oracle properties. JASA, 101, 1418–1429.
  • Tutz, G., Ulbricht, J. (2009). Penalized regression with correlation-based penalty. Statistics and Computing, 19, 239–253.
  • Friedman, J., Hastie, T., Tibshirani, R. (2010). Regularization paths for generalized linear models via coordinate descent. J. Stat. Softw., 33(1).
  • Hand, D.J., Till, R.J. (2001). A simple generalisation of the area under the ROC curve for multiple class classification problems. Machine Learning, 45, 171–186.
  • The Cancer Genome Atlas (TCGA) Research Network. Glioma RNA-seq data.
  • Bandeira, T. da R. MSc thesis, NOVA School of Science and Technology.