Article

Sparse classification with paired covariates

Armin Rauschenberger\(^{1,2}\) AR, Iuliana Ciocănea-Teodorescu\(^1\) ICT, Marianne A. Jonker\(^3\) MAJ, Renée X. Menezes\(^1\) RXM, and Mark A. van de Wiel\(^{1,4}\) MvdW

\(^1\) Department of Epidemiology and Biostatistics, Amsterdam UMC, VU University Amsterdam, Amsterdam, The Netherlands

\(^2\) Luxembourg Centre for Systems Biomedicine, University of Luxembourg, Esch-sur-Alzette, Luxembourg

\(^3\) Department for Health Evidence, Radboud University Medical Center, Nijmegen, The Netherlands

\(^4\) MRC Biostatistics Unit, University of Cambridge, Cambridge, UK

Abstract

This paper introduces the paired lasso: a generalisation of the lasso for paired covariate settings. Our aim is to predict a single response from two high-dimensional covariate sets. We assume a one-to-one correspondence between the covariate sets, with each covariate in one set forming a pair with a covariate in the other set. Paired covariates arise, for example, when two transformations of the same data are available. It is often unknown which of the two covariate sets leads to better predictions, or whether the two covariate sets complement each other. The paired lasso addresses this problem by weighting the covariates to improve the selection from the covariate sets and the covariate pairs. It thereby combines information from both covariate sets and accounts for the paired structure. We tested the paired lasso on more than 2000 classification problems with experimental genomics data, and found that for estimating sparse but predictive models, the paired lasso outperforms the standard and the adaptive lasso. The R package palasso is available from CRAN.

Full text (open access)

Rauschenberger et al. (2020). “Sparse classification with paired covariates”. Advances in Data Analysis and Classification 14:571-588. doi: 10.1007/s11634-019-00375-6. (Click here to access PDF.)