Exploring gene expression profiles using density-based dimensionality reduction methods
Abstract
Traditional statistical methods, such as principal component analysis, often fail to capture complex dependencies in gene expression data. To address this limitation, we propose a functional framework combining multidimensional scaling with a density-based version of principal component analysis. By representing gene expression profiles through estimated distributions, the method captures both distributional shape and variability across individuals. Using artificial datasets, we show that clustering performed on density-based scores accurately recovers the original class structure. We also compare our approach with two nonlinear dimensionality reduction techniques, uniform manifold approximation and projection and diffusion maps, as dimensionality increases, highlighting the importance of $L_2$ normalization in preserving discriminative power. For moderate dimensions, densities are estimated using a multivariate gamma kernel, well suited to the non-negative and asymmetric nature of transcriptomic data. Finally, we establish convergence results for the estimated inner products and prove the spectral consistency of the resulting eigenvalues and eigenvectors. The extracted principal components effectively capture both lower-order statistical moments and complex gene interaction patterns that are often inaccessible to classical linear methods.
Keywords
- Clustering
- gene-gene interactions
- multidimensional scaling
- nonparametric gamma kernel
- principal component.
Thanks
References
- [1] H. Abdi and L.J. Williams, Principal component analysis. WIREs Comput. Stat. 2 (4), 433–459, 2010.
- [2] C.C. Aggarwal, A. Hinneburg and D. A. Keim, On the surprising behavior of distance metrics in high dimensional space. In: International Conference on Database Theory, pp. 420–434, Springer, London, UK, 2001.
- [3] E. Bair, T. Hastie, D. Paul and R. Tibshirani, Prediction by supervised principal components. J. Am. Stat. Assoc. 101 (473), 119–137, 2006.
- [4] M. Belkin and P. Niyogi, Laplacian eigenmaps for dimensionality reduction and data representation. Neural Comput. 15 (6), 1373–1396, 2003.
- [5] J. Bigot, R. Gouet, T. Klein and A. López, Geodesic PCA in the Wasserstein space by convex PCA. Ann. Inst. Henri Poincaré Probab. Stat. 53 (1), 1–26, 2017.
- [6] T. Bouezmarni and O. Scaillet, Consistency of asymmetric kernel density estimators and smoothed histograms with application to income data. Econom. Theory. 21 (2), 390–412, 2005.
- [7] R. Boumaza, Distribution asymptotique de l’affinité L2 de densités gaussiennes. C. R. Acad. Sci. Paris, Ser. I. 328 (6), 527–529, 1999.
- [8] R. Boumaza, S. Yousfi and S. Demotes-Mainard, Interpreting the principal component analysis of multivariate density functions. Commun. Stat. Theory Methods. 44 (16), 3321–3339, 2015.
Details
Primary Language
English
Subjects
Semi- and Unsupervised Learning, Biostatistics, Computational Statistics, Statistical Data Science, Probability Theory, Operator Algebras and Functional Analysis
Journal Section
Research Article
Authors
Smail Yousfi
*
0000-0002-3990-6016
Algeria
Early Pub Date
July 8, 2026
Publication Date
August 17, 2026
Submission Date
December 12, 2025
Acceptance Date
June 24, 2026
Published in Issue
Year 2026 Volume: 55 Number: 4