Research Article

Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function

Volume: 1 Number: 2 December 20, 2025

Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function

Abstract

In this paper, we introduce and propose a novel subset selection of variables in the quantile regression (QR) model using information-based model selection criterion ICOMP(IFIM)misspec as the fitness function within the genetic algorithm (GA) as our optimization tool. Estimation of the coefficients of the QR model is carried out using a computationally efficient weighted least squares (WLS) method. A nonparametric kernel covariance estimator of the QR model is derived and implemented as the asymptotic covariance matrix of the model to derive and score ICOMP(IFIM)misspec. A real numerical example is carried out on a benchmark prostate data set to show the performance of the ICOMP(IFIM)misspec criterion along with AIC and SBC for selecting the best subset of variables to explain the conditional distribution of response on several quantile levels via all possible subset model selection methods and the GA. Our results show that the AIC criterion includes many redundant variables than the ICOMP-based models do, which is not so surprising. The models selected by the GA coincide with the models selected via all possible subset selection methods. However, GA is a more time-efficient and less costly procedure than all possible subset selection methods, especially in high-dimensional datasets. Our proposed quantile regression (QR) model approach outperforms the model selection as compared to the classic linear regression (LR) modeling approach.

Keywords

Supporting Institution

TUBITAK Post Doctoral Scholarship

Ethical Statement

This paper has not been submitted elsewhere

References

  1. Akaike, H. (1973). Information theory and an extension of the maximum likelihood principle. In B. Petrox &F. Csaki (Eds.), Second international symposium on information theory. (p. 267-281). Budapest.
  2. Barrodale, I., & Roberts, F. (1973). An improved algorithm for discrete L1 linear approximation. SIAM Journalof Numerical Analysis, 10(5), 839-848.
  3. Barrodale, I., & Roberts, F. (1974). Solution of an overdetermined system of equations in the L1 norm.Communications of the Association for Computing Machinery, 17, 319-320.
  4. Beaton, A. E., & Tukey, J. W. (1974). The fitting of power series, meaning polynomials, illustrated on band-spectroscopic data. Technometrics, 16, 147-185.
  5. Behl, P., Claeskens, G., & Dette, H. (2014). Focused model selection in quantile regression. Statistica Sinica,24, 601-624.
  6. Birkes, D., & Dodge, Y. (1993). Alternative methods of regression. New York: Wiley and Sons,Inc.
  7. Bowman, A. W. (1984). An investigation of the properties of some simple kernel density estimators. Journal ofthe Royal Statistical Society: Series B (Methodological), 46(3), 305–316.
  8. Bozdogan, H. (1988). Icomp: A new model-selection criteria. In H. Bock (Ed.), Classification and relatedmethods of data analysis. North-Holland.

Details

Primary Language

English

Subjects

Statistical Data Science

Journal Section

Research Article

Authors

Publication Date

December 20, 2025

Submission Date

April 29, 2025

Acceptance Date

July 22, 2025

Published in Issue

Year 2025 Volume: 1 Number: 2

APA
Bozdogan, H. (2025). Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function. Smyrna Journal of Natural and Data Sciences, 1(2), 14-34. https://izlik.org/JA82ZP53EK
AMA
1.Bozdogan H. Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function. Smyrna Journal of Natural and Data Sciences. 2025;1(2):14-34. https://izlik.org/JA82ZP53EK
Chicago
Bozdogan, Hamparsum. 2025. “Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity As the Fitness Function”. Smyrna Journal of Natural and Data Sciences 1 (2): 14-34. https://izlik.org/JA82ZP53EK.
EndNote
Bozdogan H (December 1, 2025) Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function. Smyrna Journal of Natural and Data Sciences 1 2 14–34.
IEEE
[1]H. Bozdogan, “Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function”, Smyrna Journal of Natural and Data Sciences, vol. 1, no. 2, pp. 14–34, Dec. 2025, [Online]. Available: https://izlik.org/JA82ZP53EK
ISNAD
Bozdogan, Hamparsum. “Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity As the Fitness Function”. Smyrna Journal of Natural and Data Sciences 1/2 (December 1, 2025): 14-34. https://izlik.org/JA82ZP53EK.
JAMA
1.Bozdogan H. Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function. Smyrna Journal of Natural and Data Sciences. 2025;1:14–34.
MLA
Bozdogan, Hamparsum. “Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity As the Fitness Function”. Smyrna Journal of Natural and Data Sciences, vol. 1, no. 2, Dec. 2025, pp. 14-34, https://izlik.org/JA82ZP53EK.
Vancouver
1.Hamparsum Bozdogan. Subset Selection of Best Predictors in Quantile Regression Model Using the Genetic Algorithm and Information Measure of Complexity as the Fitness Function. Smyrna Journal of Natural and Data Sciences [Internet]. 2025 Dec. 1;1(2):14-3. Available from: https://izlik.org/JA82ZP53EK