Log-linear Models and Closed Form Estimates for Missing Values in Two Dimensional Contingency Tables
Abstract
The problem of missing data is frequently encountered in scientific research due to various reasons such as nonresponse in surveys, data recording errors, data loss, or limitations inherent in the study design. Missing data mechanisms are classified into three categories: missing completely at random (MCAR), missing at random (MAR), and not missing at random (NMAR). In the categorical data analysis, in contingency tables, the direct application of log-linear models in the presence of missing observations in one or more variables may lead to biased or misleading results. Therefore, in order to obtain valid statistical inferences, the missing data problem must be addressed using appropriate methodological approaches prior to analysis. In this study, log-linear models and their closed-form estimators are examined for two-dimensional contingency tables under scenarios where missing data occur in one variable as well as in both variables simultaneously. An illustrative example is conducted using the Myocardial Infarction Complications dataset, and the results are evaluated. The findings demonstrate that closed-form estimators provide an effective and interpretable framework for analyzing contingency tables with missing data, enabling reliable inference under different missing data mechanisms.
Keywords
Categorical data, Closed-form estimates, Contingency tables, Log-linear models, Missing data
References
- [1] Peng, C. Y., Harwell, M., Liou, S. M., & Ehman, L. H. (2006). Advances in missing data methods and implications for educational research. Real Data Analysis, 3178, 102. https://api.semanticscholar.org/CorpusID:14341113
- [2] Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581–592. https://doi.org/10.1093/biomet/63.3.581
- [3] Little, R. J. A., & Rubin, D. B. (1987). Statistical analysis with missing data. John Wiley & Sons.
- [4] Allison, P. D. (2001). Missing data. In Quantitative applications in the social sciences (pp. 72–89). SAGE.
- [5] Howell, D. C. (2007). The treatment of missing data. In The Sage handbook of social science methodology (pp. 208–224).
- [6] Schafer, J. L. (1997). Analysis of incomplete multivariate data. CRC Press.
- [7] Baker, S. G., Rosenberger, W. F., & Dersimonian, R. (1992). Closed-form estimates for missing counts in two-way contingency tables. Statistics in Medicine, 11(5), 643–657. https://doi.org/10.1002/sim.4780110509
- [8] Molenberghs, G., Beunckens, C., Sotto, C., & Kenward, M. G. (2008). Every missingness not at random model has a missingness at random counterpart with equal fit. Journal of the Royal Statistical Society: Series B, 70(2), 371–388. https://doi.org/10.1111/j.1467-9868.2007.00640.x
- [9] Kim, S., Park, Y., & Kim, D. (2015). On missing-at-random mechanism in two-way incomplete contingency tables. Statistics & Probability Letters, 96, 196–203. https://doi.org/10.1016/j.spl.2014.09.016
- [10] Kim, S., Jeon, S., & Kim, D. (2020). On log-linear modeling for an incomplete two-way contingency table with one variable subject to nonresponse. Communications in Statistics—Simulation and Computation, 49(4), 973–988. https://doi.org/10.1080/03610918.2018.1441415