Research Article

Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI)

Volume: 55 Number: 3 June 30, 2026
EN

Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI)

Abstract

Standard parametric regression models for interval-censored time-to-event data frequently depend on a restrictive assumption of population homogeneity. When applied to complex, heterogeneous populations containing unobserved latent subgroups, these traditional models invariably estimate an averaged hazard function. Consequently, this homogenization systematically biases survival projections, overestimating survival times for high-risk participants while yielding overly pessimistic prognoses for low-risk individuals. To address this fundamental limitation and the known instability of traditional iterative mixture models, we introduce the prior-stratified stochastic multiple imputation algorithm, designed to disentangle latent heterogeneity and robustly impute continuous event times. Because estimating expectation-maximization mixtures from sparse interval-censored data often results in component collapse, prior-stratified stochastic multiple imputation structurally bypasses the fragile iterative expectation-maximization loop. Instead, it utilizes a one-time static Bayesian risk stratification, seeded by clinical priors, to anchor the likelihood space. This explicitly decouples the mixture into independent, strictly convex Weibull regressions, guaranteeing stable convergence even under extreme missingness. To rigorously quantify estimation uncertainty, the algorithm subsequently executes multiple repeated stochastic draws from the inferred participant-specific truncated distributions. Simulation studies demonstrate that our algorithm successfully prevents component collapse under extreme censoring (more than 40%) and structural misspecification, maintaining high phenotypic identifiability where standard unanchored mixtures fail entirely. Empirical validation utilizing a semi-synthetic primary biliary cholangitis cohort, the signal tandmobiel dental emergence dataset, and the highly sparse Finkelstein breast cancer dataset confirms the algorithm’s robust capacity to autonomously recover latent risk architectures and track non-parametric (Turnbull) ground truths without supervised labeling. Our algorithm offers a rigorously quantified methodological bridge, converting complex, heterogeneous interval-censored observations into complete datasets, thereby unlocking conventional survival analysis toolkits while safely preserving biological dimorphism.

Keywords

Supporting Institution

The University of BUrdwan

Ethical Statement

This submission was no where sent for publication or review.

References

  1. [1] D. G. Kleinbaum and M. Klein, Survival Analysis: A Self-Learning Text, Springer, New York, 1996.
  2. [2] J. Sun, Interval censoring, in Encyclopedia of Biostatistics, Wiley, New York, 3, 2090–2095, 1998.
  3. [3] K. M. Leung, R. M. Elashoff and A. A. Afifi, Censoring issues in survival analysis, Annu. Rev. Public Health 18 (1), 83–104, 1997.
  4. [4] B. W. Turnbull, Nonparametric estimation of a survivorship function with doubly censored data, J. Am. Stat. Assoc. 69 (345), 169–173, 1974.
  5. [5] D. Sinha, M. H. Chen and S. K. Ghosh, Bayesian analysis and model selection for interval-censored survival data, Biometrics 55(2), 585–590, 1999.
  6. [6] C. B. Guure, N. A. Ibrahim and M. B. Adam, Bayesian inference of the Weibull model based on interval-censored survival data, Comput. Math. Methods Med. 2013, 849520, 2013.
  7. [7] X. Lin, B. Cai, L. Wang and Z. Zhang, A Bayesian proportional hazards model for general interval-censored data, Lifetime Data Anal. 21 (3), 470–490, 2015.
  8. [8] C. Pan, B. Cai and L. Wang, A Bayesian approach for analyzing partly intervalcensored data under the proportional hazards model, Stat. Methods Med. Res. 29 (11), 3192–3204, 2020.

Details

Primary Language

English

Subjects

Biostatistics

Journal Section

Research Article

Early Pub Date

May 7, 2026

Publication Date

June 30, 2026

Submission Date

February 22, 2026

Acceptance Date

April 17, 2026

Published in Issue

Year 2026 Volume: 55 Number: 3

APA
Kapoor, S., & Gupta, A. (2026). Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI). Hacettepe Journal of Mathematics and Statistics, 55(3), 1393-1422. https://doi.org/10.15672/hujms.1892559
AMA
1.Kapoor S, Gupta A. Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI). Hacettepe Journal of Mathematics and Statistics. 2026;55(3):1393-1422. doi:10.15672/hujms.1892559
Chicago
Kapoor, Suman, and Arindam Gupta. 2026. “Stochastic Multiple Imputation for Latent Heterogeneity in Interval-Censored Data: A Prior-Stratified Approach (PS-SMI)”. Hacettepe Journal of Mathematics and Statistics 55 (3): 1393-1422. https://doi.org/10.15672/hujms.1892559.
EndNote
Kapoor S, Gupta A (June 1, 2026) Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI). Hacettepe Journal of Mathematics and Statistics 55 3 1393–1422.
IEEE
[1]S. Kapoor and A. Gupta, “Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI)”, Hacettepe Journal of Mathematics and Statistics, vol. 55, no. 3, pp. 1393–1422, June 2026, doi: 10.15672/hujms.1892559.
ISNAD
Kapoor, Suman - Gupta, Arindam. “Stochastic Multiple Imputation for Latent Heterogeneity in Interval-Censored Data: A Prior-Stratified Approach (PS-SMI)”. Hacettepe Journal of Mathematics and Statistics 55/3 (June 1, 2026): 1393-1422. https://doi.org/10.15672/hujms.1892559.
JAMA
1.Kapoor S, Gupta A. Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI). Hacettepe Journal of Mathematics and Statistics. 2026;55:1393–1422.
MLA
Kapoor, Suman, and Arindam Gupta. “Stochastic Multiple Imputation for Latent Heterogeneity in Interval-Censored Data: A Prior-Stratified Approach (PS-SMI)”. Hacettepe Journal of Mathematics and Statistics, vol. 55, no. 3, June 2026, pp. 1393-22, doi:10.15672/hujms.1892559.
Vancouver
1.Suman Kapoor, Arindam Gupta. Stochastic multiple imputation for latent heterogeneity in interval-censored data: A prior-stratified approach (PS-SMI). Hacettepe Journal of Mathematics and Statistics. 2026 Jun. 1;55(3):1393-422. doi:10.15672/hujms.1892559