Robust Subgroup Analysis for Heterogeneous Censored Data

· Source: stat.ML updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics, Mathematics & Computational Sciences · Depth: Expert, extended

Summary

A novel robust subgroup analysis approach, RISA-ADMM, is proposed for heterogeneous censored data under accelerated failure time (AFT) models. This method integrates inverse probability weighting to handle censoring, M-estimation for robustness against heavy-tailed errors and outliers, and concave pairwise fusion penalization for automatic subgroup identification and covariate effect estimation without requiring prior knowledge of subgroup memberships. Extensive simulations demonstrate RISA-ADMM's superior performance over the imputation-based BJ-ADMM, particularly under heteroscedastic and heavy-tailed error distributions (Cases 2 and 3), where it shows greater robustness, lower variability, and more accurate subgroup and parameter estimation. For instance, in the German credit dataset, RISA-ADMM correctly identified two subgroups, unlike BJ-ADMM which found none, leading to more significant covariate identification. The method also extends to diverging-dimensional settings where covariate dimensions can grow with sample size.

Key takeaway

For Data Scientists analyzing time-to-event data with censoring and suspected population heterogeneity, RISA-ADMM offers a robust alternative to imputation-based methods. You should consider this approach, especially when dealing with heavy-tailed error distributions or outliers, as it provides more accurate subgroup identification and parameter estimation. This can lead to more reliable insights for targeted decision-making in areas like credit risk or public health, where understanding distinct subpopulations is critical.

Key insights

RISA-ADMM robustly identifies data subgroups and estimates covariate effects in censored, heterogeneous data, even with heavy-tailed errors.

Principles

Method

RISA-ADMM combines inverse probability weighting, M-estimation (e.g., Huber loss), and concave pairwise fusion penalization. It uses an ADMM algorithm with iterative penalized reweighted least squares (IPRLS) for optimization and modified BIC for parameter selection.

In practice

Topics

Best for: AI Scientist, Research Scientist, Data Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by stat.ML updates on arXiv.org.