Research question
Can hierarchical co-occurrence structure be built into logistic regression while keeping the resulting risk model inspectable?
Condensed abstract
Formal Concept Analysis extracts closed itemsets from binarized health indicators. A closure penalty then encourages coefficients belonging to the same concepts to remain coherent while preserving a linear prediction function.
Key contribution
An interpretability-by-design objective that embeds FCA-derived structure directly into logistic-regression training, evaluated against linear and tree-ensemble baselines on a large survey-derived cohort.
My contribution
The published contribution statement credits Arman Salehi with conceptualization, methodology, coding, formal analysis, and writing the original draft.
Method overview
Closed itemsets are extracted from binarized indicators. The model adds a weighted within-concept coefficient-variance term to logistic loss; closure strength and minimum support are selected through cross-validation.
Interactive method figure
Structure becomes a training constraint
Illustrative coefficient coherence
Higher is better.
Dataset and evaluation
Data
Heart Disease Health Indicators derived from the 2015 Behavioral Risk Factor Surveillance System, with approximately 380,000 complete survey records after preprocessing.
Protocol
An 80/20 stratified split was used. Five-fold cross-validation within the training data selected closure strength and minimum support; the held-out test set was evaluated with discrimination, classification, imbalance-sensitive, and calibration metrics.
Main results
The held-out test set results were accuracy 0.906, AUC 0.810, precision 0.709, recall 0.544, F1 0.556, PR-AUC 0.265, and Brier score 0.078.
Baseline comparison
At the reported threshold, the FCA-constrained model had the highest precision and F1 among the listed baselines. Logistic regression and Gradient Boosting had higher AUC; Gradient Boosting also had the lower Brier score (0.0715 versus 0.078).
Limitations and failure modes
- BRFSS variables are self-reported, observational, and not a prospective clinical cohort.
- Binarization improves inspectability but discards information.
- The evaluation uses a single U.S. survey-derived dataset.
- External, subgroup, prospective, and workflow validation remain necessary before clinical use.
Reproducibility resources
The paper-associated computational material is available as a versioned Zenodo artifact. The source data are described at Heart Disease Health Indicators (BRFSS 2015).
Verification sources
Cite this work
Arman Salehi, Ashkan Heydarian, Hamid Reza Goudarzi, Zahra Farzin Rad (2026). Interpretable heart disease risk prediction via FCA-constrained logistic regression. Health Informatics Journal. https://doi.org/10.1177/14604582261444612
BibTeX
@article{salehi2026fca,
title = {Interpretable heart disease risk prediction via FCA-constrained logistic regression},
author = {Arman Salehi and Ashkan Heydarian and Hamid Reza Goudarzi and Zahra Farzin Rad},
year = {2026},
doi = {10.1177/14604582261444612},
journal = {Health Informatics Journal},
publisher = {SAGE Publications}
}Publication record last verified 2026-07-26.