Automatic Epicardial Adipose Tissue Segmentation in CT Using 3D U-Net

Jose Antonio Rangel Ramos1, Ravi Vazirani2, morena raiola2, Alvaro Navarro3, Javier Sanchez3, Borja Ibañez4
1Centro Nacional de Investigaciones Cardiovasculares Carlos III (CNIC), 2CNIC, 3Philips, 4Centro Nacional de Investigaciones Cardiovasculares (CNIC)


Abstract

Existing automated methods for Epicardial Adipose Tissue (EAT) segmentation in cardiac CT lack systematic validation across heterogeneous adiposity phenotypes and rarely report reproducibility against human observer standards. We present a 3D U-Net trained on 130 calcium scoring CT scans acquired on a Philips Brilliance 16-slice CT scanner from the PESA (Progression of Early Subclinical Atherosclerosis) CNIC-SANTANDER study, stratified by EAT volume into three categories (low, medium, high) across training (100) and validation (30) sets. Images were preprocessed with a Hounsfield Unit window with width 500 and level 50. The network was optimised with AdamW, cosine annealing, and a combined Dice-Focal loss. Augmentation included random affine transformations and Gaussian smoothing. Inference used sliding window with overlap. The model was benchmarked against SwinUNETR and a 3D U-Net with attention gates on the same dataset. Reproducibility was assessed on two independent test set (20 each): one for intra-observer and another one for inter-observer analysis, both providing ICC, Passing–Bablok regression and Bland-Altman.

On the test set, the 3D U-Net achieved the highest Dice (0.795) and lowest volume error (10.7%) among all architectures. SwinUNETR underperformed due to the tension between its global attention mechanism and patch-based inference, amplified by the limited training set size. The attention-gate variant showed the lowest Dice (0.750) and highest HD95 (35.17 mm), suggesting architectural complexity may hinder convergence with small datasets. Despite being the most compact architecture evaluated, the 3D U-Net achieved reproducibility within human observer range. In the reproducibility study, human inter-observer ICC = 0.979 (0.855; 0.994) with Bias = 2.45 cm³. The AI model achieved ICC = 0.987 and Bias = 1.56 cm³, falling within the inter-observer range and confirming comparable agreement to human experts. These results position the proposed pipeline as a robust and scalable alternative to manual EAT annotation in large-scale CT studies.