검색 상세

급성 림프모구성 백혈병 아형 분류를 위한 전이학습 기반 딥러닝 모델의 성능-효율 비교

Performance-Efficiency Comparison of Transfer Learning-Based Deep Learning Models for Acute Lymphoblastic Leukemia Subtype Classification

초록(요약문)

급성 림프모구성 백혈병(Acute Lymphoblastic Leukemia, ALL)은 신속한 형태학적 판독과 아형 지향적 해석이 중요한 혈액암이다. 그러나 고도화된 검사 장비와 전문 인력에 대한 접근성이 제한되는 자원 제약 의료환경에서는 일관된 영상 해석을 유지하기가 쉽지 않다. 본 연구는 임상 진단 체계를 대체하거나 임상 적용성을 입증하려는 것이 아니라, 공개 데이터셋 기반의 통제된 비교 실험을 통해 조기 선별과 보조 판독 관점에서 활용 가능한 성능-효율 비교 기준을 제시하는 데 목적이 있다. 본 연구는 Kaggle ALL 이미지 데이터셋[7]의 말초혈액도말 이미지 3,256장(4개 클래스: Benign, Early, Pre, Pro)을 사용하여 ResNet50, DenseNet121, EfficientNetB0, MobileNetV2, NASNetMobile, Big Transfer(BiT), Vision Transformer(ViT)의 7개 전이학습 기반 딥러닝 모델을 비교하였다. 각 모델은 Raw_Frozen, Raw_Tuned, CLAHE_Frozen, CLAHE_Tuned의 4가지 조건에서 층화 5-겹 교차검증으로 평가되었으며, 총 140회의 실험이 수행되었다. 5개 CNN은 TensorFlow/Keras로, BiT와 ViT는 PyTorch/timm으로 구현하였으므로 추론 시간 비교는 프레임워크 차이를 고려하여 제한적으로 해석하였다. 전체 모델-조건 조합 가운데 최고 성능은 ViT의 CLAHE_Tuned 조합으로 정확도 98.19 ± 0.72%, Macro-F1 97.99 ± 0.71%를 기록하였다. BiT의 Raw_Tuned 조합은 정확도 98.03 ± 0.99%, Macro-F1 97.73 ± 1.14%, 파라미터 23.51M으로 고성능 모델군 내에서 우수한 성능-효율 균형을 보였다. 5개 CNN 기준에서는 ResNet50의 Raw_Tuned가 정확도 97.70 ± 0.76%, Macro-F1 97.39 ± 0.92%로 가장 강한 CNN 기준선이었고, MobileNetV2의 Raw_Tuned는 2.26M 파라미터에서 Macro-F1 96.39 ± 1.10%를 기록하여 가장 경량한 후보로 해석되었다. EfficientNetB0의 CLAHE_Tuned는 Macro-F1 96.59 ± 0.36%로 fold 간 변동성이 가장 낮았다. 이 결과는 ALL 이미지 분류에서 전처리와 미세조정의 효과를 개별 요소가 아니라 모델 구조와 함께 해석해야 함을 보여준다. 종합적으로 ViT CLAHE_Tuned는 최고 성능, BiT Raw_Tuned는 고성능 모델군 내 성능-효율 균형, ResNet50 Raw_Tuned는 CNN 기준선, MobileNetV2 Raw_Tuned는 경량 후보, EfficientNetB0 CLAHE_Tuned는 안정성 측면의 장점을 보였다. 다만 환자 단위 분할 정보 부재, 외부 검증 미수행, 그리고 TensorFlow/Keras와 PyTorch/timm 간 구현 차이로 인해 본 결과를 임상 일반화 성능으로 확대 해석하기는 어렵다.

more

초록(요약문)

Acute lymphoblastic leukemia (ALL) requires timely morphological assessment and subtype-oriented interpretation. However, consistent image interpretation may be difficult in resource-constrained medical environments where specialized equipment and expert review are not always available. This study does not claim clinical replacement or direct diagnostic utility; instead, it provides a controlled benchmark for early screening and auxiliary reading scenarios using a public image dataset. Using the Kaggle ALL image dataset [7] containing 3,256 peripheral blood smear images from four classes (Benign, Early, Pre, and Pro), this study compared seven transfer-learning-based deep learning models: ResNet50, DenseNet121, EfficientNetB0, MobileNetV2, NASNetMobile, Big Transfer (BiT), and Vision Transformer (ViT). Four conditions were evaluated for each model (Raw_Frozen, Raw_Tuned, CLAHE_Frozen, and CLAHE_Tuned) under stratified 5-fold cross-validation, yielding 140 experiments in total. The five CNN models were implemented in TensorFlow/Keras, whereas BiT and ViT were implemented in PyTorch/timm; this framework difference was considered when interpreting inference-time results. Among all model-condition pairs, ViT with CLAHE_Tuned achieved the highest average performance with 98.19 ± 0.72% accuracy and 97.99 ± 0.71% Macro-F1. BiT with Raw_Tuned recorded 98.03 ± 0.99% accuracy and 97.73 ± 1.14% Macro-F1 with 23.51 million parameters, showing a favorable parameter-efficiency balance among high-performing models. Within the five-CNN subset, ResNet50 with Raw_Tuned remained the strongest CNN baseline (97.70 ± 0.76% accuracy and 97.39 ± 0.92% Macro-F1), MobileNetV2 with Raw_Tuned was the lightest candidate (2.26M parameters, 96.39 ± 1.10% Macro-F1), and EfficientNetB0 with CLAHE_Tuned showed the lowest fold-to-fold variation (96.59 ± 0.36% Macro-F1). These results indicate that model selection, preprocessing, and fine-tuning should be interpreted jointly rather than independently. ViT provided the highest predictive performance, BiT offered a strong high-performance efficiency balance, ResNet50 served as the most competitive CNN baseline, and MobileNetV2 showed that lightweight models can retain meaningful discriminative capacity. Because patient-wise splitting was unavailable and external validation was not conducted, the findings should be interpreted as comparative benchmark evidence rather than proof of clinical applicability.

more

목차

제 1 장 서론 1
제 1 절 연구 배경 및 동기 1
제 2 절 연구 목적 및 연구질문 3
제 3 절 논문의 구성 4
제 4 절 연구의 기여 4

제 2 장 관련 연구 5
제 1 절 급성 림프모구성 백혈병(ALL) 개요 5
제 2 절 국내외 의료 영상 분류 연구 동향 5
제 3 절 전이학습 비교 모델 아키텍처 및 이론적 배경 7
제 4 절 백혈병 분류 관련 선행 연구 9
제 5 절 의료 AI 검증 원칙 10

제 3 장 연구 방법 및 실험 설계 11
제 1 절 병리학적 특징 강화를 위한 데이터 전처리 11
제 2 절 전이학습 모델 비교를 위한 분류기 헤드 설계 13
제 3 절 실험 조건 설계 및 학습 전략 15

제 4 장 실험 환경 및 성능 평가 20
제 1 절 교차검증 프로토콜 20
제 2 절 실험 환경의 재현성 확보 20
제 3 절 평가 지표 21
제 4 절 비교 설계의 논리 21

제 5 장 결과 및 논의 23
제 1 절 실험 개요 및 조건별 평균 성능 23
제 2 절 모델별 최적 조합 및 성능 상세 24
제 3 절 성능-효율 트레이드오프 분석 25
제 4 절 성능 차이와 변동성의 해석 27
제 5 절 종합 논의 28

제 6 장 결론 및 향후 연구 29
제 1 절 연구 요약 및 시사점 29
제 2 절 향후 연구 방향 30

참고문헌 33
부록 36
부록 A. 실험 재현성 체크리스트 36
부록 B. Fold별 상세 결과 36
부록 C. 하이퍼파라미터 조합 결과 37
부록 D. 핵심 알고리즘 및 학습 파이프라인 37
부록 E. 실험 환경 재현 코드 43

more