멀티라벨 영화 장르 분류를 위한 멀티모달 융합 전략 비교 연구
A Comparative Study of Multimodal Fusion Strategies for Multi-Label Movie Genre Classification
- 주제(키워드) 멀티모달 학습 , 멀티라벨 분류 , 영화 장르 분류 , 융합 전략 , Late Fusion , multimodal learning , multi-label classification , movie genre classification , fusion strategy , Late Fusion
- 발행기관 서강대학교 AI.SW대학원
- 지도교수 박운상
- 발행년도 2026
- 학위수여년월 2026. 8
- 학위명 석사
- 학과 및 전공 AI.SW대학원 데이터사이언스 · 인공지능
- 세부분야 해당없음
- 실제URI http://www.dcollection.net/handler/sogang/000000082980
- UCI I804:11029-000000082980
- 본문언어 한국어
- 저작권 논문은 저작권에 의해 보호받습니다.
초록(요약문)
영화 장르 분류는 한 영화에 둘 이상의 장르가 부여되는 멀티라벨 분류 문제로, 멀티모달 접근이 단일 모달리티 대비 향상된 성능을 보고해 왔다. 다 만, 어떤 융합 전략이 가장 효과적인가에 대해서는 선행 연구 간 결과가 일치 하지 않으며 동일한 데이터 환경에서 융합 전략 간 상대적 효과를 통제된 형태로 비교한 사례가 충분하지 않다. 본 연구는 The Movie Database(TMDb)로부터 수집한 6,446편의 영화와 19개 장르로 구성된 멀티라벨 분류 데이터셋을 구축하고 텍스트 인코더 TF-IDF, BERT-mini, DistilBERT와 이미지 인코더 ResNet50, ConvNeXt-Tiny, 융합 전략 Late Fusion, Early Fusion, Gated Fusion의 조합을 통제된 환경에서 비교하였다. 평가 지표로 Micro-F1과 Macro-F1을, 임계값 전략으로 고정 임계값과 장르별 임계값을 함께 적용하였다. 실험 결과, DistilBERT와 ConvNeXt-Tiny 조합에 가중 평균 Late Fusion(α=0.35)을 적용한 구성이 Micro-F1 0.6463으로 최고 성능을 기록하였으며 부트스트랩 유의성 검정을 통해 단일 모달 및 다른 융합 전략 대비 우위가 통 계적으로 유의함이 확인되었다. 융합 전략의 복잡도가 증가할수록 성능이 저하되는 양상이 관찰되었으며 그 중 Gated Fusion은 단일 모달 DistilBERT 성능 에도 미치지 못하였다. 모달리티별 장르 우위 분석에서는 두 모달리티가 서로 다른 장르 집합에서 강점을 가지는 상보성이 관찰되었으며 이는 융합 이점의 주요 원인으로 해석할 수 있다.
more초록(요약문)
Movie genre classification is a multi-label classification problem in which a single film is assigned two or more genres. Multimodal approaches have reported improved performance over single-modality approaches. However, prior studies have not reached consensus on which fusion strategy is most effective, and controlled comparisons of the relative effectiveness of fusion strategies under the same data environment remain insufficient. This study constructs a multi-label classification dataset of 6,446 movies and 19 genres collected from The Movie Database(TMDb), and compares combinations of text encoders(TF-IDF, BERT-mini, DistilBERT), image encoders(ResNet50, ConvNeXt-Tiny), and fusion strategies(Late Fusion, Early Fusion, Gated Fusion) under a controlled setting. Performance is evaluated using Micro-F1 and Macro-F1, with a fixed threshold and a per- label threshold jointly applied for thresholding. The experimental results show that the configuration combining DistilBERT and ConvNeXt-Tiny with weighted-average Late Fusion(α=0.35) achieves the best performance at Micro-F1 0.6463, and a bootstrap significance test confirms that its advantage over single-modality models and other fusion strategies is statistically significant. Performance is observed to degrade as the complexity of the fusion strategy increases, and Gated Fusion fails to surpass even the single-modality DistilBERT baseline. The per-modality genre advantage analysis reveals complementarity between the two modalities, with each excelling on different subsets of genres, which can be interpreted as a primary source of the benefit obtained from fusion.
more목차
제 1 장 서론
제 1 절 연구 배경 10
제 2 절 연구 목적 및 범위 12
제 2 장 관련 연구
제 1 절 텍스트 기반 접근법 15
제 2 절 이미지 기반 접근법 17
제 3 절 멀티모달 접근법 18
제 3 장 데이터셋
제 1 절 데이터셋 개요 23
제 2 절 데이터 전처리 24
제 3 절 장르별 분포 26
제 4 장 실험 설정
제 1 절 전체 시스템 구조 30
제 2 절 단일 모달 인코더 31
(1) 텍스트 인코더 31
(2) 이미지 인코더 33
제 3 절 융합 전략 36
(1) Late Fusion 37
(2) Early Fusion 38
(3) Gated Fusion 39
제 4 절 평가 방법 40
(1) 평가 지표 40
(2) 임계값 전략 41
(3) 통계적 유의성 검정 42
제 5 장 실험 결과
제 1 절 단일 모달리티 비교 44
(1) 텍스트 인코더 비교 44
(2) 이미지 인코더 비교 46
제 2 절 융합 전략 비교 및 분석 47
(1) 융합 전략별 성능 비교 48
(2) α 민감도 분석 49
(3) Val→Test 일반화 안정성 51
(4) 통계적 유의성 검정 53
(5) Gated Fusion 성능 저하 분석 54
(6) 텍스트 인코더 확장 — Late Fusion 이점의 일반성 56
제 3 절 모달리티별 장르 우위 분석 57
제 4 절 오분류 사례 분석 60
제 6 장 결론
제 1 절 연구 요약 65
제 2 절 한계점 68
제 3 절 향후 연구 방향 69
참고문헌 71

