검색 상세

Training-Free Movie Audio Description Generation via a Semantic Exemplar Retrieval-Based Information Selection Framework

초록(요약문)

화면해설은 시청각 매체 안의 중요한 시각 정보를 내레이션을 통해 전달하는 접근성 서비스이다. 높은 품질의 화면해설을 작성하려면 상당한 전문성과 제작 시간이 요구되므로, 영화와 같은 장편 영상 콘텐츠에 대해 화면해설을 자동으로 생성하는 기술은 미디어 접근성을 확대하는 데 중요하다. 그러나 자동 영화 화면해설 생성은 여전히 어려운 과제이다. 화면해설 문장은 제한된 내레이션 구간 안에 들어가야 하므로 간결해야 하며, 동시에 관련 인물, 행동, 객체, 자세, 사건 결과와 같은 시각적으로 중요한 정보를 선택적으로 보존해야 한다. 최근 자동 화면해설 생성 연구에서는 별도의 과제별 미세조정 없이 사전학습된 시각-언어 모델과 대형 언어 모델을 활용하는 추가 학습 없는 두 단계 생성 구조가 활용되고 있다. 이 구조에서는 첫 번째 단계에서 시각-언어 모델이 영화 클립에 대한 자세한 시각 설명을 생성하고, 두 번째 단계에서 대형 언어 모델이 이를 한 문장의 화면해설로 변환한다. 그러나 두 번째 단계의 언어 모델은 여전히 자세한 시각 설명 안의 여러 정보 중 어떤 정보를 최종 문장에 남겨야 하는지를 결정해야 한다. 본 논문은 추가 학습 없는 영화 화면해설 생성에서 발생하는 두 번째 단계의 정보 선택 문제를 다룬다. 제안 방법은 현재 자세한 시각 설명의 행동 관련 내용과 의미적으로 관련 있는 작가 작성 화면해설 예시를 검색하고, 이를 두 번째 단계의 언어 모델에 문맥 내 예시로 제공한다. 이 방법은 시각-언어 모델이나 언어 모델을 미세조정하지 않는다. 대신 의미적으로 유사한 시각 사건이 작가 작성 화면해설 문장에서 어떻게 표현되는지를 보여줌으로써, 최종 화면해설 문장에 포함할 시각 정보를 선택하고 표현하는 과정을 안내한다. 제안 방법은 MAD-Eval에서 네 가지 예시 조건, 즉 예시 없음, 고정 예시, 무작위 예시, 의미 기반 예시 조건으로 평가하였다. 모든 조건은 동일한 평가 사례, 동일한 자세한 시각 설명, 최종 화면해설 생성을 담당하는 동일한 언어 모델, 동일한 디코딩 설정, 그리고 동일한 화면해설 생성 지시문과 프롬프트 형식을 사용하며, 예시 블록만 다르다. 실험 결과, 의미 기반 예시 조건은 비교한 조건 중 가장 높은 CIDEr 점수를 보였다. CIDEr는 예시 없음 조건의 23.0에서 의미 기반 예시 조건의 25.3으로 향상되었으며, 이는 10.00%의 상대 향상에 해당한다. 이는 생성된 화면해설 문장이 정답 화면해설의 핵심 정보와 더 잘 겹쳤을 가능성을 시사한다. 관련 MAD-Eval 결과와의 비교는 제안 방법이 기존 보고 결과들 사이에서 어느 정도의 성능 수준에 있는지 보여주기 위한 참고 자료로 제시하였으며, 서로 다른 시스템 간의 통제된 성능 순위를 주장하기 위한 것은 아니다. 정성 분석은 의미 기반 예시가 입력 조건부 안내로 작동할 수 있음을 보여준다. 대표 사례들은 의미 기반 예시 조건이 다른 예시 조건보다 자세한 시각 설명 안의 정답 관련 사건 패턴을 더 효과적으로 보존하는 경우를 보여준다. 이는 의미적으로 검색된 화면해설 예시가 언어 모델이 정답과 관련 있는 시각 정보를 선택하고 표현하는 데 도움을 줄 수 있음을 시사한다. 제안 방법에는 명확한 한계가 있다. 이 방법은 첫 번째 단계에서 생성된 자세한 시각 설명의 품질에 의존하며, 여러 그럴듯한 사건이 함께 존재할 때 의미 기반 예시 검색이 항상 올바른 중요도 선택을 보장하지는 않는다. 그럼에도 본 논문의 결과는 의미 기반 예시 검색이 일반 언어 모델의 화면해설식 정보 선택을 더 적절한 방향으로 유도함으로써, 추가 학습 없는 영화 화면해설 생성의 두 번째 단계 생성 과정을 지원할 수 있음을 보여준다.

more

초록(요약문)

Audio description(AD) is an accessibility service that conveys important visual information in audiovisual media through narration. Since writing high-quality AD requires substantial expertise and production time, automatic AD generation for long-form video content such as movies is important for expanding media accessibility. However, automatic movie AD generation remains challenging. An AD sentence must fit within a limited narration interval and therefore needs to be concise, while selectively preserving visually important information such as relevant characters, actions, objects, postures, and event outcomes. Recent studies on automatic AD generation have explored training-free two-stage generation frameworks that use pretrained visual-language models and large language models without task-specific fine-tuning. In this structure, the first-stage visual-language model generates a dense visual description of a movie clip, and the second-stage large language model converts it into a one-sentence AD. However, the second-stage language model must still decide which information from the dense visual description should remain in the final sentence. This thesis addresses the Stage II information selection problem in training-free movie AD generation. The proposed method retrieves human-written AD exemplars that are semantically relevant to the action-related content of the current dense visual description and provides them to the second-stage language model as in-context exemplars. The method does not fine-tune either the visual-language model or the language model. Instead, it guides the selection and phrasing of visual information for the final AD sentence by showing how semantically similar visual events are expressed in human-written AD sentences. The proposed method is evaluated on MAD-Eval under four exemplar settings: no- exemplar, fixed-exemplar, random-exemplar, and semantic-exemplar. All settings use the same evaluation instances, dense visual descriptions, language model for final AD generation, decoding configuration, and AD generation instruction and prompt format, except for the exemplar block. The experimental results show that the semantic-exemplar setting obtains the highest CIDEr score among the compared settings. CIDEr improves from 23.0 in the no-exemplar setting to 25.3 in the semantic-exemplar setting, corresponding to a 10.00% relative improvement. This suggests that the generated AD sentences show greater overlap with key information in the reference AD. Comparisons with related MAD-Eval results are presented only as a reference for situating the proposed method among previously reported results, and are not intended to claim a controlled performance ranking across different systems. Qualitative analysis shows that semantic exemplars can function as input-conditioned guidance. Representative examples demonstrate cases in which the semantic-exemplar setting preserves reference-relevant event patterns from the dense visual description more effectively than the other exemplar settings. This suggests that semantically retrieved AD exemplars can help the language model select and phrase visual information that is relevant to the reference AD. The proposed method has clear limitations. It depends on the quality of the first stage dense visual description, and semantic exemplar retrieval does not always guarantee correct salience selection when multiple plausible events are present. Nevertheless, the results show that semantic exemplar retrieval can support the second-stage generation process in training-free movie AD generation by guiding a general language model toward more appropriate AD-style information selection.

more

목차

I. Introduction 1
II. Related Work 5
2.1 Automatic Movie Audio Description Generation 5
2.2 Training-Free AD Generation with VLMs and LLMs 7
2.3 In-Context Learning and Exemplar Selection 8
III. Preliminaries 10
3.1 Visual-Language Models 10
3.2 Large Language Models and Prompt-Based Generation 11
3.3 In-Context Exemplars 11
3.4 Sentence Embeddings and Semantic Similarity 12
IV. Method 13
4.1 Overall Framework 13
4.2 Dense Visual Description Generation 15
4.3 Action-Related Query Extraction 17
4.4 Semantic Audio Description Exemplar Retrieval 18
4.5 Audio Description Generation with Retrieved Exemplars 20
V. Experiments 22
5.1 Dataset 22
5.2 Experimental Settings 22
5.2.1 Compared Settings 22
5.2.2 Implementation Details 24
5.3 Evaluation Metrics 26
5.4 Results and Analysis 27
5.4.1 Quantitative Results 27
5.4.2 Comparison with Prior Methods 28
5.4.3 Qualitative Analysis 31
VI. Discussion 34
6.1 Interpretation of Quantitative Results 34
6.2 Semantic Exemplars as Guidance for Information Selection 35
6.3 Position Relative to Prior Methods 36
6.4 Limitations 37
VII. Conclusion 41
A. Additional Qualitative Analyses 43
A.1 Failure Cases Relative to Prior Methods 43
A.2 Detailed Stage I-to-Stage II Case Studies 45

more