검색 상세

멀티 에이전트 기반 자동 생성된 중학교 영어 문항의 타당성 연구 : 읽기 영역을 중심으로

A Study on the Validity of Multi-Agent-Based Automatically Generated Middle School English Items : Focusing on the Reading Domain

초록(요약문)

본 연구는 생성형 인공지능을 활용한 중학교 영어 읽기 문항 자동 생성에 단일 LLM 기반 방식과 멀티 에이전트 기반 방식을 적용하고, 두 방식으로 생성된 문항의 타당성을 비교·분석하는 데 목적이 있다. 단일 프롬프트만으로는 여러 제약 조건을 안정적으로 충족하기 어렵다는 기존 연구의 한계에 주목하여, 본 연구는 문항 생성, 어휘 검토 및 수정, 성취기준 검토 및 수정, 최종 검토의 네 단계로 구성된 멀티 에이전트 파이프라인을 설계하였다. 2022 개정 교육과정에 따라 개발된 중학교 영어 1 교과서의 읽기 지문과 읽기 영역 성취기준(9영01-02~9영01-07)을 입력 자료로 사용하였으며, 두 방식에 동일한 입력 자료와 문항 생성 조건을 적용하였다. 기반 언어 모델로 GPT-5.4를 사용하였고, 멀티 에이전트 기반 방식은 CrewAI 프레임워크를 활용하여 구현하였다. 두 방식으로 총 108개의 문항을 생성하였으며, 전체 108개 문항과 전문가 평가 대상으로 선정된 36개 문항의 어휘 준수율을 각각 분석하였다. 또한 최종 선정된 36개 문항은 중학교 영어 교사 3인과 영어 교과서 편집자 3인으로 구성된 총 6인의 전문가가 성취기준 반영 수준, 문항 품질 및 교육 현장 활용성 측면에서 평가하였다. 전체 생성 문항 108개의 어휘 준수율은 두 방식 모두 99.75%로 나타났다. 최종 선정 문항 36개에서는 멀티 에이전트 기반 방식의 어휘 준수율이 100%로 나타났으며, 단일 LLM 기반 방식은 99.92%로 나타났다. 성취기준 반영 수준은 멀티 에이전트 기반 방식이 단일 LLM 기반 방식보다 통계적으로 유의하게 높았다(3.71 vs. 3.49, p=.028). 문항의 타당성 및 교육 현장 활용성과 관련된 6개 평가 항목의 전체 평균에서도 멀티 에이전트 기반 방식이 단일 LLM 기반 방식보다 통계적으로 유의하게 높은 점수를 보였다(3.48 vs. 3.36, p=.013). 그러나 교육 현장 활용성 항목에서는 두 방식 간 통계적으로 유의한 차이가 나타나지 않았다. 이러한 결과는 멀티 에이전트 기반 방식이 문항 생성 과정을 단계별로 분리함으로써 생성된 문항을 보다 체계적으로 검토하는 데 기여할 수 있음을 시사한다. 그러나 멀티 에이전트 구조만으로 실제 교육 현장에서 바로 활용할 수 있는 수준의 문항 품질이 보장되는 것은 아니다. 따라서 생성형 인공지능 기반 자동문항생성은 교사의 문항 개발 전문성을 대체하는 기술이라기보다, 교사의 전문적 판단과 결합할 때 교육적 유용성이 높아질 수 있는 지원 도구로 이해할 필요가 있다. 이에 인간 전문가의 최종 판단을 포함하는 인간 참여형(Human-in-the-Loop, HITL) 협업 구조가 요구된다.

more

초록(요약문)

This study aimed to apply a single-LLM-based approach and a multi-agent-based approach to the automated generation of middle school English reading items, and to compare the validity of the items generated by the two approaches. Drawing on limitations identified in prior research—particularly the difficulty of reliably satisfying multiple constraints through a single prompt—this study designed a multi-agent pipeline consisting of four stages: item generation, vocabulary review and revision, achievement-standard review and revision, and final review. Reading passages from a Middle School English 1 textbook developed under the 2022 revised national curriculum, along with the reading-domain achievement standards (9영01-02 to 9영01-07), were used as input materials. Identical input materials and generation conditions were provided to both approaches. GPT-5.4 was used as the underlying language model, and the multi-agent system was implemented using the CrewAI framework. A total of 108 items were generated. Vocabulary compliance was analyzed for all 108 items and separately for the 36 items selected for expert evaluation. The selected items were then evaluated by six experts—three middle school English teachers and three English textbook editors—in terms of achievement-standard alignment, item quality, and classroom applicability. Across all 108 generated items, both approaches achieved a vocabulary compliance rate of 99.75%. Among the 36 selected items, the multi-agent-based approach achieved a vocabulary compliance rate of 100%, compared with 99.92% for the single-LLM-based approach. The multi-agent-based approach received significantly higher scores for achievement-standard alignment than the single-LLM-based approach (3.71 vs. 3.49, p = .028). It also received a significantly higher overall score across the six criteria related to item quality and classroom applicability (3.48 vs. 3.36, p = .013). However, no statistically significant difference was found between the two approaches in classroom applicability as an individual criterion. These findings suggest that the multi-agent-based approach, by dividing the item-generation process into discrete stages, can support more systematic review of generated items. However, a multi-agent structure alone does not guarantee a level of item quality suitable for direct classroom use. Generative AI-based automated item generation should therefore be understood not as a replacement for teachers' item-development expertise, but as a tool whose educational usefulness increases when combined with teachers' professional judgment. Accordingly, a Human-in-the-Loop (HITL) collaborative structure incorporating the final judgment of human experts is required.

more

목차

Ⅰ. 서 론 1
1.1. 연구의 필요성과 목적 1
1.2. 연구 문제 3
Ⅱ. 이론적 배경 4
2.1. 자동문항생성 4
2.1.1. 거대언어모델 기반 자동문항생성 5
2.1.2. 에이전트 기반 자동문항생성 7
2.2. 중학교 영어과 교육과정에 대한 평가 8
2.2.1. 성취기준 9
2.2.2. 어휘 지침 10
2.3. 프롬프트 엔지니어링 12
Ⅲ. 연구 방법 14
3.1. 분석 대상 15
3.1.1. 분석 단원 선정 및 구성 15
3.1.2. 성취기준 15
3.1.3. 어휘 제한 기준 17
3.2. 실험 설계 17
3.2.1. 문항 생성 조건 통제 19
3.2.2. 단일 LLM과 멀티 에이전트 기반 프롬프트 설계 19
3.2.3. 최종 평가 문항 선정 23
3.3. 타당성 평가 24
3.3.1. 어휘 제한 준수율 분석 24
3.3.2. 전문가 평가 25
Ⅳ. 연구 결과 28
4.1. 문항 생성 결과 28
4.2. 교과서 어휘 제한 준수 수준 비교 결과 29
4.2.1. 전체 생성 문항의 어휘 준수율 비교 결과 30
4.2.2. 최종 선정 문항의 어휘 준수율 비교 결과 31
4.3. 전문가 평가 분석 결과 32
4.3.1. 평가자 간 신뢰도 32
4.3.2. 교육과정 성취기준 반영 수준 비교 결과 34
4.3.3. 문항의 타당성 및 교육 현장 활용성 비교 결과 35
4.3.4. 전문가 서술형 의견 분석 37
Ⅴ. 결론 및 제언 41
5.1. 요약 및 결론 41
5.2. 논의 및 제언 44
참고문헌 46
국문초록 51
부록 53

more