검색 상세

지식그래프 기반 한국어 채용공고 숙련 추출 프레임워크 : IT 직종을 중심으로

A Knowledge Graph-Based Framework for Skill Extraction from Korean Job Postings; Focusing on IT Occupations

초록(요약문)

인공지능 기술의 확산으로 노동시장에서 요구되는 숙련은 빠르게 재편되고 있다. 이러한 변화를 지속적으로 관측하려면, 채용공고에 나타나는 숙련 수요를 일관된 기준으로 추출·정렬하고 축적할 수 있는 프레임워크가 필요하다. 그러나 채용공고 기반 숙련 추출에는 세 가지 어려움이 있다. 첫째, 채용공고는 기업 소개, 복리후생 등 직무와 무관한 정보가 혼재된 비정형 텍스트로, 분석 과정에서 노이즈가 발생한다. 둘째, 벡터 검색 기반 접근은 분류 위계와 직업–숙련 관계를 명시적으로 활용하기 어렵고, 분절되어 운용되어 온 산업·직업·숙련 분류체계 위에 결과를 일관되게 적재하기도 어렵다. 셋째, 표준 벤치마크가 부재하고 숙련의 경계가 맥락에 의존하여 신뢰할 수 있는 골드 데이터셋을 구축하기 어렵다. 본 연구는 한국어 IT 채용공고를 대상으로 지식그래프와 LLM 추론을 결합한 숙련 추출·축적 프레임워크를 제안한다. 비정형 원문을 섹션별 중간표현(Posting IR)으로 구조화해 노이즈를 분리한다. 한국표준산업분류(KSIC), 한국고용직업분류(KECO), 한국숙련사전(KSD) 및 직업–숙련 연계 관계를 참조 지식그래프(Reference KG)로 통합해 산업·직업·숙련 후보를 구조적 맥락에서 생성·판정한다. LLM은 원문 근거와 후보의 분류 경로를 대조하며, 근거가 부족하면 코드를 확정하지 않는다. 판정 코드는 원문 근거·메타데이터와 함께 채용공고 지식그래프(Job Posting KG)에 추적 가능하게 적재된다. 전문가 어노테이션 기반 파일럿 골드 평가셋(N=42)을 기준으로 벡터 검색 기반 RAG와 제안한 지식그래프 기반 추론 방식을 비교하였다. 중간표현은 산업 판단 근거가 존재하는 공고의 KSIC 정확도를 0.692에서 0.923으로 향상시켰다. 제안 방식은 주 평가 지표인 KSD 소분류 수준에서 F1 0.679를 기록하여 RAG 대비 정밀도와 재현율을 함께 개선하였다. 반면 직업분류에서는 조건 간 차이가 나타나지 않아, 각 구성요소의 효과가 과제의 특성에 따라 다름을 확인하였다. 약 8,000건의 공고 분석 결과는 Job Posting KG에 축적되어 지역·직종·시점별 숙련 수요의 동향 파악을 지원한다. 다만 평가는 파일럿 규모와 IT 직종에 한정되므로 확대 검증이 향후 과제로 남는다. 그럼에도 본 연구는 비정형 공고의 구조화부터 근거 기반 추론과 관측 데이터의 지식그래프 적재까지를 하나의 파이프라인으로 연결하여, 노동시장 숙련 수요를 비교 가능한 형태로 축적할 수 있는 기반을 제시하였다.

more

초록(요약문)

The rapid diffusion of artificial intelligence is reshaping the skills demanded in the labor market. Monitoring these changes requires a framework that can extract, align, and accumulate skill demand from job postings using consistent criteria. This task presents three challenges: job postings are unstructured texts mixed with job-irrelevant information, which introduces noise; vector retrieval-based approaches have difficulty representing classification hierarchies and occupation–skill relationships, while fragmented classification systems hinder consistent accumulation; and the absence of a standard benchmark makes it difficult to construct a reliable gold dataset. This study proposes a framework that combines a knowledge graph with LLM reasoning to extract and accumulate skill-demand data from Korean IT job postings. The raw text is transformed into a section-level intermediate representation (Posting IR) to separate noise. The Korean Standard Industrial Classification (KSIC), Korean Employment Classification of Occupations (KECO), Korean Skills Dictionary (KSD), and occupation–skill relationships are integrated into a reference knowledge graph (Reference KG), within which classification candidates are generated and evaluated. The LLM compares textual evidence with candidates' classification paths and abstains from assigning a code when evidence is insufficient; the resulting codes are stored in a Job Posting Knowledge Graph (Job Posting KG) with their evidence and metadata. Compared with vector retrieval-based RAG on an expert-annotated pilot gold set (N=42), the intermediate representation improved KSIC accuracy from 0.692 to 0.923 for postings containing industry-relevant evidence, and the proposed approach achieved an F1 score of 0.679 at the KSD subcategory level, the primary metric, improving both precision and recall. In occupational classification, however, no differences were observed, indicating that component effects vary by task. Results from approximately 8,000 postings accumulate in the Job Posting KG, supporting trend monitoring by region, occupation, and time. Although limited to a pilot-scale evaluation and IT occupations, this study connects unstructured-posting structuring, evidence-grounded reasoning, and knowledge graph storage within a single pipeline, providing a foundation for accumulating skill demand in a comparable form.

more

목차

제1장 서론 1
제1절 연구 배경 및 필요성 1
제2절 연구 목적 및 연구 질문 3
제3절 연구 방법 및 평가 방향 3
제4절 연구 범위 4
제2장 선행연구 및 이론적 배경 5
제1절 채용공고 기반 숙련추출 선행연구 흐름 5
제2절 이론적 배경 7
제3절 지식그래프와 지식그래프 기반 추론 8
제4절 프롬프트 기반 단계적 추론과 중간표현 9
제5절 근거 기반 평가와 인간 주석 9
제6절 소결 13
제3장 1단계: 참조 지식그래프(Reference KG) 구축 14
제1절 참조 지식그래프의 역할 14
제2절 분류체계 15
제3절 분류체계별 지식그래프 구축 16
제4절 KECO-KSD 20
제5절 Neo 4j 스키마와 구현 21
제6절 구성 통계 22
제 4장 2단계:채용공고 지식그래프(Job Posting KG) 구축 프레임워크 23
제1절 프레임워크 개요 23
제2절 연구 데이터셋 25
제3절 규칙 기반 파싱(Rulebase Parsing) 및 Posting IR 25
제4절 질의(Query View)와 KG 기반 후보검색 28
제5절 LLM 근거 기반 판정 30
제6절 Load IR과 채용공고 지식그래프(Job Posting KG) 적재 스키마 32
제7절 채용공고 지식그래프(Job Posting KG) 활용례 34
제5장 실험설계 및 평가 37
제1절 실험 설계 개요 37
제2절 골드 평가셋 구축 38
제3절 실험 환경 및 구현 42
제4절 평가 지표 43
제5절 실험 결과 45
제6절 오류 분석 및 진단 49
제6장 결론 및 평가 55
제1절 연구 요약 55
제2절 연구의 기여 56
제3절 연구의 한계 57
제4절 향후 연구 59
참고문헌 61
부 록 64
부록 A. 정보통신 공학기술직 KECO(2025) 코드 및 허용 대분류 승계(KECO, KSD) 64
부록 B. 전체출력 결과표 65
부록 C. 골드 데이터셋 구축과정 및 루브릭 66
부록 D. 지식그래프 구축 전처리 절차 및 프롬프트 73
부록 E. 산업직업숙련 추론 프롬프트 전문 97
부록 F. 분류체계 미매핑 표현 관측 사례 106

more