검색 상세

DPU 메모리 및 병렬성을 활용한 HNSW 기반 벡터 검색 확장 기법

Scaling HNSW-Based Vector Search with DPU Memory and Parallelism

초록(요약문)

벡터 유사도 검색은 현대 AI 서비스의 핵심 구성 요소이며, HNSW는 높은 재현율 과 낮은 지연 시간으로 인해 널리 사용되고 있다. 그러나 HNSW의 메모리 집약적인 설 계는 수십억 규모 환경에서의 배포를 어렵게 만들며, 스와핑이나 원격 메모리에 의존 할경우성능이급격히저하된다.본논문은연산성능과온보드DRAM 용량이크게 향상된 최신 DPU(SmartNIC)에 주목하고, DPU를 HNSW의 확장 메모리 계층이자 병 렬 검색 엔진으로 활용하는 호스트-DPU 통합 벡터 검색 시스템인 VeX를 제안한다. VeX는 (i) 의미 구조를 보존하면서 HNSW 인덱스를 호스트와 DPU에 독립적으로 분할 및 배치하고, (ii) 이중 경로 DMA 기반 통신 설계를 통해 호스트-DPU 간 오버헤 드를 최소화하며, (iii) 이기종성을 고려한 파이프라이닝을 통해 검색, 통신, 결과 병합 단계를 중첩 수행한다. 실험 결과, 디스크 접근이 요구되는 메모리 압박 환경에서 VeX는 안정적인 Recall@100을 유지하면서 DiskANN 대비 5–10배 더 높은 처리량을 달성한다. 또한 인 덱스가 모두 메모리에 상주하는 이상적인 환경에서도, VeX는 인메모리 HNSW 대비 최대 1.9× 더 높은 질의 처리량을 보인다.

more

초록(요약문)

Vector similarity search is a core component of modern AI services, and HNSW is widely adopted due to its high recall and low latency. However, its memory-intensive design makes billion-scale deployment difficult, and performance collapses when relying on swapping or remote memory. This paper targets recent DPUs (SmartNICs) with substantially improved compute capability and onboard DRAM, and proposes VeX, a host–DPU integrated vector search system that uses the DPU as both an extended memory tier and a parallel search engine for HNSW. VeX (i) partitions and places independent HNSW indices on the host and DPU while preserving semantic structure, (ii) minimizes host–DPU overhead via a dual-path DMA-based communication design, and (iii) overlaps search, communication, and aggregation with heterogeneity-aware pipelining. Experiments show that under memory pressure requiring disk access, VeX delivers 5–10× higher throughput than DiskANN at stable Recall@100. Even in ideal settings where the index fully resides in memory, VeX outperforms in-memory HNSW by up to 1.9× in query throughput.

more

목차

그림차례 6
표차례 7
1 서론 9
2 배경및연구동기 13
2.1 Hierarchical Navigable Small Worlds (HNSW) 13
2.2 대규모 HNSW 인덱스의 메모리 압박과 Thrashing 15
2.3 DPU 기반 벡터 검색의 기회와 도전 과제 17
3 VeX의 설계 19
3.1 설계 원칙 19
3.2 VeX 개요 20
3.3 의미 구조를 고려한 HNSW 인덱스 분할 21
3.4 효율적인 호스트–DPU 통신 경로 24
3.5 탐색 과정 25
3.6 이기종성을 고려한 파이프라인 중첩 27
3.7 VeX와 스토리지 계층과의 통합 28
3.8 구현 29
4 평가 31
4.1 실험 환경 31
4.2 메모리 제약 환경에서의 성능 33
4.2.1 전체 처리량–Recall@100 비교 33
4.2.2 디스크 기반 ANN과의 비교(DiskANN ) 34
4.2.3 분산 메모리 확장 방식과의 비교(HNSW-D(k)) 34
4.2.4 벡터 차원의 영향 35
4.3 오프로딩 비율에 대한 민감도 36
4.4 호스트–DPU 간 통신 오버헤드 분석 37
4.5 Batch 파이프라이닝이 자원 활용률과 처리량에 미치는 영향 38
4.6 대규모 데이터셋 분석 38
4.7 관련연구 40
5 결론 41
참고문헌 43

more