검색 상세

FPGA-Embedded Deep Learning Architecture for Real-Time Position Estimation in PET Detector with 256-to-4 Multiplexing Readout

초록(요약문)

In radiation imaging systems, the increasing sensor integration density leads to data bandwidth saturation and reduced system scalability. Multiplexing is widely used to reduce the number of readout channels; however, high-ratio multiplexing introduces nonlinear positional distortion that degrades position estimation accuracy. Deep learning-based approaches have been investigated to compensate for this distortion, but most existing studies have been limited to post-processing in GPU environments, which restricts real-time operation. The purpose of this study was to design a high-density analog multiplexing data acquisition (DAQ) system that reduces 256-channel SiPM signals to 4 channels, to compensate for the resulting nonlinear positional distortion using multilayer perceptron (MLP)- and convolutional neural network (CNN)-based position estimation models, and to implement the trained models on an FPGA for real-time, on-device position estimation. A single detector block was constructed by arranging four detector modules in a 2 × 2 configuration, each module comprising a 14 × 14 LYSO crystal array coupled to an 8 × 8 SiPM array, yielding 784 crystal pixels and 256 channels in total. A multiplexing circuit was designed and fabricated to reduce the 256-channel SiPM signals to 32 channels through a symmetric charge division (SCD) stage based on diode Row/Column summation, followed by further reduction to 4 channels through a discretized positioning circuit (DPC) using a resistive voltage-divider network. The DAQ system was built using a Xilinx VCU118 board (Virtex UltraScale+) and four single-channel ADCs (65 MHz sampling rate, 14- bit resolution) oper xating in parallel, with real-time peak detection implemented on the FPGA. Two position estimation models, MLP and CNN, were designed in parallel for comparative evaluation. A row-column multiplexed DAQ system under the same detector configuration was used to generate reference coordinates for supervised learning, owing to its clear crystal identification. A total of 7.76 million events acquired from a 22Na point source were used as the training dataset. Both trained models were implemented on the FPGA using high-level synthesis (HLS), and real-time coordinate output was successfully verified. Performance was evaluated using the mean full width at half maximum (FWHM) of the flood-histogram crystal peaks (in normalized coordinates), the peak-to-valley ratio (PVR; dimensionless), and the Euclidean mean absolute error (MAE), while FPGA implementation efficiency was assessed using DSP utilization and inference latency. Deep learning-based position estimation substantially improved all metrics relative to conventional Anger logic. The mean crystal-peak FWHM in the flood histogram decreased from 0.2718 to 0.0091 (MLP) and 0.0095 (CNN), corresponding to improvements of approximately 96.6% and 96.5%; the FPGA implementation retained an improvement of approximately 96.4% relative to Anger logic. The PVR increased from 1.96 by factors of 19.5× (MLP) and 18.9× (CNN), with the FPGA implementation retaining factors of approximately 17.8× and 16.7×, respectively. On an identical evaluation dataset, the Euclidean MAE decreased from 0.974 to 0.0122 (MLP) and 0.0120 (CNN), an improvement of approximately 98.8% for both models. Regarding the FPGA implementation, the MLP reduced DSP utilization by approximately 44% and inference latency by approximately 35% compared with the CNN, and FPGA inference accelerated processing by approximately 30× (MLP) and 29× (CNN) relative to GPU execution. The approximately 10% performance reduction in the FPGA implementation relative to the GPU results is attributed to quantization in the HLS-based implementation and is considered acceptable for practical use. The MLP achieved superior hardware efficiency while maintaining position estimation accuracy comparable to the CNN, confirming its suitability for edge-device, real-time PET DAQ systems. This study demonstrates that the nonlinear positional distortion introduced by high-ratio multiplexing can be effectively compensated using deep learning-based position estimation, and that real-time, on- device position estimation is achievable via HLS-based FPGA implementation without external computing resources. The proposed architecture provides a scalable and cost-effective solution for real- time PET detector systems employing high-density SiPM arrays.

more

목차

1. Introduction 1
2. Materials and Methods 4
2.1 Detector Configuration 4
2.2 Multiplexing Architecture 5
2.3 FPGA-Based DAQ System 6
2.4 Dataset Acquisition and Reference Coordinate Generation 8
2.5 Deep Learning-Based Position Estimation 10
2.6 FPGA Implementation 12
2.7 Performance Evaluation 14
3. Results 15
3.1 Position Estimation Performance 15
3.2 FPGA Implementation Validation 17
3.3 FPGA Resource Utilization and Inference Latency 19
4. Discussion 20
5. Conclusion 23
Bibliography 24
Abstract in Korean 26

more