FPGA implementation of a co-designed CNN and filtering block for digit sign language recognition

Luyoh, Joel Laing (2026) FPGA implementation of a co-designed CNN and filtering block for digit sign language recognition. Masters thesis, Universiti Teknologi MARA (UiTM).
Abstract

Sign language (SL) recognition system using convolutional neural network (CNN) have shown high accuracy performance in software level but its implementation on a portable hardware remains a challenge due to CNN’s computational demands and the lack consideration of hardware constraints during development. Moreover, the challenge is also due to the complex architecture of CNN that requires hardware with large memory resources and attain high computational cost. In addition, real-time image capturing introduces noises that necessitates the need of pre-processing filters. However, similar to the challenges of CNN, the filters implementation is also typically designed only in software without the hardware consideration. Therefore, this research aims to design and develop a functional CNN-based digit SL classification model that can implemented on field-programmable gate array (FPGA) hardware and enhanced the performance by a co-design approach of incorporating pre-processing filter block to the system. The feasibility of the CNN models with co-design filters are evaluated for real-time classification. In achieving the objectives of the research, a 7-layer CNN architecture was developed in both matrix laboratory (MATLAB) and Quartus software, then trained in MATLAB using an American digit SL dataset. Four filters (gaussian, median, bilateral and sobel) were first implemented in the FPGA pipeline, then adapted the design in MATLAB for training following the co-design approach. The weights and biases of the trained models were converted to the compatibility of FPGA format and synthesized with the designed CNN system in the FPGA pipeline. The CNN is implemented on the Terasic DE2-115 FPGA development board with D8M-GPIO camera for real-time testing. The benchmark CNN model without filter achieved 92.43% training accuracy, 88.53% validation accuracy, 69.33% testing accuracy in MATLAB and 14.33% real-time testing accuracy on FPGA using 90,612 logic elements (LEs). The bilateral filter model scores the best relative improvements over the benchmark model with +12.32% in MATLAB testing accuracy and +20.94% in real-time testing accuracy on FPGA. The performance hierarchy followed by the median filter model with improvement of +11.35% in MATLAB testing accuracy and +18.14% in real-time testing accuracy on FPGA, while gaussian filter model relatively improves by +6.16% in MATLAB testing accuracy and +8.86% in real-time testing accuracy on FPGA. The sobel filter model failed completely with only 9.97% training accuracy. In hardware implementation, all successful filter models utilized 95% to 97% of the available resources in FPGA. Looking at the consistent performance hierarchy (bilateral > median > gaussian > without filter) in both MATLAB and FPGA platforms, this validates the co-design approach in CNN. This research has successfully demonstrated that the edge-preserving filter, bilateral filter in particular, implemented through the co-design approach enhance the CNN performance for digit SL classification on FPGA significantly. The implementation of FPGA provides a realization of future application-specific integrated circuit (ASIC) development, producing practical and portable communication aided device to a reality for the deaf or hard of hearing community.

Item Details
Edit Item
Edit Item
Downloads & Files
[thumbnail of 145465_fulltext.pdf]
Text
145465_fulltext.pdf
Available under License Dasar Harta Intelek UiTM (Para 6).
Download (4MB)
[thumbnail of declarationform.pdf]
Text
declarationform.pdf
Restricted to Repository staff only
Download (454kB)
Location & Physical Holdings
Digital Copy
Physical Location
Item Status
Indexing & Metrics
Download Statistics