Convolutional neural network with enhanced vision transformer model for defect classification of additive manufacturing products

Kamarulzaman, Nur Najiha (2026) Convolutional neural network with enhanced vision transformer model for defect classification of additive manufacturing products. Masters thesis, Universiti Teknologi MARA (UiTM).
Abstract

Additive manufacturing (AM), or 3D printing, refers to a process of joining materials layer by layer to create objects. The AM process faces inconsistent outcomes and poses challenges in maintaining quality control. Manual quality control can be time-consuming. Traditional defect inspection method has evolved into an automated strategy by leveraging the benefit of artificial intelligence (AI). Machine learning (ML) is a subset in the AI field, can support decision-making for process optimization, quality control, and system improvement. The existing methods used ML have only focused on the defect classification without considering local and global feature extraction. Local and global feature extraction play an important role, as different types of defects require different feature extraction. The CNN algorithm, as a local features extraction widely used to classify the defects from AM products. However, the CNN unable to identify the relationship between the extracted features, prevents the model from understanding the global context of the images. Vision transformer (ViT), is a global feature extraction model, which enables a more comprehensive understanding of the global context within an image. However, ViT remains limited in capturing local features as transformer rely heavily on global relationships. Hybrid architectures combining CNN and ViT have emerged, leveraging the advantages of both architecture of CNN ViT. However, challenges occur when ViT divides the original input images into a series of non-overlapping patches in fixed size, in which this method destructs the image local- continuity in some degree. In addition, ViT is often determined by some tokens with large amounts of information, whereas most tokens are redundant. To address the limitations, this study proposed a deep learning model introducing CNN with enhanced ViT (CNN-enhanced ViT) to classify types of AM defects including crack, stringing, and no defect arising from the AM process. The main contribution of CNN-enhanced ViT relies on its architecture. The proposed class token extracted from the CNN feature map may serve as a reference for global information. In addition, this study further contributes to enhance the ViT architecture by performing token selection using 1D maxpooling layer. On the other hand, previous study has performed token selection which involves ViT architecture only, but to the best of author’s knowledge, no studies have performed CNN-ViT together with token selection, especially in AM scope. In this study, a total of 1500 AM images were collected from the Malaysia Automotive, Robotics and IoT Institute (MARii), which consists of defect classes of crack and stringing and a no defect class. The performance of the proposed CNN-enhanced ViT model was then evaluated and compared with the CNN-ViT model and Swin transformer models using confusion matrix. The confusion matrix results showed that the proposed model CNN-enhanced ViT achieved the best results, demonstrating a value of average accuracy of 96.00% and average precision of 94.28%, sensitivity of 94.00% and F1-score of 94.01% accordingly. The results demonstrate that the proposed CNN-enhanced ViT model is potentially able to produce effective and accurate results in defect classification for AM products, offering a lightweight and scalable solution for intelligent visual inspection system. This work has potential in contributing to the advancement of defect classification systems especially in smart manufacturing industry.

Item Details
Edit Item
Edit Item
Downloads & Files
[thumbnail of 144958_fulltext.pdf]
Text
144958_fulltext.pdf
Available under License Dasar Harta Intelek UiTM (Para 6).
Download (1MB)
[thumbnail of declarationform.pdf]
Text
declarationform.pdf
Restricted to Repository staff only
Download (583kB)
Location & Physical Holdings
Digital Copy
Physical Location
Item Status
Indexing & Metrics
Download Statistics