Confidently wrong: auditing classifiers using Reversible Generator MLP (RevGEN-MLP)

Wan Hasmar Azim, Wan Muhammad Isyraf and Rudrusamy, Gobithaasan (2025) Confidently wrong: auditing classifiers using Reversible Generator MLP (RevGEN-MLP). In: International Undergraduate Research, Innovation, Invention and Design (I-URIID) 2025: e-Book of Extended Abstracts. Universiti Teknologi MARA, Negeri Sembilan, pp. 171-175. ISBN 9786299595366
Abstract

Machine learning classifiers are known to be overconfident when facing unseen or anomalous inputs. These present a dangerous vulnerability that reduces the trust in the systems that rely on artificial intelligence as it becomes increasingly widespread. We present a novel method of auditing these machine learning models by generating confidently classified anomalies using Reversible Generator Multilayer Perceptron (RevGEN-MLP). These generated anomalies act as a tool for testing the robustness of different models, identifying which models are more vulnerable to overconfidence on a specific dataset. Our method can be adapted to datasets with continuous numerical features and can be tailored for domain-specific robustness testing. We demonstrate the application of this technique on two different datasets and show that the generated anomalies not only deceive the original model but can also transfer to other classifiers. Results reveal that a classifier’s vulnerability to such anomalies can vary by dataset, with Random Forest observed to be more robust towards these generated inputs compared to other neural networks and K-Nearest Neighbour classifiers.

Item Details
Edit Item
Edit Item
Downloads & Files
[thumbnail of 144710.pdf]
Text
144710.pdf
Download (1MB)
Indexing & Metrics
Download Statistics