Transformer-Based Ensemble Learning for Robust Multi-Class Melanoma Classification

Authors

  • Deborah A. Adedigba Southampton Solent University, United Kingdom
  • Raza Hasan Southampton Solent University, United Kingdom
  • Salman Mahmood Nazeer Hussain University, Pakistan

Keywords:

Melanoma Detection, Vision Transformer (ViT), Ensemble Learning, Meta-learning, Explainable AI (XAI), Dermoscopic Images, Multi-class Classification

Abstract

Melanoma signifies the most hazardous form of skin cancer, characterized by a worldwide increase in incidence despite the improvement of dermatological screening techniques. The existing techniques of automated detection have been found to be limited in terms of global context modeling, reliance on a single architecture, and the lack of clinically interpretable models. The present research aims to bridge the existing gap by introducing a novel framework of transformer-based ensemble learning for the automated detection of melanoma, combining Vision Transformer models and CNN models via the application of a meta-learning approach. In the present study, the Vision Transformer, Swin Transformer, MaxViT, and ConvNeXt models have been combined with a ResNet50 CNN model via a Random Forest-based meta-classification approach. A balanced multi-source dataset was created by combining the PH2, MedNode, and Derm7pt datasets, covering six different classes of lesions. The images were integrated using a real image-based oversampling technique. The stacked ensemble models achieved a macro F1 score of 0.998, Balanced Accuracy of 0.999, and a macro-AUC of 0.9999, indicating a rise of 12% over the highest performing model. The statistical significance of the models' complementary errors was assessed using McNemar’s test, which indicated a statistically significant complementary error for all models (p < 0.001). The Grad-CAM and GradCAM++ approaches have been modified for the different models to ensure the provision of spatially interpretable visualizations in accordance with dermatological diagnostic criteria. The proposed framework has shown that architecturally diverse stacking ensembles, when used with appropriate balancing of the dataset and leakage-resistant cross-validation, can achieve robust multi-class classification of skin lesion images with potential to inform clinical decision-making.

Downloads

Download data is not yet available.

Author Biographies

  • Deborah A. Adedigba, Southampton Solent University, United Kingdom

    Deborah Adedigba is a recent graduate with an MSc in Applied AI and Data Science from Southampton Solent University, awarded an OFS Scholarship. She specializes in computer vision, machine learning, and advanced analytics with diverse research work spanning medical image analysis, financial forecasting, and infectious disease detection. Her expertise includes automated skin lesion detection for melanoma and tuberculosis detection from chest X-rays. Deborah has published conference papers and developed AI applications including chatbots and cryptocurrency prediction systems, demonstrating extensive experience applying deep learning techniques across healthcare and finance domains. 

  • Salman Mahmood, Nazeer Hussain University, Pakistan

    Associate Professor at Nazeer Hussain University. He specializes in developing AI and machine learning models for medical diagnostics, with a focus on performance optimization and clinical applicability.

Downloads

Published

02-07-2026

Issue

Section

Articles

How to Cite

Deborah A. Adedigba, Hasan, R., & Mahmood, S. (2026). Transformer-Based Ensemble Learning for Robust Multi-Class Melanoma Classification. Journal of Soft Computing and Data Mining, 7(2), 256-277. https://publisher.uthm.edu.my/ojs/index.php/jscdm/article/view/23538