Insect Pest Classification Using Vision Transformers
DOI:
https://doi.org/10.69511/ijdsaa.v7i2.337Keywords:
Transformer Networks, Vision Transformers, Image Classification, Insect Pest ClassificationAbstract
Transformer-based neural networks have emerged as the dominant approach for natural language processing tasks due to their strong performance in text understanding and generation. Motivated by this success, recent research has explored the application of transformer architectures to computer vision, an area traditionally dominated by Convolutional Neural Networks (CNNs). Although the adoption of transformers in vision tasks is relatively recent, early studies have demonstrated promising results, particularly in image classification. However, limited research has systematically evaluated Vision Transformer models for fine-grained insect pest classification using large-scale agricultural datasets. In this paper, a Vision Transformer (ViT)-based computer vision model is proposed for the automated identification and classification of insect pests that damage agricultural crops. Insect pests destroy nearly one-third of global agricultural production and pose significant risks to food security and public health. Automating pest identification using deep learning can significantly reduce the time and effort required compared to manual inspection. The proposed model is trained and fine-tuned on the IP102 dataset, which contains more than 75,000 images spanning 102 categories of common crop-damaging insect pests. Experimental results demonstrate that the 8×8 ViT achieved 41.75% test accuracy, 67.28% top 5 accuracy, and 85.56% ROC-AUC, compared with 47.61% accuracy, 73.76% top 5 accuracy, and 91.21% ROC-AUC achieved by the best-performing CNN, InceptionV3. These results demonstrate the potential of transformer-based architectures for fine-grained insect pest classification in the IP102 benchmark.Downloads
Published
2025-10-23
How to Cite
Khullar, R., Topham, L., & Basava Chola, C. . (2025). Insect Pest Classification Using Vision Transformers. International Journal of Data Science and Advanced Analytics, 7(2), 485–492. https://doi.org/10.69511/ijdsaa.v7i2.337
Issue
Section
Articles
License
Copyright (c) 2025 Remen Khullar, Luke Topham, Channa Basava Chola

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

International Journal of Data Science and Advanced Analytics (IJDSAA) is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License. This license allows users to copy, distribute and transmit an article, adapt the article as long as the author is attributed and the article is not used for commercial purposes.
The author(s) confirms
- The manuscript submission has not been previously published, nor is it before another journal for consideration (or an explanation has been provided in Comments to the Editor).
- The published materials used in the manuscript were obtained permission for reproduction. (if any)