Abstract
Land use scene classification (LUSC) from remote sensing imagery plays a critical role in environmental monitoring, urban planning, and sustainable resource management. In recent years, deep learning methods have significantly advanced the state-of-the-art, with Convolutional Neural Networks (CNNs) dominating the field because of their strong ability to capture local spatial features. However, the emergence of Vision Transformers (ViTs) has introduced a new paradigm that models long-range dependencies through self attention mechanisms, potentially enabling improved global context understanding. This study presents a comparative assessment of Vision Transformers and CNN-based architectures for remote sensing land use scene classification. Representative CNN models, such as AlexNet, are evaluated alongside the Vision Transformer (ViT) using benchmark remote sensing datasets, including the UC Merced (UCM) Land Use and EuroSAT Land Use datasets. The study examines classification accuracy, precision, recall, F1-score, and computational complexity to provide a comprehensive performance comparison. Experimental results demonstrate that CNNs perform robustly on datasets with limited training samples and strong local texture characteristics, whereas Vision Transformers exhibit superior performance in capturing global spatial relationships in complex scenes when sufficient training data are available. However, ViTs typically require greater computational resources and larger training datasets to achieve optimal performance. The findings of this study provide insights into the strengths and limitations of both architectures and offer guidance for selecting appropriate models for remote sensing land use scene classification applications.
Description
All articles published by IJACSA are made freely and permanently accessible online immediately upon publication, without subscription charges or registration barriers. Authors of articles published in IJACSA are the copyright holders of their articles and have granted to any third party, in advance and in perpetuity, the right to use, reproduce, or disseminate the article, provided that the original work is properly cited.
Publisher
International Journal of Advanced Computer Science and Applications
Date of publication
Summer 7-30-2026
Language
english
Persistent identifier
http://hdl.handle.net/10950/5132
Document Type
Article
Recommended Citation
Kulkarni, Arun D., "Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification" (2026). Computer Science Faculty Publications and Presentations. Paper 37.
http://hdl.handle.net/10950/5132
Publisher Citation
4. Kulkarni, A. D. (2026). Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification. International Journal of Advanced Computer Science and Applications, vol. 17, no. 7, pp. 1-11.