Abstract

Land use scene classification (LUSC) from remote sensing imagery plays a critical role in environmental monitoring, urban planning, and sustainable resource management. In recent years, deep learning methods have significantly advanced the state-of-the-art, with Convolutional Neural Networks (CNNs) dominating the field because of their strong ability to capture local spatial features. However, the emergence of Vision Transformers (ViTs) has introduced a new paradigm that models long-range dependencies through self attention mechanisms, potentially enabling improved global context understanding. This study presents a comparative assessment of Vision Transformers and CNN-based architectures for remote sensing land use scene classification. Representative CNN models, such as AlexNet, are evaluated alongside the Vision Transformer (ViT) using benchmark remote sensing datasets, including the UC Merced (UCM) Land Use and EuroSAT Land Use datasets. The study examines classification accuracy, precision, recall, F1-score, and computational complexity to provide a comprehensive performance comparison. Experimental results demonstrate that CNNs perform robustly on datasets with limited training samples and strong local texture characteristics, whereas Vision Transformers exhibit superior performance in capturing global spatial relationships in complex scenes when sufficient training data are available. However, ViTs typically require greater computational resources and larger training datasets to achieve optimal performance. The findings of this study provide insights into the strengths and limitations of both architectures and offer guidance for selecting appropriate models for remote sensing land use scene classification applications.

Description

All articles published by IJACSA are made freely and permanently accessible online immediately upon publication, without subscription charges or registration barriers. Authors of articles published in IJACSA are the copyright holders of their articles and have granted to any third party, in advance and in perpetuity, the right to use, reproduce, or disseminate the article, provided that the original work is properly cited.

Publisher

International Journal of Advanced Computer Science and Applications

Date of publication

Summer 7-30-2026

Language

english

Persistent identifier

http://hdl.handle.net/10950/5132

Document Type

Article

Publisher Citation

4. Kulkarni, A. D. (2026). Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification. International Journal of Advanced Computer Science and Applications, vol. 17, no. 7, pp. 1-11.

Share

COinS