Novendra Setyawan, Mas Nurul Achmadiah, Chi-Chia Sun, Wen-Kai Kuo
Batik is a distinctive design that represents specific traits and is a cultural legacy acknowledged by UNESCO. The identification of new batik patterns and the differentiation between pattern combinations are becoming challenging. Several studies have been conducted using the deep learning model, predominantly the Convolution Neural Network (CNN). However, CNN's ability to extract local dependencies is limited, requiring a large number of parameters and a complex architecture to accurately recognize Batik patterns that contain both local and global feature context This study describes the creation of a Multi Stage Vision Transformer (MSViT) for the purpose of classifying Batik Patterns. The construction of the Vision Transformer involves the use of a multi-stage down-sampling technique, where convolution is utilized as a down-sampling module in each stage. The goal of down-sampling is to enhance the Vision Transformer's capability to capture both local and global dependencies in the Batik feature map. The Vanilla Attention module is employed as a spatial token mixer to identify the spatial correlation among pixels. The effectiveness of the Multi Stage Transformer in classifying Batik patterns has been demonstrated through multiple tests utilizing the Batik 300 and Batik Nitik 960 datasets. The Vision Transformer's performance surpasses that of state-of-the-art CNN while requiring less compute and fewer parameters. © 2024 IEEE.
University of Muhammadiyah Malang, Department of Electrical Engineering, Indonesia; National Formosa University, Department of Electro-Optics, Taiwan; State Polytechnic of Malang, Department of Electronics Engineering, Indonesia; National Formosa University, Smart Manufacturing and Intelligent Machinery Research Center, Taiwan; National Taipei University, Department of Electrical Engineering, Taiwan