Preview of the new IC2 website. It is not public yet and is hidden from search engines.

Publications

Adapting Vision Transformer Models for Thai Cat Breed Classification via Dataset Augmentation and Class Cardinality Reduction

N Sarutipaisan, P Maksap, C Keawphinit, B Watanapa

Web Intelligence

Abstract

This paper presents an approach to adapting a pre-trained Vision Transformer (ViT) model for cat breed classification in a Thai context. Using the Cat Breeds Refined dataset from Kaggle as a foundation, we extend the dataset by collecting additional images of Thai cat breeds — including Khao Manee, Korat, Suphalak, and Konja — using automated web crawling and manual verification. This expansion is motivated by the need to cover Thai cat breeds that are culturally significant and commonly kept as pets in Thailand but are absent from the original dataset, growing the class set from 47 to 52 breeds. To improve model generalization, we apply a data augmentation pipeline consisting of random horizontal and vertical flipping, random rotation, and color jitter. We evaluate two model configurations: a broad 52-class model covering international and Thai breeds, achieving an accuracy of 92.63% (precision 0.9293, recall 0.9263, F1 0.9260); and a focused 19-class model aligned with breeds practically encountered in Thailand, achieving 96.67% (precision 0.9676, recall 0.9667, F1 0.9666). These findings demonstrate that aligning class definitions with the target deployment context, combined with domain-specific data augmentation, is an effective strategy for adapting general-purpose vision models to region-specific classification tasks.

Authors: Nudhana Sarutipaisan, Prechaya Maksap, Chawisa Keawphinit, Bunthit Watanapa

DOI · Google Scholar