THCS-ViT: A Hybrid Lightweight Vision Transformer with Multi-Scale Feature Fusion for Thermal Facial Emotion Recognition
More details
Hide details
1
Department of Computer Science, Faculty of Electrical Engineering and Computer Science, Lublin University of Technology, Nadbystrzycka 36B Street, Lublin, Poland
Publication date: 2026-09-01
Corresponding author
Kamil Pietrak
Department of Computer Science, Faculty of Electrical Engineering and Computer Science, Lublin University of Technology, Nadbystrzycka 36B Street, Lublin, Poland
Adv. Sci. Technol. Res. J. 2026;
KEYWORDS
TOPICS
ABSTRACT
Thermal facial images pose serious challenges for classical convolutional networks and standard transformer architectures due to their low contrast, significant sensor noise, and lack of rich textural details. In this paper, we propose THCS-ViT (Thermal Hybrid Cross-Scale Vision Transformer), a hybrid architecture integrating a lightweight convolutional stem inspired by MobileNetV3, a Cross-Scale Attention Fusion (CSAF) module, and a modified MobileViT-S backbone. The model was trained on the KTFEv2 (Korean Thermal Facial Emotion Dataset v2) dataset and evaluated through cross-dataset experiments on the Rainbow subset of the Comprehensive Facial Thermal Dataset (CFTD). The proposed architecture achieved a mean accuracy of 93.8% ± 0.6% (5-fold cross-validation) on KTFEv2 and 84.36% on CFTD Rainbow, with only 1.85 million parameters and 0.68 GFLOPs. These results outperform selected baseline models in terms of both accuracy and computational efficiency. Ablation studies, statistical tests, and explainability analyses confirm the contribution of individual components. The low parameter count suggests potential for edge deployment, though hardware validation remains future work. The proposed architecture offers a promising balance between accuracy and computational efficiency, with the CSAF module and thermal bias providing complementary benefits for thermal emotion recognition.