Dissimilarity-generalized minimum spanning tree-based isolation forest
More details
Hide details
1
Lublin University of Technology, Nadbystrzycka 36B, 20-618 Lublin, Poland
Corresponding author
Łukasz Gałka
Lublin University of Technology, Nadbystrzycka 36B, 20-618 Lublin, Poland
KEYWORDS
TOPICS
ABSTRACT
Anomaly detection is one of the major challenges in modern computer systems. The collection of increasingly large amounts of data creates a growing need for unsupervised techniques. In this article, the Dissimilarity-Generalized Minimum Spanning Tree-Based Isolation Forest (DG-MSTIF) is proposed. The method is based on constructing forests of isolation trees formed using a minimum spanning tree algorithm. One potential limitation of the baseline method is its exclusive use of the Euclidean distance, which may not be optimal for all datasets, particularly in higher-dimensional data spaces. Therefore, a generalization of the method is introduced in which different dissimilarity measures are applied, including Euclidean, Manhattan, Chebyshev, Minkowski, cosine, correlation, and Canberra. The measured area under the precision–recall curve (PR AUC) values showed that the use of appropriate dissimilarity measures can improve detection effectiveness. In addition, execution time measurements made it possible to compare the computational performance depending on the selected measure. The study indicates that an appropriate choice of dissimilarity measure can lead to improved detection performance.