Main Article Content

Heterogeneous borda ensemble feature selection: An enhancement to machine learning approach for uniform resource locator phishing


P. O. Olabisi
B. M. Olukoya
G. O. Ogunleye
O. J. Adetunji

Abstract

Phishing attacks, which involve deceptive methods to harvest sensitive user information, pose significant threats to online transactions. Despite ongoing efforts to combat these cyberattacks, traditional solutions have struggled to address the increasing sophistication of phishing schemes. Machine learning (ML) techniques as emerging tools for detecting these threats are particularly more effective when combined with robust feature selection methods. The authors therefore introduced a novel approach called Heterogeneous Borda Ensemble Feature Selection (HBEFS), wherein three filter-based statistical techniques generate primary subsets of phishing indicators. The variables selected by each technique are aggregated to form the final baseline features. The proposed method efficiently identifies essential phishing features. The use of a random forest (RF) model with HBEFS indicators achieved an impressive performance of 97.4%. Also, classical models exhibit significant improvement when evaluated on the baseline features. Comparing the performance of models using HBEFS variables, individual statistical techniques, and other studies, the HBEFS consistently outperforms them. These research findings necessitate the need for innovative feature selection strategies to enhance potency of ML classification for phishing detection.


Journal Identifiers


eISSN: 2437-2110
print ISSN: 0189-9546