A Hierarchical Attention Network for Explainable Phishing URL Detection on Nigerian Financial Platforms
Abstract
Phishing remains the major channel of cyber-fraud in Nigeria, where fraudulent uniform resource locators (URLs) are used to pretend to be legitimate banks, mobile money services, and digital payments channels in order to steal users' credentials, Bank Verification Number (BVN) and card details. Current detection solutions used by Nigerian financial organisations depend heavily on blacklists, heuristics, and traditional machine learning classifiers working as black boxes, which do not give any explanation of why a particular URL was classified as malicious and show poor generalization to emerging phishing schemes. In this paper, we propose a Hierarchical Attention Network (HAN) that views a URL as a two-level hierarchy of domain and path, where each of them is segmented into tokens and apply attention both at token and segment levels to detect phishing URLs and highlight the specific parts of a URL contributing to this classification. Our model is based on 164,548 Nigerian bank-related URLs extracted from the 822,010-row Kaggle Phishing and Legitimate URLs dataset with 80/20 stratified train/test split. The baseline is a flattened BiLSTM network without attention layers. The baseline gave 94.37% accuracy (phishing F1 score = 0.9421), while our HAN scored 94.05% accuracy (phishing F1 score = 0.9373) along with higher stability of validation accuracy curve and smaller false-positive rate, meaning that slight drop in accuracy is compensated by increased precision and interpretability. The visualisation of the attention weights showed that the model always gives the most weight to the structural anomalies like abnormal subdomain tokens, combined or hyphenated strings, and non-standard prefixes like 'secure' or 'mail', but not to brand names of the banks themselves.
Full Text:
PDFReferences
Adebowale, M. A., Lwin, K. T., & Hossain, M. A. (2023). Intelligent phishing detection scheme using deep learning algorithms. Journal of Enterprise Information Management, 36(3), 747 to 766.
Aljofey, A., Jiang, Q., Qu, Q., Huang, M., & Niyigena, J. P. (2020). An effective phishing detection model based on character level convolutional neural network from URL. Electronics, 9(9), 1514.
Asiri, S., Xiao, Y., Alzahrani, S., Li, S., & Li, T. (2023). A survey of intelligent detection designs of HTML URL phishing attacks. IEEE Access, 11, 6421 to 6443.
Ayorinde, A. S. (2025). Explainable deep learning models for detecting sophisticated cyber enabled financial fraud across multi layered FinTech infrastructure. International Journal of Cybersecurity and Digital Forensics, 5(3), 241 to 263.
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.
Banik, S., Dandyala, S. S. M., & Nadimpalli, S. V. (2022). Heuristic based detection techniques. International Journal of Advanced Engineering Technologies and Innovations, 1(2), 352 to 362.
Chai, Y., Zhou, Y., Li, W., & Jiang, Y. (2021). An explainable multi modal hierarchical attention model for developing phishing threat intelligence. IEEE Transactions on Dependable and Secure Computing, 19(2), 790 to 803.
Chanaa, A. (2021). E learning text sentiment classification using hierarchical attention network (HAN). International Journal of Emerging Technologies in Learning, 16(13), 157 to 167.
Fajar, A., Yazid, S., & Budi, I. (2024). Comparative analysis of black box and white box machine learning model in phishing detection. arXiv preprint arXiv:2412.02084.
Fatokun, J. O., Sikiru, M., Balogun, F., & Okorie, D. D. (2025). Fraud detection in Nigerian investment advisory sector using machine learning algorithms. FUDMA Journal of Sciences, 9(10), 5 to 11.
George, M. Z. H., Alam, M. K., & Hasan, M. T. (2025). Machine learning for fraud detection in digital banking: A systematic literature review. arXiv preprint arXiv:2510.05167.
Gupta, B. B., Yadav, K., Razzak, I., Psannis, K., Castiglione, A., & Chang, X. (2021). A novel approach for phishing URLs detection using lexical based machine learning in a real time environment. Computer Communications, 175, 47 to 57.
Hu, G., Lu, G., & Zhao, Y. (2021). Bidirectional hierarchical attention networks based on document level context for emotion cause extraction. Findings of the Association for Computational Linguistics: EMNLP 2021, 558 to 568.
Jain, A. K., Parashar, S., Katare, P., & Sharma, I. (2020). Phishskape: A content based approach to escape phishing attacks. Procedia Computer Science, 171, 1102 to 1109.
Jaziriyan, M. M., & Ghaderi, F. (2023). Automatic post editing of hierarchical attention networks for improved context aware neural machine translation. Journal of AI and Data Mining, 11(1), 95 to 102.
Karim, A., Shahroz, M., Mustofa, K., Belhaouari, S. B., & Joga, S. R. K. (2023). Phishing detection system through hybrid machine learning based on URL. IEEE Access, 11, 36805 to 36822.
Korkmaz, M., Sahingoz, O. K., & Diri, B. (2020). Detection of phishing websites by using machine learning based URL analysis. Proceedings of the 11th International Conference on Computing, Communication and Networking Technologies (ICCCNT), 1 to 7.
Lim, W. H., Liew, W. F., Lum, C. Y., & Lee, S. F. (2020). Phishing security: Attack, detection, and prevention mechanisms. Proceedings of the International Conference on Digital Transformation and Applications (ICDXA).
Ngo, V. D., Vuong, T. C., Van Luong, T., & Tran, H. (2024). Machine learning based intrusion detection: Feature selection versus feature extraction. Cluster Computing, 27(3), 2365 to 2379.
Nigeria Inter Bank Settlement System (NIBSS). (2023). 2023 annual fraud landscape report. Nigeria Inter Bank Settlement System Plc.
Ozcan, A., Catal, C., Donmez, E., & Senturk, B. (2021). A hybrid DNN LSTM model for detecting phishing URLs. Neural Computing and Applications, 35(7), 4957 to 4973.
Rashid, J., Mahmood, T., Nisar, M. W., & Nazir, T. (2020). Phishing detection using machine learning technique. Proceedings of the 2020 1st International Conference of Smart Systems and Emerging Technologies (SMART TECH), 43 to 46.
Tajaddodianfar, F., Stokes, J. W., & Gururajan, A. (2020). Texception: A character or word level deep learning model for phishing URL detection. Proceedings of ICASSP 2020, 2857 to 2861.
Zamir, A., Khan, H. U., Iqbal, T., Yousaf, N., Aslam, F., Anjum, A., & Hamdani, M. (2020). Phishing web site detection using diverse machine learning algorithms. The Electronic Library, 38(1), 65 to 80.
Refbacks
- There are currently no refbacks.