AI phishing detection tool

Citation

Abstract

Phishing attacks remain one of the most critical cybersecurity threats, exploiting malicious URLs to deceive users and extract sensitive information. Traditional detection approaches, such as blacklist-based and rule-based systems, are increasingly ineffective against rapidly evolving and previously unseen phishing techniques. To address this challenge, this study proposes a hybrid deep learning framework that integrates a supervised transformer-based model (RoBERTa) with an unsupervised Autoencoder for anomaly detection using structured URL features. The RoBERTa model captures contextual and semantic patterns from raw URLs, while the Autoencoder identifies deviations through reconstruction error, enabling detection of unknown phishing behaviors. The framework is evaluated using standard metrics including accuracy, precision, recall, F1-score, and AUC, along with cross-validation to ensure robustness. Experimental results demonstrate that the hybrid model achieves near-perfect performance, significantly outperforming individual approaches while maintaining strong generalization capability. The combination of supervised and unsupervised learning enhances detection reliability and reduces false positives. Overall, this research provides an effective, scalable, and adaptive solution for real-world phishing detection in dynamic cybersecurity environments.

Description

This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 56-57).

Publisher Link

Type

Thesis

Creative Commons license

Attribution-NonCommercial-NoDerivatives 4.0 International

Except where otherwise noted, this item's license is described as

Attribution-NonCommercial-NoDerivatives 4.0 International