Authors: Swati, Disha Vishwakarma, Anuradha Singh, Anusha Rathore
Abstract: The cybersecurity menace that is phishing attacks continues to thrive although online criminals are always intent on developing deceitful links that are exact replicas of genuine websites. Conventional machine learning (ML) methods were found to require use of human-coded lexical features which often fail in detecting complex subtleties of such forms of deception as, for example, homoglyph attacks, coded URLs or dynamically changing redirections. The authors of the research presented a new lightweight technique based on transformers for detecting phishing URLs, which uses context-sensitive URL tokenization and the self-attention methods, enabling detecting both structural features and semantic meaning of the URLs without the need for manual feature extraction. The authors employed the following methods: subword tokenization, transformer-based learning of the meanings of the words in context and some changes relying on adversarial thinking which helped achieve all of its goals, including better performance in capturing previously unknown forms of phishing. The authors evaluated their system on the publicly available datasets, and their results were comparable to those produced by people on regular systems, which indicates the high efficiency of the proposed technology. The conducted research proves that new transformer techniques are better at overcoming modern obfuscation techniques than conventional methods.
International Journal of Science, Engineering and Technology