• Home
  • Search
  • A lexical approach for classifying malicious URLs
  • Open Access IconOpen Access
  • Cite Icon68
  • https://doi.org/10.1109/hpcsim.2015.7237040Copy DOI Icon

A lexical approach for classifying malicious URLs

  • Jul 1, 2015
  • Michael Darling +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Given the continuous growth of malicious activities on the internet, there is a need for intelligent systems to identify malicious web pages. It has been shown that URL analysis is an effective tool for detecting phishing, malware, and other attacks. Previous studies have performed URL classification using a combination of lexical features, network traffic, hosting information, and other strategies. These approaches require time-intensive lookups which introduce significant delay in real-time systems. In this paper, we describe a lightweight approach for classifying malicious web pages using URL lexical analysis alone. Our goal is to explore the upper-bound of the classification accuracy of a purely lexical approach. We also aim to develop a scalable approach which could be used in a real-time system. We develop a classification system based on lexical analysis of URLs. It correctly classifies URLs of malicious web pages with 99.1% accuracy, a 0.4% false positive rate, an F1-Score of 98.7, and 0.62 milliseconds on average. Our method also outperforms similar approaches when classifying out-of-sample data.

Similar Papers
  • Book Chapter
  • Citations14

Measurement Study on Malicious Web Servers in the .nz Domain

  • Jan 01, 2009
  • Christian Seifert +4
  • Book Chapter
  • Citations3

Clustering Client Honeypot Data to Support Malware Analysis

  • Jan 01, 2010
  • Yaser Alosefer +1
  • Conference Article
  • Citations1

Detection of Malicious Webpages Using Deep Learning

  • Dec 15, 2021
  • A K Singh +1
  • Conference Article
  • Citations14

Malicious Webpage Classification Based on Web Content Features using Machine Learning and Deep Learning

  • Oct 26, 2022
  • A Saleem Raja +3
  • Conference Article
  • Citations20

Detecting Obfuscated JavaScript Malware Using Sequences of Internal Function Calls

  • Mar 28, 2014
  • Alireza Gorji +1
  • Book Chapter
  • Citations1

Towards a Lexical Analysis on Chinese Middle Constructions

  • Jan 01, 2018
  • Lulu Wang
  • Conference Article
  • Citations2

POSTER

  • Oct 12, 2015
  • Toshiki Shibahara +4
  • Conference Article
  • Citations6

An efficient visitation algorithm to improve the detection speed of high-interaction client honeypots

  • Nov 02, 2011
  • Hong-Geun Kim +4
  • Research Article
  • Citations8

Malicious web pages: What if hosting providers could actually do something…

  • Jun 19, 2015
  • Computer Law & Security Review
  • Huw Fryer +2
  • Research Article
  • Citations17

A drift aware adaptive method based on minimum uncertainty for anomaly detection in social networking

  • Aug 19, 2020
  • Expert Systems with Applications
  • Emad Mahmodi +2
  • Book Chapter
  • Citations5

Identification of Malicious Web Pages by Inductive Learning

  • Jan 01, 2009
  • Peishun Liu +1
  • Research Article
  • Citations15

JsSandbox: A Framework for Analyzing the Behavior of Malicious JavaScript Code using Internal Function Hooking

  • Jan 01, 2012
  • KSII Transactions on Internet and Information Systems
  • Hyoung Chun Kim
  • Research Article
  • Citations119

Automated testing of graphics shader compilers

  • Oct 12, 2017
  • Proceedings of the ACM on Programming Languages
  • Alastair F Donaldson +3
  • PDF
  • Research Article
  • Citations2

Malicious and benign websites classification using machine learning methods

  • Aug 06, 2020
  • Theoretical and Applied Cybersecurity
  • M Lavreniuk +1
  • Conference Article
  • Citations6

Research on Prevention Solution of Advanced Persistent Threat

  • Jan 01, 2014
  • Xiaomei Liu
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.