Abstract
Cyber crimes including computer virus/malwares, spam, illegal sales, and phishing websites are proliferated aggressively via the disguised Uniform Resource Locators (URL). Although numerous studies were conducted for the URL classification task, the traditional URL classification solutions retreated due to the hand-crafted feature engineering and the boom of newly generated URLs. In this paper, we study the representation learning of URLs, and explore the URL classification using deep learning. Specifically, we propose URL2vec to extract both the structural and lexical features of URLs, and apply temporal convolutional network (TCN) for the URL classification task. The experimental results show that URL2vec outperforms both word2vec and character-level embedding for URL representation, and TCN achieves the best performance than baselines with the precision up to 95.97%.
| Original language | English |
|---|---|
| Title of host publication | 2019 IEEE International Conference on Intelligence and Security Informatics, ISI 2019 |
| Editors | Xiaolong Zheng, Ahmed Abbasi, Michael Chau, Alan Wang, Lina Zhou |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 74-79 |
| Number of pages | 6 |
| ISBN (Electronic) | 9781728125046 |
| DOIs | |
| State | Published - Jul 2019 |
| Event | 17th IEEE International Conference on Intelligence and Security Informatics, ISI 2019 - Shenzhen, China Duration: 1 Jul 2019 → 3 Jul 2019 |
Publication series
| Name | 2019 IEEE International Conference on Intelligence and Security Informatics, ISI 2019 |
|---|
Conference
| Conference | 17th IEEE International Conference on Intelligence and Security Informatics, ISI 2019 |
|---|---|
| Country/Territory | China |
| City | Shenzhen |
| Period | 1/07/19 → 3/07/19 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 16 Peace, Justice and Strong Institutions
Keywords
- Cyber Crime
- Temporal Convolutional Network
- URL Classification
- URL2vec
Fingerprint
Dive into the research topics of 'Leverage temporal convolutional network for the representation learning of URLs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver