A Supervised Approach for Automatic Web Documents Topic Extraction Using Well-Known Web Design Features
Autor: | Amirreza Shirani, Ahmad Zaeri, Kazem Taghandiki |
---|---|
Rok vydání: | 2016 |
Předmět: |
Topic model
Information retrieval Computer science Supervised learning 020206 networking & telecommunications 02 engineering and technology computer.software_genre Computer Science Applications Education Information extraction Web mining Web design Web page 0202 electrical engineering electronic engineering information engineering 020201 artificial intelligence & image processing Data mining Web crawler computer Data Web |
Zdroj: | International Journal of Modern Education and Computer Science. 8:20-27 |
ISSN: | 2075-017X 2075-0161 |
DOI: | 10.5815/ijmecs.2016.11.03 |
Popis: | The aim of this paper is to propose an efficient method for identification of web document topics which is often considered as one of the debatable challenges in many information retrieval systems. Most of the previous works have focused on analyzing the entire text using time-consuming methods and also many of them have used unsupervised approaches to identify the main topic of documents. However, in this paper, it is attempted to exploit the most widely-used Hyper-Text Markup Language (HTML) features to extract topics from web documents using a supervised approach. Hiring an interactive crawler, we firstly try to analyze HTML structures of 5000 webpages in order to identify the most widely-used HTML features. In the next step, the selected features of 1500 webpages are extracted using the same crawler. Suitable topics are given to each web document by users in a supervised learning process. A topic modeling technique is used over extracted features to build four classifiersC4.5, Decision Tree, Naive Bayes and Maximum Entropywhich are separately adopted to train and test our data. The results of classifiers are compared and the high accurate classifier is selected. In order to examine our approach in a larger scale, a new set of 3500 web documents is evaluated using the selected classifier. Results show that the proposed system provides remarkable performance which is able to obtain 71.8% recognition rate. |
Databáze: | OpenAIRE |
Externí odkaz: |