An Efficient Approach for Web Indexing of Big Data through Hyperlinks in Web Crawling

Autor: R. Suganya Devi, D. Manjula, R. K. Siddharth
Jazyk: angličtina
Rok vydání: 2015
Předmět:
Zdroj: The Scientific World Journal, Vol 2015 (2015)
Druh dokumentu: article
ISSN: 2356-6140
1537-744X
DOI: 10.1155/2015/739286
Popis: Web Crawling has acquired tremendous significance in recent times and it is aptly associated with the substantial development of the World Wide Web. Web Search Engines face new challenges due to the availability of vast amounts of web documents, thus making the retrieved results less applicable to the analysers. However, recently, Web Crawling solely focuses on obtaining the links of the corresponding documents. Today, there exist various algorithms and software which are used to crawl links from the web which has to be further processed for future use, thereby increasing the overload of the analyser. This paper concentrates on crawling the links and retrieving all information associated with them to facilitate easy processing for other uses. In this paper, firstly the links are crawled from the specified uniform resource locator (URL) using a modified version of Depth First Search Algorithm which allows for complete hierarchical scanning of corresponding web links. The links are then accessed via the source code and its metadata such as title, keywords, and description are extracted. This content is very essential for any type of analyser work to be carried on the Big Data obtained as a result of Web Crawling.
Databáze: Directory of Open Access Journals