AN IMPROVED APPROACH OF WRAPPER GENERATION TECHNIQUES FOR WEB SOURCES
| Author(s) | : | Sweta |
| Institution | : | Dept. of P.G. Studies and Research in Computer Science, GUK |
| Published In | : | Vol. 5, Issue 4 — April 2018 |
| Page No. | : | 1392-1395 |
| Domain | : | Engineering |
| Type | : | Research Paper |
| ISSN (Online) | : | 2348-4470 |
| ISSN (Print) | : | 2348-6406 |
The World Wide Web has more and more online Web databases which can be searched through Web queryinterfaces. All the Web databases make up the deep Web. Often the retrieved information is enwrapped in Web pages inthe form of data records. These special Web pages are generated dynamically and are hard to index by traditionalcrawler based search engines, such as Google and Yahoo. The topic of Web data extraction has received a lot ofattention in recent years and most of the proposed solutions are based on analyzing the HTML source code or the tagtrees of the Web pages. Web data extraction is the process of extracting user required information from websites. Theweb document contains data which is not in structured format. Specific data is able to be extracted from all these Websources in order to be used by other users or applications. The word web data extraction means the extraction of datathat is present in the web documents in HTML format and removing the unwanted things such as tags, advertisements,videos and so on from web sources.
Sweta, “AN IMPROVED APPROACH OF WRAPPER GENERATION TECHNIQUES FOR WEB SOURCES”, International Journal of Advance Engineering and Research Development (IJAERD), Vol. 5, Issue 4, pp. 1392-1395, April 2018.








