실시간 웹 크롤링 분산 모니터링 시스템 설계 및 구현

Design and Implemention of Real-time web Crawling distributed monitoring system

초록

We face problems from excessive information served with websites in this rapidly changing information era. We find little information useful and much useless and spend a lot of time to select information needed. Many websites including search engines use web crawling in order to make data updated. Web crawling is usually used to generate copies of all the pages of visited sites. Search engines index the pages for faster searching. With regard to data collection for wholesale and order information changing in realtime, the keyword-oriented web data collection is not adequate. The alternative for selective collection of web information in realtime has not been suggested. In this paper, we propose a method of collecting information of restricted web sites by using Web crawling distributed monitoring system (R-WCMS) and estimating collection time through detailed analysis of data and storing them in parallel system. Experimental results show that web site information retrieval is applied to the proposed model, reducing the time of 15-17%.

키워드

웹 크롤링; 빅 데이터; 하둡; 스파크; 카프카; 병렬시스템; 모니터링; Web Crawling; big data; hadoop; spark; kafka; Parallel systems; Monitoring
제목
실시간 웹 크롤링 분산 모니터링 시스템 설계 및 구현
제목 (타언어)
Design and Implemention of Real-time web Crawling distributed monitoring system
저자
김영아; 김계희; 김현주; 김창근
DOI
10.22156/CS4SMB.2019.9.1.045
발행일
2019-01
저널명
융합정보논문지
권
9
호
1
페이지
45 ~ 53