[python][flask] webpage-scraper

webpage-scraper

webpage-scraper-master.zip

webpage-scraper is a flask based application which allows the users to :

  • Input URL with the freedom of inputting it with/ without the protocol and sub-domain specifiers.
  • Fetch a list of URLs to all the images on the webpage with an option to download all the images in a directory with name specified by the user.
  • Get a list of all the hyperlinks on the webpage. Save them into a text file with a name specified by the user.
  • Get the indented html source code of the webpage and save it in a .html file with a user-provided name.
  • Fetch the text on the webpage stripping the html code. Save it in a text file with a filename of user's choice.
  • The database is deployed on mLab and uses MongoDB for fast access to long list of images, hyperlinks and text for a URL that has been requested by some other user in the past, thus, reducing processing time for subsequent users.

Pre- requisites

To install requirements:

[sudo] pip install requirements

If you don't have pip installed, this Python installation guide can guide you through the process.

To install MongoDB Community Edition:

Make sure you have MongoDB installed

Getting started

git clone http://github.com/mansimarkaur/webpage-scraper 
cd webpage-scraper
python crawler_flask.py

Open http://127.0.0.1:5000/ in your browser. Input URL and have fun ????

[출처] https://github.com/mansimarkaur/webpage-scraper

 

 

 

 

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수

등록된 글이 없습니다.

대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED