Anonymous Web Scraping using Python and Tor

Requirements

Python

c01python-logo-master-v3-TM.png

 

socks module provides a standard socket-like interface for Python for tunneling connections through SOCKS proxies.
$ sudo pip install pysocks

Requests is an HTTP library, written in Python, that is wrapper over urllib*, and is very pythonic in use.
$ sudo pip install requests

c02requests-sidebar.png

 

Fedora

$ sudo yum install tor

Ubuntu

$ sudo apt-get install tor

c032000px-Tor-logo-2011-flat_svg.png

경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

 

After doing all the required installations. To start the tor and let run in background run following command.
$ tor &

By default tor uses port# 9050 if not mentioned otherwise. You can check if the process is listening using command netstat.
$ netstat -tupln

look for process listening on port# 9050 c04untitled.png

 

Now for the programming part do following open up the python interpreter and run commands as follows.
>>> import socks
>>> import socket
>>> socks.setdefaultproxy(proxy_type=socks.PROXY_TYPE_SOCKS5, addr="127.0.0.1", port=9050)

socks.setdefaultproxy sets a default proxy which all further socksocket objects will use, unless explicitly changed.
>>> socket.socket = socks.socksocket

socks.socksocket returns a socket object which is assigned to socket.socket which opens a socket. Now all connections made by the script will be done using this socket.

>>> import requests
>>> print requests.get("http://icanhazip.com").text
176.10.99.203

Try opening that url http://icanhazip.com from your browser as well. This website shows your public IP address. You will see different IP address in browser and in program output. Now you can change above script and write your webscraping or webcrawling program around it and make your python program run anonymously on the internet.

 

[출처] https://deshmukhsuraj.wordpress.com/2015/03/08/anonymous-web-scraping-using-python-and-tor/

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
22 [파이썬으로 웹 크롤러 만들기] 크롤링 시작하기(2/3) file 졸리운_곰 2017.02.16 709
21 [파이썬으로 웹 크롤러 만들기] 크롤링 시작하기(1/3) file 졸리운_곰 2017.02.16 772
20 Python으로 RESTAPI 이용하기 졸리운_곰 2017.02.11 462
19 [Python] 네이버 블로그 xml-rpc 클라이언트 만들기 – 1 졸리운_곰 2017.01.07 563
18 웹 크롤링 by python file 졸리운_곰 2016.12.14 1451
17 [Scrapy] 웹사이트 크롤링해서 DB 저장 하기(분양정보수집사례) file 졸리운_곰 2016.11.29 740
16 [Scrapy] 웹사이트 크롤링해서 파일 저장 하기(분양정보수집사례) file 졸리운_곰 2016.11.29 1721
15 [python] BeautifulSoup으로 웹에 있는 데이터 긁어오기 졸리운_곰 2016.11.15 586
14 [python] httplib — HTTP protocol client¶ 졸리운_곰 2016.11.15 548
13 python torrent 자동 다운로드 : How to automatically search and download torrents with Python and Scrapy 졸리운_곰 2016.11.02 875
12 Apache와 Python 연동하기 졸리운_곰 2016.10.16 1612
11 HTML Scraping [파이썬 웹 스크래핑 간단 메뉴얼] 졸리운_곰 2016.07.30 766
» Anonymous Web Scraping using Python and Tor file 졸리운_곰 2016.05.04 1284
9 Using TOR with Python file 졸리운_곰 2016.05.04 1911
8 Crawling anonymously with Tor in Python 졸리운_곰 2016.05.04 542
7 scrapy : 크롤링 해보기 file 졸리운_곰 2016.04.21 719
6 Python으로 공공데이터 오픈 API 활용하기 졸리운_곰 2016.01.14 1225
5 Title: HTML Scraper using pyhton 졸리운_곰 2015.09.01 400
4 위대한 LG SMART SMA LG 빅데이터 플랫폼 [파이썬 웹 크롤러] 웹 스파이더 file 졸리운_곰 2015.05.02 1154
3 위대한 LG SMART SMA LG 빅데이터 플랫폼 [소셜 웹 마이닝] 데이터 마이닝, 웹 마이닝 file 졸리운_곰 2015.05.02 775
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED