증권뉴스 데이터 수집(1.5/3.0)

 

2017. 5. 13. 21:06
URL 복사

 

경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

 
 

1/3 편에서 제공한 코드가 구조화가 잘 되어 있지 않아서
기능별로 리팩토링 했습니다.

import time
import re
import requests

def get_date():
    """
    수집 대상 날짜를 사용자로부터 키보드로 입력 받아 돌려준다.
    :return: 날짜
    """

    target_date = input("Enter date to Crawl news article urls (YYYYMMDD): ")

    return target_date

def create_output_file(target_date) :
    """
    출략 파일을 생성하고 파일 객체를 돌려준다.
    :param target_date: 날짜 
    :return: 파일객체
    """

    output_file_name = "article_urls" + target_date + ".txt"
    output_file = open(output_file_name, "w", encoding="utf-8")

    return output_file

def get_html(target_date, page_num):
    """
    주어진 날짜와 페이지 번호에 해당하는 페이지 URL에 접근하여 HTML을 돌려준다.
    :param target_date: 
    :param page_num: 
    :return: 
    """
    user_agent = "'Mozilla/5.0"
    headers ={"User-Agent" : user_agent}

    page_url = "http://news.naver.com/main/list.nhn?sid2=258&sid1=101&mid=shm&mode=LS2D&date=" + \
               str(target_date) + "&page=" + str(page_num) + ""

    response = requests.get(page_url, headers=headers)
    html = response.text

    return html

def ext_news_article_urls(html):
    """
    주어진 html에서 기사 url을 추출하여 돌려준다.
    :param html: 
    :return: 
    """

    url_frags = re.findall('<a href="(.*?)"', html)
    news_article_urls=[]

    for url_frag in url_frags:
        if  "sid1=101&sid2=258" in url_frag and "aid" in url_frag:
            news_article_urls.append(url_frag)
        else :
            continue

    return news_article_urls

def write_news_article_urls(output_file, urls):
    """
    기사 URL들을 출력 파일에 기록한다.
    :param output_file: 
    :param urls: 
    :return: 
    """
    for url in urls:
        print(url, file=output_file)

def pause():
    """
    2초동안 쉰다.
    :return: 
    """
    time.sleep(2)

def close_output_file(output_file):
    """
    출력파일을 닫는다.
    :param output_file: 
    :return: 
    """
    output_file.close()

def main():
    """
    사용자로부터 수집대상 날짜를 입력받아 해당 날짜의 네이버 경제 뉴스 기사 URL을 수집한다.
    :return: 
    """

    target_date = get_date()
    output_file = create_output_file(target_date)
    page_num = 1
    max_page_num = 100

    while True :
        html = get_html(target_date, page_num)

        if page_num>=max_page_num:
            break

        urls = ext_news_article_urls(html)
        write_news_article_urls(output_file, urls)
        page_num+=1
        pause()

    close_output_file(output_file)

main()



본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
42 Django에서 MySQL DB를 연동하기 pycharm file 졸리운_곰 2018.04.10 732
41 Python Flask 로 간단한 REST API 작성하기 file 졸리운_곰 2018.04.07 496
40 증권뉴스데이터 수집(3/3편) 졸리운_곰 2018.02.18 709
39 증권뉴스데이터 수집(2/3편) 졸리운_곰 2018.02.18 437
» 증권뉴스 데이터 수집(1.5/3.0) 졸리운_곰 2018.02.18 482
37 증권뉴스 데이터 수집(1/3) file 졸리운_곰 2018.02.18 601
36 python 활용 웹 사이트가 존재하는지 체크 : Python check if website exists 졸리운_곰 2018.01.16 446
35 파이썬3을 이용하여 코인원,빗썸,코빗의 가상화폐 시세정보를 불러오는 프로그램을 만들었다. file 졸리운_곰 2017.12.02 716
34 네이버 실시간 검색어를 자동 추출하는 방법 file 졸리운_곰 2017.11.14 631
33 Cinema 3 - (Extremely Simplified) Example of Microservices in Python file 졸리운_곰 2017.08.03 395
32 나만의 웹 크롤러 만들기 with Requests/BeautifulSoup file 졸리운_곰 2017.07.08 586
31 Web Scraping using Python / FinAlgML(놀러온특강) Python을 통한 웹 스크래핑 및 DB화 file 졸리운_곰 2017.07.08 728
30 PiP - Python in PHP 졸리운_곰 2017.05.06 893
29 Developing a RESTful micro service in Python file 졸리운_곰 2017.03.06 1120
28 BitTorrent 프로토콜의 동작원리 file 졸리운_곰 2017.02.26 1014
27 Torrent의 원리 file 졸리운_곰 2017.02.26 1445
26 How to automatically search and download torrents with Python and Scrapy 졸리운_곰 2017.02.26 774
25 Web scraping, article extraction and sentiment analysis with Scrapy, Goose and TextBlob 졸리운_곰 2017.02.26 395
24 [Python] 네이버 주식 종목별 일별 데이터 가져오기 file 졸리운_곰 2017.02.24 1869
23 [파이썬으로 웹 크롤러 만들기] 크롤링 시작하기(3/3) file 졸리운_곰 2017.02.16 705
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED