- 전체
- Python 일반
- Python 수학
- Python 그래픽
- Python 자료구조
- Python 인공지능
- Python 인터넷
- Python SAGE
- wxPython
- TkInter
- iPython
- wxPython
- pyQT
- Jython
- django
- flask
- blender python scripting
- python for minecraft
- Python 데이터 분석
- Python RPA
- cython
- PyCharm
- pySide
- kivy (python)
Python 인터넷 증권뉴스 데이터 수집(1.5/3.0)
2018.02.18 23:26
증권뉴스 데이터 수집(1.5/3.0)
2017. 5. 13. 21:06
1/3 편에서 제공한 코드가 구조화가 잘 되어 있지 않아서
기능별로 리팩토링 했습니다.
import time
import re
import requests
def get_date():
"""
수집 대상 날짜를 사용자로부터 키보드로 입력 받아 돌려준다.
:return: 날짜
"""
target_date = input("Enter date to Crawl news article urls (YYYYMMDD): ")
return target_date
def create_output_file(target_date) :
"""
출략 파일을 생성하고 파일 객체를 돌려준다.
:param target_date: 날짜
:return: 파일객체
"""
output_file_name = "article_urls" + target_date + ".txt"
output_file = open(output_file_name, "w", encoding="utf-8")
return output_file
def get_html(target_date, page_num):
"""
주어진 날짜와 페이지 번호에 해당하는 페이지 URL에 접근하여 HTML을 돌려준다.
:param target_date:
:param page_num:
:return:
"""
user_agent = "'Mozilla/5.0"
headers ={"User-Agent" : user_agent}
page_url = "http://news.naver.com/main/list.nhn?sid2=258&sid1=101&mid=shm&mode=LS2D&date=" + \
str(target_date) + "&page=" + str(page_num) + ""
response = requests.get(page_url, headers=headers)
html = response.text
return html
def ext_news_article_urls(html):
"""
주어진 html에서 기사 url을 추출하여 돌려준다.
:param html:
:return:
"""
url_frags = re.findall('<a href="(.*?)"', html)
news_article_urls=[]
for url_frag in url_frags:
if "sid1=101&sid2=258" in url_frag and "aid" in url_frag:
news_article_urls.append(url_frag)
else :
continue
return news_article_urls
def write_news_article_urls(output_file, urls):
"""
기사 URL들을 출력 파일에 기록한다.
:param output_file:
:param urls:
:return:
"""
for url in urls:
print(url, file=output_file)
def pause():
"""
2초동안 쉰다.
:return:
"""
time.sleep(2)
def close_output_file(output_file):
"""
출력파일을 닫는다.
:param output_file:
:return:
"""
output_file.close()
def main():
"""
사용자로부터 수집대상 날짜를 입력받아 해당 날짜의 네이버 경제 뉴스 기사 URL을 수집한다.
:return:
"""
target_date = get_date()
output_file = create_output_file(target_date)
page_num = 1
max_page_num = 100
while True :
html = get_html(target_date, page_num)
if page_num>=max_page_num:
break
urls = ext_news_article_urls(html)
write_news_article_urls(output_file, urls)
page_num+=1
pause()
close_output_file(output_file)
main()
[출처] 증권뉴스 데이터 수집(1.5/3.0)|작성자 엉드루
본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
댓글 0
| 번호 | 제목 | 글쓴이 | 날짜 | 조회 수 |
|---|---|---|---|---|
| 11 |
[PyQT5] PyQt5 Mainwindow에 Qt Designer를 사용한 graph Widget 추가
| 졸리운_곰 | 2024.06.06 | 476 |
| 10 |
[PyQT5] [Python with Pyqt5] Widget에다가 그래프 넣기 (Feat. Matplotlib)
| 졸리운_곰 | 2024.06.06 | 437 |
| 9 |
[PyQT5] 파이썬(Python)PyQt5 - QMessageBox 사용하기
| 졸리운_곰 | 2024.06.06 | 455 |
| 8 |
[PyQT5] UI Designer 에서 Tab Widget 생성 하기
| 졸리운_곰 | 2024.06.06 | 420 |
| 7 |
[PyQT5] UI Designer 에서 리사이즈 시 같이 확장하기
| 졸리운_곰 | 2024.06.06 | 537 |
| 6 |
[PyQT5] UI Designer 에서 Grid Layout 배치 해보기
| 졸리운_곰 | 2024.06.06 | 474 |
| 5 |
[pyQT] 「Python : PyQt5」 Qt Designer : .ui → .py 변환
| 졸리운_곰 | 2024.06.05 | 624 |
| 4 |
[pyQT] Dev/python/ [pyqt5] 프로그램창을 항상 가장 위에 있게 하면서 동시에 타이틀 바도 없게 하려면?
| 졸리운_곰 | 2024.06.01 | 542 |
| 3 |
[pyQT] QtDesigner에서 리소스 편집기 사용하기
| 졸리운_곰 | 2024.05.28 | 499 |
| 2 |
[pyQT] Qt Resource 파일 (.qrc) 적용방법
| 졸리운_곰 | 2024.05.28 | 421 |
| 1 | [pyQT] from PyQt5.QtChart import QLineSeries, QChart, QValueAxis, QDateTimeAxis ImportError: DLL load failed while importing QtChart: 지정된 모듈을 찾을 수 없습니다. | 졸리운_곰 | 2024.01.28 | 430 |

