- 전체
- Python 일반
- Python 수학
- Python 그래픽
- Python 자료구조
- Python 인공지능
- Python 인터넷
- Python SAGE
- wxPython
- TkInter
- iPython
- wxPython
- pyQT
- Jython
- django
- flask
- blender python scripting
- python for minecraft
- Python 데이터 분석
- Python RPA
- cython
- PyCharm
- pySide
- kivy (python)
Python 인터넷 [Python, 인터넷] 네이버 뉴스 기사 크롤링
2021.05.23 17:41
[Python, 인터넷] 네이버 뉴스 기사 크롤링
네이버 뉴스 기사 크롤링
네이버 뉴스에 접속하여, 원하는 키워드와 원하는 기간을 설정하여 나오는 모든 기사 검색 결과를 크롤링하는 방법입니다. 기사 타이틀, 기사 등록일, 언론사 및 정확하지는 않지만 기자 이름과 기자 이메일까지 가져옵니다 마지막으로 엑셀 및 csv로 저장합니다.
#코드
import requests
from bs4 import BeautifulSoup
import math
import pandas
import re
resultList = []
url = "https://search.naver.com/search.naver?"
params = {
"where": 'news',
# 네이버 기사 검색 값
"query": '매틱 네트워크 스테이킹',
# 페이지네이션 값
"start": 0,
# "nso": 'so:r,p:1y,a:all'
}
# nso: so: r, p: 1y, a: all -> 최근 1년
# nso: so: r, p: 6m, a: all -> 최근 6개월
# nso: so: r, p: 1d, a: all -> 1일
# 없으면 전체 검색
# headers={'User-Agent': 'Mozilla/5.0'} -> 안티 크롤링 회피
raw = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, params=params)
html = BeautifulSoup(raw.text, "html.parser")
# 검색결과 html body
articles = html.select("ul.type01 > li")
# 전체 기사 수
totalCount = html.select("div.section_head > div.title_desc > span")[0].text.split(' / ')[1][:-1]
for i in range(0, math.floor(int(totalCount)/10)+1)):
if i == 0:
params['start'] = i
else:
params['start'] = i * 10 + 1
raw = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}, params=params)
html = BeautifulSoup(raw.text, "html.parser")
articles = html.select("ul.type01 > li")
for ar in articles:
# 제목 값
title = ar.select_one("a._sp_each_title").text
# 검색된 기사의 url을 가져와서 다시 html을 get
articleUrl = ar.find("a")["href"]
innerRaw = requests.get(articleUrl, headers={'User-Agent': 'Mozilla/5.0'})
# 가져온 기사 html중 '기사', '@' string을 모두 가져온다
innerHtml = BeautifulSoup(innerRaw.text, "html.parser")
reporter = innerArticles = innerHtml(text=re.compile("기자"))
reporterEmail = innerArticles = innerHtml(text=re.compile("@"))
# 언론사 값
source = ar.select_one("span._sp_each_source").text
# 등록일 값
date = ar.select_one("dd.txt_inline").text.split(" ")[1]
res = {"title": title, "company": source,
"url": articleUrl, "date": date, "reporter": reporter, "reporterEmail": reporterEmail}
resultList.append(res)
# 검색된 기사 갯수
resultList.append({"totalCount": totalCount})
df = pandas.DataFrame(resultList)
df.to_csv('blockChain_articles.csv')
df.to_excel('blockChain_articles.xlsx')
[출처] https://kyounghwan01.github.io/blog/etc/python/naver-news-crawling/#%E1%84%8F%E1%85%A9%E1%84%83%E1%85%B3
본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
댓글 0
| 번호 | 제목 | 글쓴이 | 날짜 | 조회 수 |
|---|---|---|---|---|
| 11 |
[PyQT5] PyQt5 Mainwindow에 Qt Designer를 사용한 graph Widget 추가
| 졸리운_곰 | 2024.06.06 | 476 |
| 10 |
[PyQT5] [Python with Pyqt5] Widget에다가 그래프 넣기 (Feat. Matplotlib)
| 졸리운_곰 | 2024.06.06 | 437 |
| 9 |
[PyQT5] 파이썬(Python)PyQt5 - QMessageBox 사용하기
| 졸리운_곰 | 2024.06.06 | 455 |
| 8 |
[PyQT5] UI Designer 에서 Tab Widget 생성 하기
| 졸리운_곰 | 2024.06.06 | 420 |
| 7 |
[PyQT5] UI Designer 에서 리사이즈 시 같이 확장하기
| 졸리운_곰 | 2024.06.06 | 537 |
| 6 |
[PyQT5] UI Designer 에서 Grid Layout 배치 해보기
| 졸리운_곰 | 2024.06.06 | 474 |
| 5 |
[pyQT] 「Python : PyQt5」 Qt Designer : .ui → .py 변환
| 졸리운_곰 | 2024.06.05 | 624 |
| 4 |
[pyQT] Dev/python/ [pyqt5] 프로그램창을 항상 가장 위에 있게 하면서 동시에 타이틀 바도 없게 하려면?
| 졸리운_곰 | 2024.06.01 | 542 |
| 3 |
[pyQT] QtDesigner에서 리소스 편집기 사용하기
| 졸리운_곰 | 2024.05.28 | 499 |
| 2 |
[pyQT] Qt Resource 파일 (.qrc) 적용방법
| 졸리운_곰 | 2024.05.28 | 421 |
| 1 | [pyQT] from PyQt5.QtChart import QLineSeries, QChart, QValueAxis, QDateTimeAxis ImportError: DLL load failed while importing QtChart: 지정된 모듈을 찾을 수 없습니다. | 졸리운_곰 | 2024.01.28 | 430 |

