[R lang 크롤링] R - 크롤링

rvest 라이브러리 설치 및 임포트

 

install.packages("rvest")

library(rvest)

 


 

만약 오류가 나면 iconv 인코딩 하면된다.

# UTF-8 되어있으면 문제 없음

html <- read_html(iconv("주소",  from = 'euc-kr',to='cp949'),encoding='cp949')

character 인코딩 체계가 어떻게 되어있는지 확인(확률값)

guess_encoding(html)

 

 

xpath 텍스트 뽑아내기

html_nodes(html,xpath='//*[@id="old_content"]/table/tbody/tr/td[2]/text()')

 


 

예시 : 인공지능 신문 - 의료 AI 기사 크롤링

 

# html 주소 불러오기

html <- read_html("http://www.aitimes.kr/news/articleList.html")

# url 저장
# 한개를 찾을땐 node / 여러개 찾을땐 nodes 
# #은 id를 의미 / .은 class를 의미함

url <- html_nodes(html,".list-titles")%>%
  html_nodes('a')%>%
  html_attr('href')

# 기사 내용 불러오기

news <- c()

for(i in 1:length(url)){

  html <- read_html(paste0("http://www.aitimes.kr",url[i]))

  text <- html_node(html,'#article-view-content-div')%>%

    html_text()

  news <- c(news,text)

}

# 기사 내용 정제작업

t <- c()

for(i in 1:length(news)){

  x22 <- SimplePos22(news[i])

  x22

  str_match(x22, "[A-z가-힣]+/N")

  word_nn <- as.vector(na.omit(str_match(x22, "([A-z가-힣]+)/N")[,2]))

  t <- c(t,word_nn)

  word <- table(t)

}

# 워드클라우드로 띄워보기

wordcloud2(word)

[출처] https://truman.tistory.com/161

 

 

 

 

경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

[R lang 크롤링] R - 크롤링

rvest 라이브러리 설치 및 임포트

 

install.packages("rvest")

library(rvest)

 


 

만약 오류가 나면 iconv 인코딩 하면된다.

# UTF-8 되어있으면 문제 없음

html <- read_html(iconv("주소",  from = 'euc-kr',to='cp949'),encoding='cp949')

character 인코딩 체계가 어떻게 되어있는지 확인(확률값)

guess_encoding(html)

 

 

xpath 텍스트 뽑아내기

html_nodes(html,xpath='//*[@id="old_content"]/table/tbody/tr/td[2]/text()')

 


 

예시 : 인공지능 신문 - 의료 AI 기사 크롤링

 

# html 주소 불러오기

html <- read_html("http://www.aitimes.kr/news/articleList.html")

# url 저장
# 한개를 찾을땐 node / 여러개 찾을땐 nodes 
# #은 id를 의미 / .은 class를 의미함

url <- html_nodes(html,".list-titles")%>%
  html_nodes('a')%>%
  html_attr('href')

# 기사 내용 불러오기

news <- c()

for(i in 1:length(url)){

  html <- read_html(paste0("http://www.aitimes.kr",url[i]))

  text <- html_node(html,'#article-view-content-div')%>%

    html_text()

  news <- c(news,text)

}

# 기사 내용 정제작업

t <- c()

for(i in 1:length(news)){

  x22 <- SimplePos22(news[i])

  x22

  str_match(x22, "[A-z가-힣]+/N")

  word_nn <- as.vector(na.omit(str_match(x22, "([A-z가-힣]+)/N")[,2]))

  t <- c(t,word_nn)

  word <- table(t)

}

# 워드클라우드로 띄워보기

wordcloud2(word)

[출처] https://truman.tistory.com/161

 

 

 

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
공지 오라클 기본 샘플 데이터베이스 졸리운_곰 2014.01.02 86130
공지 [SQL컨셉] 서적 "SQL컨셉"의 샘플 데이타 베이스 SAMPLE DATABASE of ORACLE 가을의 곰을... 2013.02.10 78632
공지 [G_SQL] Sample Database 가을의 곰을... 2012.05.20 95353
33 [spark][sparksql][odbc][jdbc] JDBC and ODBC drivers and configuration parameters file 졸리운_곰 2021.04.14 4127
32 [spark][pyspark][php] Natively Connect to Spark Data in PHP 졸리운_곰 2021.04.14 1702
31 [SPARK][Python][pySpark][아콘 소프트][나무기술] Real-world Python workloads on Spark: Standalone clusters : 스파크 예제 논란, driver-host 불필요 file 졸리운_곰 2021.04.03 1691
30 [Spark] Apache Spark Cluster(Standalone) 스파크 클러스터 스텐드 얼론 구축 졸리운_곰 2021.03.28 1523
29 [Spark][머신러닝] Apache Spark-Python vs Scala 성능 비교 file 졸리운_곰 2021.03.21 1152
28 [Spark][MSA] Apache Spark - Key/Value Paris (Pair RDD) 졸리운_곰 2021.03.21 1583
27 [Spark][머신러닝] Apache Spark - RDD (Resilient Distributed DataSet) Persistence file 졸리운_곰 2021.03.21 1312
26 [Spark][머신러닝] Apache Spark - RDD (Resilient Distributed DataSet) 이해하기 - #2 file 졸리운_곰 2021.03.21 1085
25 [Spark][머신러닝] Apache Spark - RDD (Resilient Distributed DataSet) 이해하기 - #1 file 졸리운_곰 2021.03.21 1678
24 [Spark][머신러닝] Apache Spark 소개 - 스파크 스택 구조 file 졸리운_곰 2021.03.21 1318
23 [Spark] cache()와 persist()의 차이 file 졸리운_곰 2021.03.16 1550
22 [Spark] Spark - RDD vs Dataframes vs Datasets 우리는 언제, 왜 RDD, Dataframes, Datasets를 사용해야 할까? file 졸리운_곰 2021.03.15 1317
21 [Spark & Oracle] Reading Data From Oracle Database With Apache Spark file 졸리운_곰 2021.03.15 1148
20 [spark][pySpark] 스파크 튜토리얼 - 스파크 SQL file 졸리운_곰 2021.03.15 1856
19 [spark][flask][python] Machine learning at Scale using Pyspark & deployment using AzureML/Flask file 졸리운_곰 2021.03.14 2034
18 [pySpark, 파이썬 spark] Best Practices Writing Production-Grade PySpark Jobs file 졸리운_곰 2021.03.14 1691
17 [apache spark] 아파치 스파크 Data Sharing between multiple Spark Jobs in Databricks file 졸리운_곰 2021.03.13 1792
16 [Apache Spark] Spark SQL 아파치 스파크 SQL 개요 졸리운_곰 2021.03.13 1282
15 [spark] Apache Livy: A REST Interface for Apache Spark file 졸리운_곰 2021.03.12 1354
14 [spark] Spark 및 Oracle 데이터베이스 file 졸리운_곰 2021.03.06 1410
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED