Web scraping in R: A tutorial using Super Bowl Data

David Radcliffe

January 18, 2016

Hadley Wickham created an R package called rvest that makes it easy to scrape static web pages. In this note, I will show how to use rvest to extract a table with Super Bowl data from a web page.

We will use the optional R packages rvest, stringr, and tidyr. These packages must be installed before running this tutorial.

install.packages('rvest')
install.packages('stringr')
install.packages('tidyr')

Once the packages have been installed, they can be loaded using the library command.

library(rvest)
library(stringr)
library(tidyr)

We use the read_html function to read a web page. This function is provided by the xml2 package, which was loaded automatically when we loaded rvest.

url <- 'http://espn.go.com/nfl/superbowl/history/winners'
webpage <- read_html(url)

Next, we use the functions html_nodes and html_table (from rvest) to extract the HTML table element and convert it to a data frame.

sb_table <- html_nodes(webpage, 'table')
sb <- html_table(sb_table)[[1]]
head(sb)
##                               X1            X2
## 1 Super Bowl Winners and Results          <NA>
## 2                            NO.          DATE
## 3                              I Jan. 15, 1967
## 4                             II Jan. 14, 1968
## 5                            III Jan. 12, 1969
## 6                             IV Jan. 11, 1970
##                              X3                            X4
## 1                          <NA>                          <NA>
## 2                          SITE                        RESULT
## 3 Los Angeles Memorial Coliseum  Green Bay 35, Kansas City 10
## 4           Orange Bowl (Miami)      Green Bay 33, Oakland 14
## 5           Orange Bowl (Miami) New York Jets 16, Baltimore 7
## 6  Tulane Stadium (New Orleans)   Kansas City 23, Minnesota 7

We remove the first two rows, and set the column names.

경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

sb <- sb[-(1:2), ]
names(sb) <- c("number", "date", "site", "result")
head(sb)
##   number          date                          site
## 3      I Jan. 15, 1967 Los Angeles Memorial Coliseum
## 4     II Jan. 14, 1968           Orange Bowl (Miami)
## 5    III Jan. 12, 1969           Orange Bowl (Miami)
## 6     IV Jan. 11, 1970  Tulane Stadium (New Orleans)
## 7      V Jan. 17, 1971           Orange Bowl (Miami)
## 8     VI Jan. 16, 1972  Tulane Stadium (New Orleans)
##                          result
## 3  Green Bay 35, Kansas City 10
## 4      Green Bay 33, Oakland 14
## 5 New York Jets 16, Baltimore 7
## 6   Kansas City 23, Minnesota 7
## 7       Baltimore 16, Dallas 13
## 8            Dallas 24, Miami 3

It is traditional to use Roman numerals to refer to Super Bowls, but Arabic numerals are more convenient to work with. We will also convert the date to a standard format.

sb$number <- 1:49
sb$date <- as.Date(sb$date, "%B. %d, %Y")
head(sb)
##   number       date                          site
## 3      1 1967-01-15 Los Angeles Memorial Coliseum
## 4      2 1968-01-14           Orange Bowl (Miami)
## 5      3 1969-01-12           Orange Bowl (Miami)
## 6      4 1970-01-11  Tulane Stadium (New Orleans)
## 7      5 1971-01-17           Orange Bowl (Miami)
## 8      6 1972-01-16  Tulane Stadium (New Orleans)
##                          result
## 3  Green Bay 35, Kansas City 10
## 4      Green Bay 33, Oakland 14
## 5 New York Jets 16, Baltimore 7
## 6   Kansas City 23, Minnesota 7
## 7       Baltimore 16, Dallas 13
## 8            Dallas 24, Miami 3

The result column should be split into four columns – the winning team’s name, the winner’s score, the losing team’s name, and the loser’s score. We start by splitting the results column into two columns at the comma. This operation uses the separate function from the tidyr package.

sb <- separate(sb, result, c('winner', 'loser'), sep=', ', remove=TRUE)
head(sb)
##   number       date                          site           winner
## 3      1 1967-01-15 Los Angeles Memorial Coliseum     Green Bay 35
## 4      2 1968-01-14           Orange Bowl (Miami)     Green Bay 33
## 5      3 1969-01-12           Orange Bowl (Miami) New York Jets 16
## 6      4 1970-01-11  Tulane Stadium (New Orleans)   Kansas City 23
## 7      5 1971-01-17           Orange Bowl (Miami)     Baltimore 16
## 8      6 1972-01-16  Tulane Stadium (New Orleans)        Dallas 24
##            loser
## 3 Kansas City 10
## 4     Oakland 14
## 5    Baltimore 7
## 6    Minnesota 7
## 7      Dallas 13
## 8        Miami 3

Finally, we split off the scores from the winner and loser columns. The function str_extract from the stringr package finds a substring matching a pattern. In this case, the pattern is a sequence of 1 or more digits at the end of a line.

pattern <- " \d+$"
sb$winnerScore <- as.numeric(str_extract(sb$winner, pattern))
sb$loserScore <- as.numeric(str_extract(sb$loser, pattern))
sb$winner <- gsub(pattern, "", sb$winner)
sb$loser <- gsub(pattern, "", sb$loser)
head(sb)
##   number       date                          site        winner
## 3      1 1967-01-15 Los Angeles Memorial Coliseum     Green Bay
## 4      2 1968-01-14           Orange Bowl (Miami)     Green Bay
## 5      3 1969-01-12           Orange Bowl (Miami) New York Jets
## 6      4 1970-01-11  Tulane Stadium (New Orleans)   Kansas City
## 7      5 1971-01-17           Orange Bowl (Miami)     Baltimore
## 8      6 1972-01-16  Tulane Stadium (New Orleans)        Dallas
##         loser winnerScore loserScore
## 3 Kansas City          35         10
## 4     Oakland          33         14
## 5   Baltimore          16          7
## 6   Minnesota          23          7
## 7      Dallas          16         13
## 8       Miami          24          3

Our data frame is looking pretty good, so we write it to a CSV (comma-separated value) file.

write.csv(sb, 'superbowl.csv', row.names=F)

 

[출처] https://rpubs.com/Radcliffe/superbowl

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
공지 오라클 기본 샘플 데이터베이스 졸리운_곰 2014.01.02 86144
공지 [SQL컨셉] 서적 "SQL컨셉"의 샘플 데이타 베이스 SAMPLE DATABASE of ORACLE 가을의 곰을... 2013.02.10 78639
공지 [G_SQL] Sample Database 가을의 곰을... 2012.05.20 95368
51 오라클 Select 해서 Update 하기 졸리운_곰 2018.01.22 1212
50 디폴트 세팅의 함정과 오라클 파라미터 file 졸리운_곰 2017.07.15 1685
49 그림으로 배우는 ‘공정쿼리와 인덱스 생성도’ file 졸리운_곰 2017.07.15 1538
48 오라클 랜덤 함수와 사용자 정의 함수 file 졸리운_곰 2017.07.15 3232
47 개발자들이 자주 접하는 오라클 에러 메세지 file 졸리운_곰 2017.07.15 4000
46 오라클 DICTIONARY를 활용한 DB툴 프로그램 ‘FreeSQL’ file 졸리운_곰 2017.07.15 1182
45 알면 유용한 오라클 기능들 file 졸리운_곰 2017.07.15 1565
44 알면 유용한 오라클 기능 ‘GATHER_PLAN_STATISTICS’ file 졸리운_곰 2017.07.15 1859
43 개발자들의 영원한 숙제 ‘NULL 이야기’ file 졸리운_곰 2017.07.15 4343
42 오라클 플랜을 보는 법 file 졸리운_곰 2017.07.15 1880
41 반드시 알아야 하는 오라클 힌트절 7가지 file 졸리운_곰 2017.07.15 1566
40 그림으로 배우는 ‘오라클 조인의 방식’ 이야기 file 졸리운_곰 2017.07.15 1912
39 오라클 옵티마이저 ‘CBO와 RBO’ 이해하기 file 졸리운_곰 2017.07.09 1801
38 만능 쿼리와 한 방 쿼리 file 졸리운_곰 2017.07.09 1514
37 퀴리 최적화 및 튜닝을 위한 오라클 공정쿼리 작성법 file 졸리운_곰 2017.07.09 1590
36 누구도 알려주지 않았던 ‘오라클 쿼리 작성의 비법’ file 졸리운_곰 2017.07.09 1313
35 누구도 알려주지 않았던 ‘오라클 인덱스 생성도’의 비밀 file 졸리운_곰 2017.07.09 1600
34 oracle 패스워드 유효기간 만료 복구 졸리운_곰 2017.05.06 1648
33 Oracle Instant Client 이용하여 Toad 사용하기 file 졸리운_곰 2017.05.06 1221
32 [Oracle] 1. 각 테이블 코멘트 조회하기 / 2. 테이블의 컬럼정보 조회하기 졸리운_곰 2017.04.17 1190
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED