- 전체
- Sample DB
- database modeling
- [표준 SQL] Standard SQL
- G-SQL
- 10-Min
- ORACLE
- MS SQLserver
- MySQL
- SQLite
- postgreSQL
- 데이터아키텍처전문가 - 국가공인자격
- 데이터 분석 전문가 [ADP]
- [국가공인] SQL 개발자/전문가
- NoSQL
- hadoop
- hadoop eco system
- big data (빅데이터)
- stat(통계) R 언어
- XML DB & XQuery
- spark
- DataBase Tool
- 데이터분석 & 데이터사이언스
- Engineer Quality Management
- [기계학습] machine learning
- 데이터 수집 및 전처리
- 국가기술자격 빅데이터분석기사
- 암호화폐 (비트코인, cryptocurrency, bitcoin)
[SPARK][Python][pySpark][아콘 소프트][나무기술] How to Run a Spark Standalone Job
How to Run a Spark Standalone Job
Overview
This is a minimal Spark script that imports PySpark, initializes a SparkContext and performs a distributed calculation on a Spark cluster in standalone mode.
Who is this for?
This how-to is for users of a Spark cluster that has been configured in standalone mode who wish to run Python code.
Spark Standalone Summary
Before you start
To execute this example, download the cluster-spark-basic.py example script to the cluster node where you submit Spark jobs.
For this example, you’ll need Spark running with the standalone scheduler. You can install Spark using an enterprise Hadoop distribution such as Cloudera CDH or Hortonworks HDP. Some additional configuration might be necessary to use Spark in standalone mode.
Modifying the script
After downloading the cluster-spark-basic.py example script open the file in a text editor on your cluster. Replace HEAD_NODE_HOSTNAME with the hostname of the head node of the Spark cluster.
# cluster-spark-basic.py from pyspark import SparkConf from pyspark import SparkContext conf = SparkConf() conf.setMaster('spark://HEAD_NODE_HOSTNAME:7077') conf.setAppName('spark-basic') sc = SparkContext(conf=conf) def mod(x): import numpy as np return (x, np.mod(x, 2)) rdd = sc.parallelize(range(1000)).map(mod).take(10) print rdd
Let’s analyze the contents of the spark-basic.rst example script. The first code block contains imports from PySpark.
The second code block initializes the SparkContext and sets the application name.
The third code block contains the analysis code that calculates the modulus of a range of numbers up to 1000 using the NumPy package and returns/prints the first 10 results.
Note: you may have to install NumPy with acluster conda install numpy.
Running the job
You can run this script by submitting it to your cluster for execution using spark-submit or by running this command
python cluster-spark-basic.py
The output from the above command shows the first ten values that were returned from the cluster-spark-basic.py script.
16/05/05 22:26:53 INFO spark.SparkContext: Running Spark version 1.6.0 [...] 16/05/05 22:27:03 INFO scheduler.TaskSetManager: Starting task 0.0 in stage 0.0 (TID 0, localhost, partition 0,PROCESS_LOCAL, 3242 bytes) 16/05/05 22:27:04 INFO storage.BlockManagerInfo: Added broadcast_0_piece0 in memory on localhost:46587 (size: 2.6 KB, free: 530.3 MB) 16/05/05 22:27:04 INFO scheduler.TaskSetManager: Finished task 0.0 in stage 0.0 (TID 0) in 652 ms on localhost (1/1) 16/05/05 22:27:04 INFO cluster.YarnScheduler: Removed TaskSet 0.0, whose tasks have all completed, from pool 16/05/05 22:27:04 INFO scheduler.DAGScheduler: ResultStage 0 (runJob at PythonRDD.scala:393) finished in 4.558 s 16/05/05 22:27:04 INFO scheduler.DAGScheduler: Job 0 finished: runJob at PythonRDD.scala:393, took 4.951328 s [(0, 0), (1, 1), (2, 0), (3, 1), (4, 0), (5, 1), (6, 0), (7, 1), (8, 0), (9, 1)]
Troubleshooting
If something goes wrong consult the FAQ / Known issues page.
[출처] https://docs.anaconda.com/anaconda-cluster/howto/spark-basic/
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
댓글 0
| 번호 | 제목 | 글쓴이 | 날짜 | 조회 수 |
|---|---|---|---|---|
| 공지 | 오라클 기본 샘플 데이터베이스 | 졸리운_곰 | 2014.01.02 | 86134 |
| 공지 | [SQL컨셉] 서적 "SQL컨셉"의 샘플 데이타 베이스 SAMPLE DATABASE of ORACLE | 가을의 곰을... | 2013.02.10 | 78635 |
| 공지 | [G_SQL] Sample Database | 가을의 곰을... | 2012.05.20 | 95360 |

