Flume or Kafka? Try both!

 

In the last few month, I’ve spent a lot of time answering one question: Flume or Kafka?

Its a pretty tough choice.
Flume has been around for a while, there are many users and happy production systems, its well supported by major Hadoop vendors, there are well known best practices on how to configure and tune it and its integration with both Hadoop and common data sources is unparalleled. If you are not a developer, its easy to configure a good data pipeline with Flume without writing a single line of code.

Kafka is slightly newer, but it offers impressive latency and throughput, scalability and high availability. Its a very well designed system. In addition, Kafka makes it easy pull data out of it into variety of systems – batch, streaming, applications, dashboards and even the command line.

In many cases the decision comes down to specific priorities:
Do you prefer possibility of integration with many data targets? Or an easy way to send data to Hadoop?
Do you prefer the flexibility of writing your own data producers or consumers? Or do you prefer a configuration-only solution without any programming involved?
Is it important for you to control every aspect of performance, availability and scalability? Or do you want a system that is mostly well-tuned out of the box, but may not squeeze every bit of possible performance?

After few of those discussions, we realized that this is silly. Why do we need to choose between a system that is great for developers and a system that is great for administrators? Why choose between fantastic scalability and flexibility and a great collection of data integration sources and sinks?

Why can’t we have it all?

 

This is why we helped write Flafka (unofficial code name) – a set of Kafka source and sink for Flume.

Flafka serves two important use-cases:

1. You have Kafka. You have apps pushing data into Kafka. But you need to get the data to Hadoop. Not just HDFS. You need HBase and Solr too. Maybe you even need to tweak the data a bit on the way – mask sensitive data for example or flag suspicious events.  Use Kafka-source with HDFS, HBase or Solr sinks. And add an interceptor to massage the data for you.

2. You need to land data from HTTP, log files, syslog or vmstat to Kafka. You don’t want to write your own producer and rather use existing tried-and-true solution. Use a Flume source with a Kafka sink to land data in Kafka without having to figure out how to write a producer.

경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

Since Flume is configuration-only, Flafka allows non-developers to easily use Kafka and to integrate Kafka with Hadoop.
But we didn’t want to sacrifice flexibility for simplicity, so the Kafka source will support any configuration that a Kafka consumer will accept, and Kafka sink can be configured just like any producer. Out of the box, we tuned Flafka for reliability rather than pure speed, reasoning that preventing data loss is our most important mission. If you have different trade-offs and preferences, it is easy to test them with Flafka.

Like all other Flume sources and sink, Flafka is designed to work in batches – you can configure batch sizes and the Kafka source will write data in batches to the Flume channel, and the Kafka sink will read data in batches from the channel and send them to Kafka. This design will increase throughput and improve resource utilization, but will also increase latency. You can tune the balance between throughput and latency to match your requirements by experimenting with different batch sizes.

You can even configure multiple Kafka sources to read from the same topic. If you configure all of the with the same groupId, each source will read a different set of partitions, which will lead to improved throughput. This is true even if the sources are configured on different agents. If one of the agents crashes, the remaining agents will automatically re-balance the partitions between them, which reduces delays and data loss.

At the moment, there is no Flume release with Flafka, so you will need to build Flafka from Flume’s source repository and build it:

git clone https://github.com/apache/flume.git flume-local
cd flume-local
git checkout trunk
mvn package -DskipTests


After building, simply copy the jars from flume-ng-dist/target/apache-flume-1.6.0-SNAPSHOT-bin/apache-flume-1.6.0-SNAPSHOT-bin/lib  to /usr/lib/flume-ng/lib and restart your Flume agent.

To get you started, here is an example of how I configure Flume to read data from a Kafka topic and write it to an HDFS directory. The directory will have the same name as the topic, with subdirectories for each day of data:

 tier1.sources  = source1
 tier1.channels = channel1
 tier1.sinks    = sink1
 
 tier1.sources.source1.type = org.apache.flume.source.kafka.KafkaSource
 tier1.sources.source1.zookeeperConnect = shapira-1:2181
 tier1.sources.source1.topic = shapira
 tier1.sources.source1.groupId = flume
 tier1.sources.source1.channels = channel1
 tier1.sources.source1.interceptors = i1
 tier1.sources.source1.interceptors.i1.type = timestamp
 tier1.sources.source1.kafka.consumer.timeout.ms = 100
 
 tier1.channels.channel1.type   = memory
 tier1.channels.channel1.capacity = 10000
 tier1.channels.channel1.transactionCapacity = 1000
 
 
 tier1.sinks.sink1.type         = hdfs
 tier1.sinks.sink1.hdfs.path    = /tmp/shapira/kafka/%{topic}/%y-%m-%d
 tier1.sinks.sink1.hdfs.rollInterval = 5
 tier1.sinks.sink1.hdfs.rollSize = 0
 tier1.sinks.sink1.hdfs.rollCount = 0
 tier1.sinks.sink1.hdfs.fileType = DataStream
 tier1.sinks.sink1.channel      = channel1 

And here’s a Flume configuration for getting data from vmstat to Kafka:

 tier1.sources  = source1
 tier1.channels = channel1
 tier1.sinks    = sink1
 
 tier1.sources.source1.type = exec
 tier1.sources.source1.command = /usr/bin/vmstat 1
 tier1.sources.source1.channels = channel1
 
 tier1.channels.channel1.type   = memory
 tier1.channels.channel1.capacity = 10000
 tier1.channels.channel1.transactionCapacity = 1000
 
 
 tier1.sinks.sink1.type         = org.apache.flume.sink.kafka.KafkaSink
 tier1.sinks.sink1.topic = sink1
 tier1.sinks.sink1.brokerList = kafkagames-1:9092,kafkagames-2:9092
 tier1.sinks.sink1.channel = channel1
 tier1.sinks.sink1.batchSize = 20

I hope you’ll enjoy the ease of use that Flume brings to Kafka. Let us know your experiences in the comments, and open Jira tickets if you run into issues, have ideas for improvements or want to contribute a patch.

 

[출처] http://ingest.tips/2014/09/26/trying-to-decide-between-flume-and-kafka-try-both/

 

 

 

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
공지 오라클 기본 샘플 데이터베이스 졸리운_곰 2014.01.02 86180
공지 [SQL컨셉] 서적 "SQL컨셉"의 샘플 데이타 베이스 SAMPLE DATABASE of ORACLE 가을의 곰을... 2013.02.10 78666
공지 [G_SQL] Sample Database 가을의 곰을... 2012.05.20 95401
42 블록체인 기반 플랫폼 비즈니스를 이해하자 졸리운_곰 2017.09.10 1554
41 다운타임 없는 서비스 구현 패턴 file 졸리운_곰 2017.09.10 1229
40 테이블의 수직분할과 수평분할에 대한 이해 file 졸리운_곰 2017.09.10 3908
39 정규화와 응집도에 대한 고찰 file 졸리운_곰 2017.05.28 1782
38 머신러닝 새 도전…“클라우드를 벗어나라" file 졸리운_곰 2017.05.28 1711
37 보안성 높이는 공공 거래장부 블록체인 file 졸리운_곰 2017.05.28 1757
36 디지털, 속도의 전쟁 VS 데이터, 품질의 전쟁 file 졸리운_곰 2017.05.05 1515
35 04. 데이터 모델링의 3단계 진행 file 졸리운_곰 2016.03.15 1659
34 데이터베이스 설계의 기본 원리.pdf file 졸리운_곰 2016.03.15 2353
33 실체유형(Entity Type) 정의 사항 및 도출 file 졸리운_곰 2015.05.21 2198
32 마농의 SQL 백문백답: 단순하고 쉽게 작성하는 SQL 노하우 [1회] file 졸리운_곰 2015.05.21 1829
31 sql개발자-sql전문가자격시험 시험 예제.pdf file 졸리운_곰 2015.02.15 2302
30 데이터베이스 선정에는 비밀이 있다 - 4부 졸리운_곰 2015.01.15 2096
29 데이터베이스 선정에는 비밀이 있다 - 3부 졸리운_곰 2015.01.15 2042
28 데이터베이스 선정에는 비밀이 있다 - 2부 졸리운_곰 2015.01.15 1827
27 데이터베이스 선정에는 비밀이 있다 - 1부 졸리운_곰 2015.01.15 2391
26 지금 우리에게 필요한 것은 데이터베이스 성능 최적화이다 (2부) 졸리운_곰 2015.01.15 1484
25 우리에게 필요한 것은 데이터베이스 성능 (1부) 졸리운_곰 2015.01.15 1778
24 21회 결과 secret 졸리운_곰 2014.11.10 0
23 [데이터아키텍쳐준전문가] 시험 fail 자료 secret 졸리운_곰 2014.08.31 0
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED