5 basic steps for Key-value store database design

by Laszlo Wagner

Nowadays NoSql databases are very popular. There are lots of vendors with different technical approach. CAP theorem applies for most of them:

  • Consistency – we can see same data at the same time on all nodes
  • Availability - every request is guaranteed to be processed and responded regardless of success factor
  • Partition tolerance - system remains accessible if some messages are lost or failure of part of the system (e.g. we lost some nodes)

Sometimes we can tune two features of the three. However tons of documents can be found about NoSQL technology but data modeling part is not the best studied part. I also found different implementations supporting different techniques, having some sort of advantages or shortcomings. I picked one of my favorite key value store implementation, Redis, to show some real working modeling techniques. Our demo database must serve a Web application, store users and their click actions.

  1. Define objects

In the first step of data modeling we need to identify objects. In the relational database design they will be tables most likely. In Redis they can be strings, lists, hashes, etc. Regardless the physical implementation we need to separate objects by good names / keys. We can identify USER and CLICK as objects in our example. We will use USER: and CLICK: as a prefix. This will be part of our naming convention.

  1. Define identifiers

Since Redis is a key value store we need to define keys to store data. We defined two objects above we always have to think how they are connected. It seems obvious if we store users we want to know where and when they clicked. This is the purpose of our demo database. We know from the application designer the LOGIN_ID is a unique identifier for the user. So we will store user data in this format: USER:<LOGIN_ID> e.g. USER:100AB. That was easy. For CLICK object we have multiple options to define key as identifier.

  1. Analyze requirement for compound keys

For the key definition of the CLICK object we need to know what the main purpose is of the objects and what the requirements are. We got the information from the application designer we need to answer to this question: How many times a certain user clicked to an element of the application. To fulfill this requirement we create key like this CLICK:<LOGIN_ID>:<ELEMENT_ID> e.g. CLICK:100AB:Button_Submit

  1. Define the physical implementation

To have real data in the database we have to define and choose physical implementation of our objects. For this we have to know the database implementation what kind of options we have. Redis is not a plain key-value store, but a data structures server, supporting different kind of values. Most obvious choices are String – store one value, Hash – store multiple fields. In our example we will use hash for USER and string for CLICK.

  1. Create mockup to test your database design

So let’s just create a working version and test it with the application. We will use HMSET for USER and SET for CLICK to create them.

  • HMSET USER:100AB Name “John Doe” Email “john@doe.com” LastLogin  “2015-05-05 05:05:05”
  • SET CLICK: 100AB:Button_Submit 1

We would like to support counting of click actions by database we just need to increment the number of clicks.

경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

  • INCR CLICK: 100AB:Button_Submit

NoSQL databases are not the best in aggregation however there are some very good features in Redis to overcome this shortcoming. We can support the application with top 10 lists by different dimension. For examples:

Top 10 clicked elements by user:

  • Store data: ZINCRBY CLICK: 100AB 1 Button_Submit
  • Query data: ZREVRANGE CLICK: 100AB 0 9 WITHSCORES

Top 10 active users by number of click:

  • Store data: ZINCRBY CLICK:ActiveUsers 1 USER:100AB
  • Query data: ZREVRANGE CLICK:ActiveUsers 0 9 ->
                       HGET <result> Name -> “John Doe”

Here you can see the beauty of our design. We can easily “join” click information to users in order to get the name of the active person.

Surely there are some other conceptual techniques and principles of NoSQL data modeling like denormalization, index tables, composite keys, etc. However the common things, I found most important, we need to analyze the requirements and transform them into a physical model corresponding to the database implementation.

 

[출처] https://www.linkedin.com/pulse/5-basic-steps-key-value-store-database-design-laszlo-wagner

 

 

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
공지 오라클 기본 샘플 데이터베이스 졸리운_곰 2014.01.02 86856
공지 [SQL컨셉] 서적 "SQL컨셉"의 샘플 데이타 베이스 SAMPLE DATABASE of ORACLE 가을의 곰을... 2013.02.10 79157
공지 [G_SQL] Sample Database 가을의 곰을... 2012.05.20 95893
78 [Kafka] Kafka 한번 살펴보자... Quickstart file 졸리운_곰 2021.06.18 1313
77 Java Kafka Producer, Consumer 예제 구현 Java를 이용하여 Kafka Producer와 Kakfa Consumer를 구현해보자. file 졸리운_곰 2021.06.18 1172
76 Beginner’s Guide to Understand Kafka file 졸리운_곰 2021.06.18 1575
75 [Kafka] Kafka 설치/실행 및 테스트 file 졸리운_곰 2021.06.18 1066
74 [java] [kafka] [Kafka] 개념 및 기본예제 file 졸리운_곰 2021.06.16 2170
73 Getting started with Apache Kafka in Python file 졸리운_곰 2020.09.10 2262
72 [Kafka] 다운로드 및 Quick Start file 졸리운_곰 2020.09.07 1899
71 [Kafka] 기본 개념잡기 file 졸리운_곰 2020.09.07 1713
70 Flume Integration with Kafka file 졸리운_곰 2019.04.16 2086
69 빅데이터: 플럼(Flume) 토폴로지 설계 file 졸리운_곰 2019.04.16 1604
68 실시간 처리를 위한 분산 메시징 시스템 카프카(Kafka) file 졸리운_곰 2018.05.12 1412
67 Flume과 Kafka를 사용한 초당 100만개 로그 수집 테스트 file 졸리운_곰 2018.05.12 1402
66 웹 크롤링 / web crwaling / web scraping / 웹 스크래핑 file 졸리운_곰 2017.07.09 1848
65 빅데이터 단지 몇퍼센트의 예측 정확성을 위하여 장애로 가득찬 빅데이터 시스템을 도입하여야 하는가에 대한 의문! file 졸리운_곰 2017.03.20 1611
64 빅데이터: 플럼(Flume) 토폴로지 설계 file 졸리운_곰 2017.03.20 1422
63 [실시간 분석 시스템] Apache Flume를 활용한 데이터 수집(1) file 졸리운_곰 2017.03.06 1322
62 [실시간 분석 시스템] 데이터 수집 #2 Apache Sqoop을 활용하여 RDBMS 데이터 수집(2) file 졸리운_곰 2017.03.06 1386
61 [실시간 분석 시스템] 데이터 수집 #2 Apache Sqoop을 활용하여 RDBMS 데이터 수집(1) file 졸리운_곰 2017.03.06 1109
60 [실시간 분석 시스템] 데이터 수집 #1 오픈 소스 수집기 비교 file 졸리운_곰 2017.03.06 1751
59 [실시간 분석 시스템] 일단 데이터 들여다 보기 file 졸리운_곰 2017.03.06 1947
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED