NLP 참고 인터넷 문서 정리 [자연어처리] [한글 자연언어처리]
2019.12.11 17:32
NLP 참고 인터넷 문서 정리 [자연어처리] [한글 자연언어처리]
NLP
? A curated list of speech and natural language processing resources
? An easy introduction to Natural Language Processing
? Introduction to Natural Language Processing for Text
? Introduction To Natural Language Processing | Machine Learning Projects | Eduonix
? A Review of the Neural History of Natural Language Processing
? Natural Language Processing (NLP) Tutorial | Data Science Tutorial | Simplilearn
? Over 200 of the Best Machine Learning, NLP, and Python Tutorials — 2018 Edition
? Natural Language Processing Tutorial Part 1 | NLP Training Videos | Text Analysis
? Natural Language Processing Tutorial Part 2 | NLP Training Videos | Text Analysis
? Natural Language Processing Tutorial Part 3 | NLP Training Videos | Text Analysis
? Natural Language Processing Tutorial Part 4 | NLP Training Videos | Text Analysis
? Natural Language Processing Tutorial Part 5 | NLP Training Videos | Text Analysis
? Natural Language Processing Tutorial Part 6 | NLP Training Videos | Text Analysis
? Natural_language_Processing_self_study
? Extracting meaningful text from webpages
? Extracting (meaningful) text from webpages - II
? Part 1: For Beginners - Bag of Words 캐글뽀개기 6월 이상열
? Writers Choose Their Favorite Words 쓰이는 단어의 종류를 통해 글 쓴 사람 예측?
? Algorithms for text fingerprinting?
? 하나의 차트로 이해하는 민주당과 공화당이 세계를 보는 다른 시각
? Ask HN: What are the best tools for analyzing large bodies of text?
? Special Section: Reconceiving Text Analytics
? ExoBrain
? 인간-기계 지식소통을 위한 자연어 QA 워크샵 – 엑소브레인 인공지능
? 한자로
? Making Apps Understand Natural Language
? Automatically spotting interesting sentences in parliamentary debates
? Bag of Words Meet Bags of Popcorn - (1) Part 1: Bag of Words
? WHERE TECHNOLOGY MEETS BUSINESS. TYING TEXT ANALYTICS TO YOUR BUSINESS GOALS
? For 40 years, computer scientists looked for a solution that doesn’t exist edit distance
? Deep Learning for NLP Best Practices
? DAWG data structure in Word Judge
? A Simple Artificial Intelligence Capable of Basic Reading Comprehension
? Extracting Structured Data From Recipes Using Conditional Random Fields
? How To Create Natural Language Semantic Search For Arbitrary Objects With Deep Learning
? politeness - Write in a more polite, friendly tone
? Understanding Natural Language with Deep Neural Networks Using Torch
? An Inside View of Language Technologies at Google
? Google Cloud에서 Natural Language API 정리
? Google Cloud 서비스 계정키 얻기 및 GCS 공유하기
? Understanding Convolutional Neural Networks for NLP
? 자연어 처리 문제를 해결하는 CONVOLUTIONAL NEURAL NETWORKS 이해하기
? Convolutional Methods for Text
? 텍스트 처리와 관련해서는 LSTM/GRU를 비롯한 RNNs 가 대세지만 CNN도 장점이 있는데 이를 잘 정리한 글
? RNN이 순서에 영향을 받지만 CNN은 단어의 의미에 영향을 주는 데에 있어 조금 멀리 떨어져 있는 문장에서의 단어 등이 역할을 할 수 있음
? 전체를 한꺼번에 보게 하는 데에는 더 유리
? NLP 전반에 대한 이해와 DNN 종류들의 장단점 등도 잘 파악할 수 있는 매우 좋은 글
? Convolutional Sequence-to-Sequence Learning (2017)
? Convolutional Sequence-to-Sequence Learning (2017)
? (NLP 처음 접하시는 분들을 위한)
1. RNN enc-dec 부터 conv seq2seq 까지 간단한 흐름 정리
1. conv s2s 이해를 위해 읽어야 할 논문 10+ 편
? Learning Deep Structured Semantic Models for Web Search using Clickthrough Data
? collocations.de - Association Measures
? Lecture 4: Evaluating language models
? An Experimental Study on Open Source Korean Morphological Analyzers for Evaluating Noun Extraction
? Episode 22: 자연언어처리 특집 1부 – 마이크로소프트 NLP연구실의 김용범님과 함께
? Espresso - AIR LAB, Changwon National University
? 악평생성기 (Bad Comment Generator using RNN) _ 송치성
? Bad Comment Generator using RNN
? Generating text using a Recurrent Neural Network
? 딥엘라스틱 - 검색 + 로봇 저널리즘 + 인지신경언어학 + 딥러닝NLP
? PHP + MySQL 언어 식별기(Language Detection) 개발기
? word-rnn - a fork of Andrej Karpathy's wonderful char-rnn
? Introducing DeepText: Facebook's text understanding engine
? 페이스북, ‘사람 수준으로’ 내용을 이해하는 딥텍스트 A.I. 공개
? 니코니코동화의 공개코멘트 데이터를 Deep Learning로 해석하기
? Introducing Cloud Natural Language API, Speech API open beta and our West Coast region expansion
? ko_restoration - Module for restoring Korean text working with KomornaPy
? Exploring Session Context using Distributed Representations of Queries and Reformulations
? 사용자의 쿼리 세션데이터와, 문서클릭데이터로 CNN으로 쿼리의 word-embedding을 만듦
? 쿼리와 관계를 벡터로 변환
? 두 쿼리의 관계벡터는 단순히 두 쿼리벡터의 뺴기(차이?)로 간단하지만
? 이러한 관계벡터들을 클러스터링하니, 쿼리 변환의 의도가 클러스터링 됨
? 동일의도인데, 다른 모양의 쿼리변환
? 검색 의도를 좁히는 쿼리변환
? 의도를 아예 점프하는 쿼리변환
? BabelNet
? An Intuitive Natural Language Understanding System
? Korean Treebank Annotations Version 2.0
? sample EUC-KR encoded
? An NLP Approach to Analyzing Twitter, Trump, and Profanity
? Deep Learning Cases: Text and Image Processing
? CS 124: From Languages to Information
? NLP Seminar Schedule — Winter 2019
? PyData Paris 2016 - Statistical Topic Extraction
? brat rapid annotation tool online environment for collaborative text annotation
? brat rapid annotation tool (brat) - for all your textual annotation needs
? 자료실
? 확률문법
? korean.abcthesaurus.com 동의어 사전
? Microsoft Concept Graph Preview For Short Text Understanding
? en.wikipedia.org/wiki/Precision_and_recall
? 실제와 예측이 일치; True Positive / Negative
? 실제와 예측이 불일치; False Positive / Negative
? 발생했다고 예측 Positive, 발생하지 않았다고 예측 Negative
? 정밀도와 재현율
? accuracy, precision, recall의 차이
? 정확도(accuracy)와 정밀도(precision)의 차이
? en.wikipedia.org/wiki/Sensitivity_and_specificity
? #2.6. Accuracy, Precision, Recall
? 입개발자를 위한 Accuracy, Precision, Recall
? Classification &Clustering 모델 평가
? Fighting Financial Fraud with Targeted Friction
? Beyond Accuracy: Precision and Recall
? Comparison of the best NSFW Image Moderation APIs 2018
? Understand Classification Performance Metrics
? 민감도와 특이도 (sensitivity and specificity)
? Natural Language Understanding with Distributed Representation
? Repository for PyCon 2016 workshop Natural Language Processing in 10 Lines of Code
? Deep Learning the Stock Market
? NLP: Everyday, Analytical &Unusual Uses
? Welcome to Railroad Diagram Generator! BNF rule to diagram
? Is Google Hyping it? Why Deep Learning cannot be Applied to Natural Languages Easily
? ratsgo.github.io/blog/categories
? Information Extraction with Reinforcement Learning
? Last Words: Computational Linguistics and Deep Learning
? PDP(연결주의)쪽 룸멜허트나 맥클랜드의 연구들 - 신경망 기반 의미론 모형
? 인간 언어와 관련한 인지과학적 연구 - 어떻게 언어를 학습하고 개념들이 조직화되는가라는 관점
? Computational Linguistics and Deep Learning
? 4 APPROACHES TO NATURAL LANGUAGE PROCESSING &UNDERSTANDING
? Distributional: 최근 유행하는 ML. 폭은 넓힐 수 있지만, 깊이는 잡지 못함
? Frame-based: 마빈 민스키. 논리적 semantics에 강점. 확고한 supervision이 존재해야 한다는 단점
? Model-theoretical: Q/A와 rich semantics의 장점. (프레임 기반보다 더한) labor-intensive and narrow in scope
? Interactive learning: language as a cooperative game between speaker and listener
? Syntax – what is grammatical? : “no compiler errors”
? Semantics – what is the meaning?: “no implementation bugs”
? Pragmatics – what is the purpose or goal?: “implemented the right algorithm.”
? Deep Learning for Text Understanding from Scratch
? Mining English and Korean text with Python
? nlp4kor
? Building A Gigaword Corpus Lessons on Data Ingestion, Management, and Processing for NLP
? Teaching Machines to Describe Images with Natural Language Feedback
? Sang-Kil Park's Jupyter Notebooks
? An Adversarial Review of “Adversarial Generation of Natural Language”
? Deep Learning for Speech and Language
? deep learning nlp best practices
? Speech and Language Processing (3rd ed. draft)
? Memory Augmented Neural Networks for Natural Language Processing
? Natural Language Processing Tasks and Selected References
? 자연언어처리(NLP)를 위한 언어학 기초
? 담화분석
? 화용론
? 의미론
? 통사론
? 구와 문장
? 형태론
? 단어의 형성
? 언어의 기원
? Deep Learning for NLP, advancements and trends in 2017
? AI: NLP
? Experiments Codes for Bi-directional Block Self-attention
? Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling
? 주어진 시퀀스를 여러 개의 Block 으로 나누고 intra-block SAN으로 local context 를 모델링한 뒤, inter-block SAN으로 long-range dependency 를 모델링
? 기존의 Self-Attention Network (SAN) 이 너무 메모리를 많이 쓰는 점을 개선
? 많은 NLP 분야에서 Self-attention 기법들이 (특히 번역 분야에서는) 표준으로 자리잡고 후속 연구가 활발히 이루어지고 있는 걸로 보임
? (ex. Non-autoregressive transformer, Masked self-attention, Directional self-attention)
? Understanding and Applying Self-Attention for NLP - Ivan Bilan
? How to solve 90% of NLP problems: a step-by-step guide
? 파이썬자연어처리
? Text Analysis Developers’ Workshop 2018 참석 후기
? Text Analysis in Excel: Real world use-cases
? Auto Tagging Stack Overflow Questions
? A Neural Network Model That Can Reason - Prof. Christopher Manning
? Compositional Attention Networks for Machine Reasoning
? CodeSandbox - an online editor that helps you create web applications, from prototype to deployment
? Team AURA - 1st Meeting Summary
? NLP Tutorial with Deep Learning using tensorflow
? NLP Tutorial with Deep Learning using tensorflow
? Natural Language Processing Tutorial for Deep Learning Researchers TensorFlow and Pytorc
? NLP's ImageNet moment has arrived
? Feature-wise transformations - A simple and surprisingly effective family of conditioning mechanisms
? PyConKr 2018 Why I learn, How I learn
? Analogy and Analogical Reasoning
? handwritten Hangul Datasets: PE92, SERI95, and HanDB
? How NLP is Automating the complete Text Analysis Process for Enterprises?
? 강화학습을 자연어 처리에 이용할 수 있을까? (보상의 희소성 문제와 그 방안)
? NLP's ImageNet moment has arrived
? 시간 문제에 불과하다는 결론, BERT의 등장으로 현실에 가까워짐(ELMO - LSTM / OpenAI의 GPT, BERT - Transformer)
? Pre-trained Models의 fine-tuning은 필수, 인간이 언어를 이해한다는 것이 그저 엄청난 계산에 불과할 뿐이라는 사실(정말인가?)
? 이제 계산량을 줄이는 방법이 아니라 계산량을 늘리고 계산 속도를 높이는 방향이 옳을 지도 모름
? DLK2NLP: Day-by-day Line-by-line Keras-based Korean NLP
? 3i4K - Intonation-aided intention identification for Korean
? KorEmo - 5-class Korean emotion classifier 감정분류
? raws - Real-time Automatic Word Segmentation (for user-generated texts) 한영 noisy text segmentation
? NLP Guide: Identifying Part of Speech Tags using Conditional Random Fields
? Industrial strength Natural Language Processing
? A Review of the Neural History of Natural Language Processing
? Analyzing open-ended text? Its easier than you think!
? Fast Word Segmentation of Noisy Text
? Solving NLP task using Sequence2Sequence model: from Zero to Hero
? Natural Language Processing is Fun! How computers understand Human Language
? Natural Language Processing in Python
? The 7 NLP Techniques That Will Change How You Communicate in the Future
? (Part I)
? Natural Language Understanding benchmark
? NLU / Intent Detection Benchmark by Intento, August 2017
? 콜라 좀… 쉽게 담을 수 없나요, 쓰앵님 메뉴 검색을 위해 초중종성 분리 검색 개발
? Machine Learning with Python: NLP and Text Recognition
? Concrete solutions to real problems
? OpenAI GPT-2: Understanding Language Generation through Visualization
? Better Language Models and Their Implications GPT-2 based artificial news
? The Illustrated GPT-2 (Visualizing Transformer Language Models)
? OpenGPT-2: We Replicated GPT-2 Because You Can Too
? The Illustrated GPT-2 (Visualizing Transformer Language Models)
? Fine-Tuning GPT-2 from Human Preferences
? Text generation with a Variational Autoencoder
? Sentence Simplification with Seq2Seq
? Integrating Transformer and Paraphrase Rules for Sentence Simplification
? 10 Exciting Ideas of 2018 in NLP
? #자연어, #시퀀스를 위한 #재귀신경망 성능향상 기법! 대공개!! 첫번째!
? Justin J. Nguyen: Exposing Dark Data in the enterprise with custom NLP | PyData Miami 2019
? Natural language processing of customer reviews
? Introduction to Natural Language Processing (NLP) and Bias in AI
? nlp_applications ipynb
? NLP 101: 딥러닝과 자연어 처리 학습을 위한 자료 저장소
? Natural Language Processing RoadMap - 2019
? NLP HighlightsPro - Allen Institute for Artificial Inte Seattle, United States
? SKC_Text_Preprocessing - SKC 텍스트 전처리 강의
? 딥 러닝 자연어 처리를 학습을 위한 파워포인트. (Deep Learning for Natural Language Processing)
? Distilling knowledge from Neural Networks to build smaller and faster models
띄어쓰기
? Sentence boundary disambiguation
? python-crfsuite를 사용해서 한국어 자동 띄어쓰기를 학습해보자
? 딥러닝 한글 자동띄어쓰기 모형 성능 향상 및 API 업데이트
? KoSpacing : 한글 자동 띄어쓰기 패키지 공개
? KoSpacing - R package for automatic Korean word spacing
BERT
? Open Sourcing BERT: State-of-the-Art Pre-training for Natural Language Processing
? BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
? BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
? BERT TensorFlow code and pre-trained models for BERT
? BERT – STATE OF THE ART LANGUAGE MODEL FOR NLP
? Language Learning with BERT - TensorFlow and Deep Learning Singapore
? BERT-NER - Use google BERT to do CoNLL-2003 NER !
? Bert state Of The Art pre Training for nlp Post
? bert-multiple-gpu - A multiple GPU support version of BERT
? NVIDIA Achieves 4X Speedup on BERT Neural Network
? The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning)
? Multi-label Text Classification using BERT – The Mighty Transformer
? Visualization tool for Transformer-based language representation models (demonstrated on BERT)
? Transformer-Encoder-with-Char
? Language Model Overview: From word2vec to BERT
? BERT Explained: State of the art language model for NLP
? Efficient Training of Bert by Progressively Stacking
? Source code for "Efficient Training of BERT by Progressively Stacking"
? 카톡 데이터는 어떻게 정제할 수 있을까? - Dialog-BERT 만들기 1편
? 누가누가 잘하나! 대화체와 합이 잘 맞는 Tokenizer를 찾아보자! - Dialog-BERT 만들기 2편
? 카톡 대화 데이터를 BERT로 잘 학습시킬 수 있을까? - Dialog-BERT 만들기 3편
? A Simple Guide On Using BERT for Binary Text Classification
? Fast implementation of BERT inference directly on NVIDIA (CUDA, CUBLAS) and Intel MKL
? MULTI GPU환경에서 ETRI 한국어 BERT모델 활용한 Korquad 학습 방법
? nlp-api - ETRI KoBERT에서 사용하기 위해 만든 Mecab 형태소 분석기 API
? AI도 한글 공부가 필요해! 국내 유일의 한국어 데이터셋 코쿼드(KorQuAD) 2.0 이야기
? Korean BERT pre-trained cased (KoBERT)
? Google Brain’s XLNet bests BERT at 20 NLP tasks
? XLNet: Generalized Autoregressive Pretraining for Language Understanding(19.06.25)
? 실제 코드로 보는 XLNet (Code Review)
? A Simple Explanation of XLNet
? 파이콘 2019 100억건의 카카오톡 데이터로 똑똑한 일상대화 인공지능 만들기
? Smaller, faster, cheaper, lighter: Introducing DistilBERT, a distilled version of BERT
? exBERT - A Visual Analysis Tool to Explore Learned Representations in Transformers Models
? More on Transformers: BERT와 친구들
? KoreanCharacterBert - Korean BERT model using character tokenizer
Book
? Neural Network Methods for Natural Language Processing
? A Primer on Neural Network Models for Natural Language Processing
? Quantitative corpus linguistics with R: a practical introduction
? Speech and Language Processing (3rd ed. draft)
Category
? text categorization; 예를 들어 100만개의 상품 description이 있고, 이걸 supervised를 위한 document로 사용해, 나중에 들어오는 상품 description을 통해 cateogory 판별
? naive bayes
? gensim, model.docvecs e.g. model.docvecs.most_similar([1,2,3]) -> 문서 태그가 '10000'이면 model.docvecs['10000']으로 해당 docvec을 가져옴
? most_similar 호출 시 파라미터로써 벡터(numpy array)의 리스트 혹은, 문서의 태그들이 담긴 리스트 전달 가능
? 결과 값으로 문서의 태그 및 유사도를 반환
? doc2vec
? 낮은 정확도
? 기본적으로 word co-occurrence 에 기반하고 있고 각 word 는 word embedding 에 의한 vector 사용
? 이 vector들의 단순 합은 ambiquity 문제가 경험적으로 발생
? document 단위가 짧으면 짧은 대로 , 쿼리 스트링이 짧으면 짧은 대로 또 ambiguity 문제가 발생
? word2vec
? doc2vec과 유사
? 전체 corpus 에 대해 모델을 만든 후, predict 할 때 description 보다 제목 같이 짧으면서 컨텍스트를 담고 있는 것으로 입력을 주면 좀 나음
? 이미 카테고리 도메인이 결정된 경우 LDA/LSI 가 더 좋은 방법일 수 있음
? LDA / LSI 는 각각의 카테고리를 반영하는 토큰의 기여도를(weight) 확률분포로 표현
? LDA 경우 더 많이 기여하고 있는 워드 순
? LSI 의 경우 positive 기여도 뿐만 아니라 negative 기여도 확률을 결과로 반환
? 그러므로 쿼리스트링이 있을 때 가장 확률 높은 카테고리 계산 가능
? TFIDF
? feature 수가 많다 해도 document similarity 를 계산하는 게 아니라 카테고리를 분류하기 위함이기 때문에 dimension 문제가 크지 않을 수 있음
? TFIDF 로 weighting 한 벡터들을 가지고 클러스터링
? 실제 label 가지고 TFIDF weight 가 각 label 을 얼마나 잘 구분하고 있는지 feasibility 를 판단할 수도 있음
? 혹은 각 카테고리별로 모델을 만들어서 dictionary 를 작게 만들어 feature 수를 줄일 수도 있음
? 각각의 dictionary 셋과 워드에 대한 TFIDF weight 를 가지고 카테고리별로 representing 한 워드들을 뽑아볼 수도 있음
? Category Theory for Programmers: The Preface
? Category Theory for Scientists (Old Version)
? Logic, Languages, Compilation, and Verification
? ‘뉴욕타임스’, 머신러닝 기반 자동 태그 시스템 개발
? Fast &easy baseline text categorization with vw
? 글쓰기 화면에서 카테고리 자동 추천하는 모델 만들기 fasttext
ChatBot
? DEEP LEARNING FOR CHATBOTS, PART 1 – INTRODUCTION
? 딥러닝 챗봇, PART 1 – INTRODUCTION (한글번역)
? DEEP LEARNING FOR CHATBOTS, PART 2 – IMPLEMENTING A RETRIEVAL-BASED MODEL IN TENSORFLOW
? 딥러닝 챗봇 , PART 2 – IMPLEMENTING A RETRIEVAL-BASED MODEL IN TENSORFLOW(한글번역)
? Deep leaning for Chatbot Developers
? Clippy’s Back: The Future of Microsoft Is Chatbots
? Build a bot without coding - Launch a full-featured chatbot in 7 minutes
? Microsoft Bot Framework 관련 강좌
? Botkit - Building Blocks for Building Slack Bots
? AWS Lambda와 API Gateway로 Slack Bot 만들기
? Your next shopping experience starts with a text
? x.ai is a personal assistant who schedules meetings for you
? Kino - My Personal Assistant (개인용 Slack Bot을 통한 Quantified Self 프로젝트)
? wit.ai
? Wit.ai stories/conversational app demo
? The White House's New Facebook Messenger Bot Makes It Easy To Send A Message To Obama
? Wonder is a bot that will remember anything for you
? Introducing the Bots Landscape: 170+ companies, 63220 billion in funding, thousands of bots
? 지적 대화를 위한 깊고 넓은 딥러닝 Pycon APAC 2016
1. 이미지(사람의 얼굴 사진)을 이해하고 스스로 만드는 모델
? github.com/carpedm20/DCGAN-tensorflow
? Question Answering, Language Model
? Teaching Machines to Read and Comprehend
? Neural Variational Inference for Text Processing
? Advanced Natural Language Processing Tools for Bot Makers – LUIS, Wit.ai, Api.ai and others
? The Rise of Chat Bots: Useful Links, Articles, Libraries and Platforms
? Know Your Bot, Part II: Slack, The Bot Paradise
? Know Your Bot, Part I: Telegram And Twitter
? s2 lab1-1: API.ai concept and terms
? s2 lab1-2: API.ai making bot demo
? Multi-domain Neural Network Language Generation for Spoken Dialogue Systems(NAACL-HLT 2016)
? code
? Dialog System - http://nlp.postech.ac.kr/research/dialog_system
? Do-it-yourself NLP for bot developers
? Making Friends With Artificial Intelligence: Eric Horvitz at TEDxAustin
? Facebook steps in to prove the value of chatbots with Tommy Hilfiger
? The rise of bots... acquisitions!
? 라이크 어 Poncho: JiveScript 날씨 챗봇
? Heek is a chatbot that can build you a website
? 챗봇 개발 프레임워크 ChatFlow, 베타버전 출시
? Build a restaurant reservation Messenger bot using IBM Watson with no code
? Deep Q&A
? 챗봇 시작해보기
? A developer's guide to chatbots
? TF-KR Conf 2 강의 2: 조재민, Developing Korean chatbot 101
? Developing Korean Chatbot 101
? Retrieval-Based Conversational Model in Tensorflow (Ubuntu Dialog Corpus)
? 20170121 한국인공지능협회 - 제7차 오픈세미나 - 챗봇 (1/5)
? 20170121 한국인공지능협회 - 제7차 오픈세미나 - 챗봇 (2/5)
? 20170121 한국인공지능협회 - 제7차 오픈세미나 - 챗봇 (3/5)
? 20170121 한국인공지능협회 - 제7차 오픈세미나 - 챗봇 (4/5)
? 20170121 한국인공지능협회 - 제7차 오픈세미나 - 챗봇 (5/5)
? KahWee Teng: Coding Chat Bots - JSConf.Asia 2016
? 카카오톡 자동응답 API를 이용하여 카카오톡 봇 만들기
? The Conversational Intelligence Challenge
? Visual Dialog Challenge 2018
? Natural Language Pipeline for Chatbots
? Contextual Chat-bots with Tensorflow
? How To Build an Interactive Chatbot for Twitter Direct Messages
? 왓슨으로 쉽게 개발하는 카카오톡 챗봇 1. Watson Conversation 서비스로 인공지능 대화 서비스 만들기
? Node.js Facebook 챗봇 빠른시작: 369봇 만들기
? Stephanie - YOUR VIRTUAL ASSISTANT!
? Sounder API - the Sounder Library API, which is an abstraction of the Sounder Algorithm
? USAGE
? Deal or no deal? Training AI bots to negotiate
? Deal or No Deal? End-to-End Learning for Negotiation Dialogues pytorch
? Creating AnswerBot with Keras and TensorFlow (TensorBeat)
? Own ChatBot Based on Recurrent Neural Network
? Chatbots: Theory and Practice
? Show me red! – feat. 서울 시립 미술관 데이터를 사용한 챗봇 만들기
? Python과 Tensorflow를 활용한 AI Chatbot 개발 및 실무 적용
? Neural Network Dialog System Papers
? 강화학습 챗봇 Dialogue Policy Optimization
? Python과 Tensorflow를 활용한 Al 챗봇 개발 4강
? Python과 Tensorflow를 활용한 AI Chatbot 개발 및 실무 적용
? Seq2Seq를 활용한 간단한 Q/A 봇을 만들어보자
? Retrieval-Based Conversational Model in Tensorflow (Ubuntu Dialog Corpus)
? Automated Text Classification Using Machine Learning
? sunwoobot - 선우봇 카카오 i 오픈빌더 챗봇
? A Repository of Conversational Datasets PolyAI 공개. Reddit, OpenSubtitles, AmazonQA 등에서 모은 수억 건의 대화 데이터셋
? Chatbot
? Plato Research Dialogue System
? Introducing the Plato Research Dialogue System: A Flexible Conversational AI Platform
? Not another Conversational AI report
? How do Dialogue Systems decide what to say or which actions to take?
Python
? Building AI Chat bot using Python 3 &TensorFlow
? Chat bot making process using Python 3 &TensorFlow
? 신정규 : Creating AI chat bot with Python 3 and Tensorflow - PyCon APAC 2016
? Create a Chatbot for Telegram in Python to Summarize Text
? python으로 telegram bot 활용하기
? 1 기본 설정편
? 2 채널편
? 3 챗봇편
? Learn to build your first bot in Telegram with Python
? Building a Telegram Bot �� to Automate Web Processes Using Python, Selenium and Telegram
? [카카오톡 대화 생성기(http://jsideas.net/python/2017/04/05/kakao_rnn.html)
? ChatOps with PowerShell - Matthew Hodgkins
? Building a Simple Chatbot from Scratch in Python (using NLTK)
? A Transformer Chatbot Tutorial with TensorFlow 2.0
? RASA - Create assistants that go beyond basic FAQs
? Building a chatbot with Rasa
? Building a Conversational Chatbot for Slack using Rasa and Python -Part 1
? How to build a voice assistant with open source Rasa and Mozilla tools
Classification
? Implementing a CNN for Text Classification in TensorFlow
? Convolutional Neural Network for Text Classification in Tensorflow
? IMPLEMENTING A CNN FOR TEXT CLASSIFICATION IN TENSORFLOW (한글 번역)
? CNNs for sentence classification
? 합성곱 신경망(CNN) 딥러닝을 이용한 한국어 문장 분류
? MIT 6.S191 Lecture 2: Sequence Modeling with Neural Networks
? Free Code Friday - Better and Faster Machine Learning Classifiers in Python
? “What is Relevant in a Text Document?”
? 예를 들어, 카테고리가 있는 뉴스문서 학습데이터가 있는 경우 문서를 분류하는 분류기를 만들 때
? 문서에서 어떤 단어가 어떤 클래스로 분류하는데 얼만큼의 영향이 있었는지 역으로 추적하기가 쉽지 않음(Maximum Entropy 같은 걸 사용하는 것이 아니라면)
? 이를 역으로 추적하는 방법에 대한 논문
? Text Classification using Neural Networks
? Text Classification using Algorithms
? Text Classifier Algorithms in Machine Learning
? Tensorflow Text Classification – Python Deep Learning
? lime
? On Building a “Fake News” Classification Model *update
? Automated Text Classification Using Machine Learning
? TRAIN ONCE, TEST ANYWHERE: ZERO-SHOT LEARNING FOR TEXT CLASSIFICATION
? Zero Shot Learning : 학습 데이터없이 텍스트 분류 모델 만들기
? Zero Shot Learning은 학습을 하지 않고 데이터세트의 구성원을 추론할 수 있는 방법
? 대부분 하나의 데이터 세트에서 습득한 지식을 다른 학습 세트에 적용 할 수 있는 일부 형태의 transfer learning에 의해 성취됩니다
? 지금까지 imagenet 데이터세트의 지식을 새로운 것에 사용할 수 있는 비전 작업을 위해 여러 개의 Zero Shot Learning 방법을 제안했지만 텍스트 분류를 위한 건 최초
? 큰 노이즈의 데이터세트에서 문장과 해당 범주 간의 관계를 학습하여 새로운 범주 또는 새 데이터세트로 일반화
? TRY OUR CUSTOM CLASSIFIER DEMO
? Alisa Dammer - Baby steps in short-text classification with python
? Actionable and Political Text Classification Using Word Embeddings and LSTM
? Pycon Ireland 2017: Text Classification with Word Vectors &Recurrent Neural Networks - Shane Lynn
? Machine Learning - Text Classification with Python, nltk, Scikit &Pandas
? Introduction to Natural Language Processing with Python - Asyncjs
? Patrick Harrison | Modern NLP in Python
? Advanced Python 2: Advanced Text Processing
? Creating a simple text classifier using Google CoLaboratory Google CoLaboratory 환경에서 Scikit Learn를 사용하여 간단한 2진 텍스트 분류자를 만드는 방법
? Text Classification with TensorFlow Estimators
? Multi-Class Text Classification with Scikit-Learn
? Multi Label Text Classification with Scikit-Learn
? Recurrent Neural Network for Text Calssification
? Introducing state of the art text classification with universal language models
? Evaluating Classifiers: Confusion Matrix for Multiple Classes
? The last 3 years in Text Classification
? Automated Text Classification Using Machine Learning
? Introducing Custom Classifier – Build Your Own Text Classification Model Without Any Training Data
? Practical Text Classification With Python and Keras
? Multi-Class Text Classification with SKlearn and NLTK in python| A Software Engineering Use Case
? Tutorial on Text Classification (NLP) using ULMFiT and fastai Library in Python
? Democratizing NLP content modeling with transfer learning using GPUs - Sanghamitra Deb
? The State of Transfer Learning in NLP
? NLP Classification Tutorial with PyTorch CBOW, CNN, DCNN, RNN, LSTM
? Practical Text Classification With Python and Keras
? Multi-Class Text Classification Using PySpark, MLlib &Doc2Vec
? Using Doc2Vec to classify movie reviews
? A Basic NLP Tutorial for News Multiclass Categorization
? Natural Language Processing, Support Vector Machine, TF- IDF, deep learning, Spacy, Attention LSTM
? 헤드 라인과 간단한 설명을 기반으로 뉴스 유형을 식별하여 Python에서 텍스트 데이터의 멀티 클래스 분류 방법을 이해
Clustering
? dbscan
? Finding Topics in Harry Potter using K-Means Clustering
? 언론사가 알아야 할 알고리즘
? Comparing different clustering algorithms on toy datasets
? Text Clustering : Get quick insights from Unstructured Data 1
? Text Clustering : Get quick insights from Unstructured Data 2
? 14 Great Articles and Tutorials on Clustering
? The 5 Clustering Algorithms Data Scientists Need to Know
? Understanding Hate Speech on Reddit through Text Clustering
Conference
? JSALT 2019 Montréal: Dive into Deep Learning for Natural Language Processing
? LangCon
? 텐서플로 월드2019 행사 핵심요약 2. NLP가 대세입니다!
Corpus
? CORPORA AND OTHER LANGUAGE AND SPEECH DATA UNDER DICE
? UTagger + KorpuSQL을 이용해서 코퍼스 구축하기
? KorpuSQL 클릭만으로 간편하게 코퍼스 구축하기
? 인공지능 씨앗 한글 말뭉치, 2007년 멈춰선 까닭
? ④ 송철의 국립국어원장 "한국어 AI 시대의 기초는 말뭉치..제2의 세종계획 추진해야"
? 언제까지 포털 영어사전만 쓸 건가요? – 말뭉치(코퍼스)를 활용한 영어 글쓰기 기초 편
? NIA(National Information Society Agency) Dictionary
? 신조어 포함된 형태소사전 공개..빅데이터 분석 정확도↑
? 국어사전 데이터
? Korean Parallel corpora (of https://sites.google.com/site/koreanparalleldata/)
? Facebook, NYU expand available languages for natural language understanding systems
? koSentences - a large-scale web corpus of Korean text
Disambiguation
? Automatic disambiguation of English puns
? Discovering Types for Entity Disambiguation
Doc2Vec
? REDDIT 2 VEC - Use Doc2Vec to get SubReddit Suggestions
Filtering
Knowledge
? :BaseKB Gold Ultimate is now available in AWS
? Knowledge-Based Trust: Estimating the Trustworthiness of Web Sources
Language Model LM
? 언어 모델링은 음성-텍스트, 대화식 시스템, 텍스트 요약과 같은 여러 가지 자연어 처리 작업에 핵심적인 문제
? Text Generation
? 텍스트 생성은 언어 모델링 문제의 유형
? 잘 학습된 언어 모델은 텍스트에서 사용된 단어의 이전 순서를 기반으로 단어의 발생 가능성을 학습
? 언어 모델은 문자 수준, n-gram 수준, 문장 수준 또는 단락 수준에서 조작 가능
? WHAT EVERY NLP ENGINEER NEEDS TO KNOW ABOUT PRE-TRAINED LANGUAGE MODELS
? Language modeling a billion words
? Perplexed by Game of Thrones. A Song of N-Grams and Language Models
? Character-Aware Neural Language Models
? Character-Aware Neural Language Models
? CNN과 Highway Network를 사용 (입력은 LSTM)해서 State-of-Art의 성과
? 기존보다 크게 감소된 Parameter로 높은 성능을 내어, 휴대폰과 같은 Model Size가 중요한 영향을 미치는 곳에 적합
? Word Embedding 시 형태소 tagging 필요하지 않음
? 형태소 정보들이 많은 언어에서 기존보다 높은 성능 (언어 종속성 낮음)
? How to Develop a Word Embedding Model for Predicting Movie Review Sentiment keras, word2vec
? MUSE: Multilingual Unsupervised and Supervised Embeddings
? LSTM and QRNN Language Model Toolkit
? Generating Drake Rap Lyrics using Language Models and LSTMs
? Recurrent Neural Networks: The Powerhouse of Language Modeling
? Character-Aware Neural Language Models
? LAMA: LAnguage Model Analysis
LDA Latent Dirichlet Allocation
? Latent Dirichlet Allocation (LDA) with Python
? Latent Dirichlet Allocation, LDA
? word2vec, LDA, and introducing a new hybrid algorithm: lda2vec
? LDA in Python – How to grid search best topic models?
? Scikit Learn은 Latent Dirichlet allocation(LDA), LSI, Non-Negative Matrix Factorization과 같은 알고리즘을 사용하여 주제 모델링을 위한 편리한 인터페이스를 제공
? 이 튜토리얼에서는 최상의 LDA 토픽 모델을 작성하고 결과를 의미있는 결과로 보여주는 방법
? Language Modelling and Text Generation using LSTMs — Deep Learning for NLP
? 최첨단의 RNN을 구현하고 학습하여 자연어 텍스트를 생성하는 언어 모델을 만드는 방법을 설명
? 이 모델의 목적은 일부 입력 텍스트가 있는 경우 새 텍스트를 생성
? Topic Modeling and Latent Dirichlet Allocation (LDA) in Python
? The Hottest Topics In Machine Learning - Analyzing machine learning trends in research
Library
? 날개셋
? 오픈 한글
? 은전한닢 프로젝트 - 검색에서 쓸만한 오픈소스 한국어 형태소 분석기를 만들자!
? elasticsearch-analysis-seunjeon 5.0.0.0 배포합니다
? AllenNLP - An open-source NLP research library, built on PyTorch
? An open-source NLP research library, built on PyTorch
? crf
? Autosub - Command-line utility for auto-generating subtitles for any video file
? CLaF: Clova Language Framework https://naver.github.io/claf
? decaNLP - The Natural Language Decathlon: A Multitask Challenge for NLP
? The Natural Language Decathlon
? fastText is a library for efficient learning of word representations and sentence classification
? C++, 추가적인 의존 라이브러리 없음
? Deep Learning 기반의 분류기와 정확도는 비슷하면서도 속도가 빠름
? multi-core CPU 상에서 10억개 이상의 단어를 10분 내로 학습하고, 50만개의 문장을 1분안에 312k개의 클래스로 분류 가능
? Bag of Tricks for Efficient Text Classification
? our fast text classifier fastText is often on par with deep learning classifiers in terms of accuracy, and many orders of magnitude faster for training and evaluation
? We can train fastText on more than one billion words in less than ten minutes using a standard multicore CPU, and classify half a million sentences among 312K classes in less than a minute.
? Enriching Word Vectors with Subword Information
? Facebook’s Artificial Intelligence Research lab releases open source fastText on GitHub
? Introduction to Natural Language Processing with fastText
? FastText.zip: Compressing text classification models
? Aligning the fastText vectors of 78 languages
? Introduction to Natural Language Processing with fastText
? fastText4j - Java port of C++ version of Facebook Research fastText
? scikit-learn wrappers for Python fastText
? FastText Tutorial - How to Classify Text with FastText
? models.fasttext – FastText model gensim example
? 글쓰기 화면에서 카테고리 자동 추천하는 모델 만들기
? Attention API로 간단히 어텐션 사용하기 gluonNLP
? go-freeling - Golang Natural Language Processing
? hangul-toolkit - 한글 자모 분해, 조합(오토마타), 조사 붙이기, 초/중/종 분해조합, 한글/한자/영문 여부 체크 등을 지원
? InferSent - semantic sentence 표현을 제공하는 sentence embedding 방법
? Kanji recognition - implementation of Nei Kato's directional feature extraction algorithm
? khaiii
? kakao의 오픈소스 Ep9 - Khaiii : 카카오의 딥러닝 기반 형태소 분석기
? 카카오 형태소 분석기(khaiii) 설치와 은전한닢(mecab) 형태소 분석기 비교
? 카카오 형태소 분석기(khaiii) 분석 시간 및 딥러닝 모델 성능 비교
? Kiwi - 지능형 한국어 형태소 분석기(Korean Intelligent Word Identifier)
? 지능형 한국어 형태소 분석기 ver 0.3 - 알고리즘 최적화 &메모리 풀
? knwl - A Javascript Natural Language Parser
? KoalaNLP = Korean + Scala + NLP. 한국어 형태소 및 구문 분석기의 모음입니다
? KoParadigm: Korean Inflectional Paradigm Generator
? paradigm은 용언 활용 테이블을 뜻하는 언어학 용어. 예를 들어, 영어의 go는 go, went, going, goes 등과 같이 어형이 변화
? 한국어는 그 변화양상이 복잡. 동사/어미의 종류와 소리에 따라 규칙이 복잡. 그 규칙들을 테이블로 정리해 공개
? KorpuSQL
? Koshort - Koshort은 한국어 NLP를 위한 high-level API 프로젝트입니다
? LASER - Zero-shot transfer across 93 languages: Open-sourcing enhanced LASER library
? Mecab
? mit-nlp
? NGT - Neighborhood Graph and Tree for Indexing High-dimensional Data
? word embeddings와 같은 고차원 데이터에서 k nearest item을 근사적으로 빠르게 찾는 라이브러리
? annoy와 비슷하지만 graph tree 기반 indexing
? parserator - a framework for making parsers using natural language processing (NLP) methods
? PyText - a deep-learning based NLP modeling framework built on PyTorch
? PyText - A natural language modeling framework based on PyTorch https://fb.me/pytextdocs
? Open-sourcing PyText for faster NLP development
? Introducing PyText - Facebook’s New Framework for Better NLP Development
? Rouzeta - 유한 상태 기반의 한국어 형태소 분석기
? Simplenlg - a simple Java API designed to facilitate the generation of Natural Language
? Stanford Natural Language Processing Group
? corenlp
? Stanford CoreNLP – a suite of core NLP tools
? StanfordNLP: A Python NLP Library for Many Human Languages
? Stanford CS224U: Natural Language Understanding | Spring 2019
? StarSpace - Learning embeddings for classification, retrieval and ranking
? tacit - Text Analysis,Collection and Interpretation Tool
? Text Understanding from Scratch
Library Java
? Autocomplete words with spring boot and redis 자동완성
? KLAY - Korean Language AnalYzer (한국어 형태소 분석기)
? lucene-Korean-Analyzer Lucene Analyzer For Korean
? 03. Solr 5.0.0 - 아리랑(arirang) 한글 형태소 분석기 적용
Library JavaScript
? TajaJS is a simple Hangul library in JavaScript
Library Python
? 13 Deep Learning Frameworks for Natural Language Processing in Python
? Approximate nearest neighbor methods and vector models – NYC ML meetup
? Approximate Nearest Neighbors
? Document Clustering with Python
? Mining English and Korean text with Python
? A Python script to check if a character is or a text contains emoji
? flair - A very simple framework for state-of-the-art Natural Language Processing (NLP)
? Tadej Magajna - State of the Art NLP with Flair
? Keyword finder: automatic keyword extraction from text
? KoNLPy: Korean NLP in Python
? 자바, 미안하다! Korean NLP with Python
? MAC OSX에서 konlpy 설치 시 ImportError: No module named 'jpype' 오류 해결
? korean - A library for Korean morphology
? gist.github.com/allieus/0e8b609fe146ad63462ca81c70b2f5a2
? Koshort - a Python project for Korean natural language processing... or maybe Korean domestic cat
? Goorm - A little word cloud generator in Python - Korean wrapper
? krtpy - Korean Romanization/Hangulization utility written in python
? kss - Korean Sentence Splitter
? NLTK
? book
? github.com/zerosum99/python_nltk
? Tutorial 5: Analyzing text using Python NLTK
? NLTK with Python 3 for Natural Language Processing
? NLTK Text Processing Tutorial Series
? Computing Document Similarity with NLTK (March 2014)
? Tokenizing Words Sentences with Python NLTK
? Natural Language Processing (NLP) Tutorial with Python &NLTK
? ParlAI (pronounced “par-lay”) - a framework for dialog AI research, implemented in Python
? ParlAI: A new software platform for dialog research
? pyeunjeon (python + eunjeon) 은전한닢 프로젝트와 mecab 기반의 한국어 형태소 분석기의 독립형 python 인터페이스
? py-hanspell - 파이썬 한글 맞춤법 검사 라이브러리. (네이버 맞춤법 검사기 사용)
? PyStruct - Structured Learning in Python
? soynlp 단어 추출/ 토크나이저 / 품사판별/ 전처리 기능을 제공
? spaCy - a library for industrial-strength natural language processing in Python and Cython
? NLP (SpaCy) 총 4개의 챕터, SpaCy 패키지 사용 방법
? spaCy Cheat Sheet: Advanced NLP in Python
? dependency parse tree visualization
? spaCy: Industrial-strength NLP
? dependency parse tree visualization
? Neural coref - State-of-the-art coreference resolution based on neural nets and spaCy
? 신경망과 spaCy를 이용한 coreference resolution library
? State-of-the-art neural coreference resolution for chatbots
? Prodigy: A new tool for radically efficient machine teaching
? yujuwon.tistory.com/m/tag/spaCy
? Machine Learning for Text Classification Using SpaCy in Python
? Natural Language in Python using spaCy: An Introduction
? TextBlob Sentiment: Calculating Polarity and Subjectivity python
? Natural Language Basics with TextBlob
? Text Generation With LSTM Recurrent Neural Networks in Python with Keras
? UTagger
? vwnlp - Solving NLP problems with Vowpal Wabbit: Tutorial and more
Library R
? KoNLP - R package for Korean NLP http://cran.r-project.org/web/packages/KoNLP/index.html
? KoNLP v.0.80.0 버전 업(on CRAN now)
? KoNLPer - KoNLP 결과를 보내주는 flask with r 서버 dockerize http://konlper.duckdns.org/list
? KoSpacing - Automatic Korean word spacing with R
? KoSpacing : 한글 자동 띄어쓰기 패키지 공개
Library Scala
? Open Korean Text Processor - An Open-source Korean Text Processor
? twitter-korean-text - 트위터에서 만든 한국어 처리기
LSA
? Latent Semantic Variable Models
? Word vectors using LSA, Part - 2
? 숨은의미분석 LSA(Latent Semantic Analysis)
LSH
? LSH (Locality sensitive hashing)
MOOC, Lecture
? A Primer on Neural Network Models for Natural Language Processing
? List of free resources to learn Natural Language Processing
? Learn Natural Language Processing
? CS224d: Deep Learning for Natural Language Processing
? CS224d 2017 video subtitles translation project for everyone
? CS224n: Natural Language Processing with Deep Learning
? CS 224N: TensorFlow Tutorial
? Lecture Collection | Natural Language Processing with Deep Learning (Winter 2017)
? CS224n: Natural Language Processing with Deep Learning | Winter 2019
? CS224U: Natural Language Understanding
? Distributional word representations
? Deep Learning for Natural Language Processing: 2016-2017
? Lecture 8 - Generating Language with Attention Chris Dyer
? CS4650 and CS7650 ("Natural Language") at Georgia Tech
? CS 447: Natural Language Processing
? CS 20SI: Tensorflow for Deep Learning Research
? YSDA Natural Language Processing course
? NLP_COURSE: A Deep Learning YSDA Natural Language Processing Course By GitHub
Named Entity
? Named Entity Recognition: Examining the Stanford NER Tagger
? 한국어 개체명 인식 기술(Named Entity Recognition)
? NeuroNER - A Named-Entity Recognition Program based on Neural Networks and Easy to Use
? Entity extraction using Deep Learning
? 기사의 각 단어를 organisation, person, miscellaneous 및 other의 네가지 범주로 태그
? 그런 다음 기사에서 가장 두드러진 조직과 이름을 찾아 딥러닝 모델은 각 단어를 위의 4가지 범주로 분류
? 그런 다음 원치 않는 태깅을 필터링하고 가장 유명한 이름과 조직을 찾는 규칙 기반 접근 방식
? Named Entity Recognition: Milestone Models, Papers and Technologies
? Introduction to Named Entity Recognition
? Named Entity Recognition (NER), Meeting Industry’s Requirement by Applying state-of-the-art Deep
? Named Entity Recognition with NLTK and SpaCy
? Multilingual Named Entity Recognition: Research to Reality
? etagger - reference tensorflow code for named entity recognition
News
? 마커, “뉴스, 다 읽지 마세요. 형광펜 처리된 중요한 부분만 보세요”
? “수 없이 쏟아지는 읽을거리, 중요한 것만 밑줄 쳐 드립니다”, 마커 정철현 대표
? ‘뉴욕타임스’, 머신러닝 기반 자동 태그 시스템 개발
? 지난 26년간 언론에서 가장 중요한 정보원은 누구였을까?
? 세월호 참사 1년 동안의 언론보도를 통해 드러난 언론매체의 정치적 경도
? 세월호 참사 1년 동안의 언론보도를 통해 드러난 언론매체의 정치적 경도
? 네이버 뉴스 댓글 ‘남성’ 많고 ’10대·여성’ 적고
? 김경훈: 뉴스를 재미있게 만드는 방법; 뉴스잼 - PyCon APAC 2016
? 20160813, PyCon2016APAC 뉴스를 재미있게 만드는 방법; 뉴스잼
? ‘2억9천만원 아파트’ 기사에 달린 댓글로 본 사회학
? Google starts highlighting fact-checks in News
? Extract News In Three Words Using Triples
Ontology
? jena Ontology API와 sparQL을 사용하여 검색시스템 만들기
Paper
? Semantics, Representations and Grammars for Deep Learning
? Language Understanding for Text-based Games Using Deep Reinforcement Learning
? Linguistic Knowledge as Memory for Recurrent Neural Networks
? Recent Trends in Deep Learning Based Natural Language Processing
? 57 SUMMARIES OF MACHINE LEARNING AND NLP RESEARCH
? Paper in Natural Laguage Processing
? Paper Digest: EMNLP 2019 Highlights
Parser
? Announcing SyntaxNet: The World’s Most Accurate Parser Goes Open Source
? Google 자연어 처리 오픈소스 SyntaxNet 공개
? Google SyntaxNet 설치하기(Ubuntu / Mac)
? SyntaxNet in context: Understanding Google's new TensorFlow NLP model
? Structured Training for Neural Network Transition-Based Parsing
? github.com/dsindex/syntaxnet
? github.com/dsindex/parsing-syntaxnet
? An Upgrade to SyntaxNet, New Models and a Parsing Competition
? Grammatical Framework - A programming language for multilingual grammar applications
? Syntactic Parsing of Web Queries with Question Intent
? Novel Modeling of Syntactic Parsing for Web Queries
? The Yahoo Query Treebank, V. 1.0
? Language Data - Yahoo Answers Query Treebank, version 1.0
? Phoenix Server - a Galaxy-wrapped version of the Phoenix robust semantic CFG parser
? SLING: A Natural Language Frame Semantic Parser
? SQLova - a neural semantic parser translating natural language utterance to SQL query
QA Question Answer
? SQuAD - The Stanford Question Answering Dataset
? BiDAF - Bi-Directional Attention Flow for Machine Comprehension
? SQuAD - The Stanford Question Answering Dataset
? www.facebook.com/groups/AIKoreaOpen/permalink/1207284209305687
? Implementation of Dynamic memory networks by Kumar et al. http://arxiv.org/abs/1506.07285
? Generating Factoid Questions With Recurrent Neural Networks: The 30M Factoid Question-Answer Corpus
? Deep Language Modeling for Question Answering using Keras
? Deep Language Modeling for Question Answering using Keras
? FRDF Frame Semantic-based QA system
? FRDF: Frame-semantic-based QA system
? KBQA: An Online Template Based Question Answering System over Freebase
? KBQA: Learning Question Answering over QA Corpora and Knowledge Bases
? Question Answering System using Multiple Information Source and Open Type Answer Merge
? SearchQA
? START - Natural Language Question Answering System
? TriviaQA: A Large Scale Dataset for Reading Comprehension and Question Answering
? Reading Wikipedia to Answer Open-Domain Questions
? SIGIR2017에서 발표한 RNN을 이용한 자연어 질의 변환
? PR-037: Ask me anything: Dynamic memory networks for natural language processing
? Learning to reason by reading text and answering questions
? Learning to reason by reading text and answering questions
? MRQA 2018: Machine Reading for Question Answering
? Transparency-by-Design networks (TbD-nets)
? Building a Question-Answering System from Scratch— Part 1
? 2018 06-11-active-question-answering
? Bilinear attention networks for visual question answering
? Presenting Multitask Learning as Question Answering: The Natural Language Decathlon
? ATOMIC An Atlas of Machine Commonsense for If-Then Reasoning
Sentiment
? A comparison of open source tools for sentiment analysis
? 감정어휘 평가사전과 의미마디 연산을 이용한 영화평 등급화 시스템
? TextBlob Sentiment: Calculating Polarity and Subjectivity python
? Natural Language Basics with TextBlob
? Modern Methods for Sentiment Analysis
? LSTM Networks for Sentiment Analysis
? Sentiment Analysis using LSTM network
? KOrean Sentiment Analysis Corpus, KOSAC
? Naver sentiment movie corpus v1.0
? Naver Movie Sentiment Classification
? The emotional arcs of stories are dominated by six basic shapes
? The emotional arcs of stories are dominated by six basic shapes
? dracula.sentimentron.co.uk/sentiment-demo
? Sentiment Analysis and Aspect classification for Hotel Reviews
? Exploring Sentiment in Literature with Deep Learning
? Learning when to skim and when to read
? 감성분석 API
? Sentiment analysis on Twitter using word2vec and keras
? Sentiment analysis on forum articles using word2vec and Keras
? TWITTER SENTIMENT ANALYSIS USING COMBINED LSTM-CNN MODELS
? How to Develop an N-gram Multichannel Convolutional Neural Network for Sentiment Analysis
? 5 Things You Need to Know about Sentiment Analysis and Classification
? Sentiment analysis in Korean
? Detecting Sarcasm with Deep Convolutional Neural Networks
? Basic Data Cleaning/Engineering Session Twitter Sentiment Data
? Perform sentiment analysis with LSTMs, using TensorFlow
? Sentence classification by MorphConv
? How to build a simple text classifier with TF-Hub
? 예제의 텍스트 임베딩 함수가 estimator로 바로 피딩되는 바람에 feature vector 자체에 접근 불가능
? 이를 해결한 방법 demo_sentence_feature.ipynb
? Simplifying Sentiment Analysis using VADER in Python (on Social Media Text)
? Sentiment analysis : Frequency-based models
? Sentiment analysis : Frequency-based models
? Sentiment analysis : Machine-Learning approach
? Sentiment analysis : Machine-Learning approach
? A Beginner’s Guide on Sentiment Analysis with RNN
? Sentiment Classification with Natural Language Processing on LSTM
? Sentiment Analysis: Concept, Analysis and Applications
Similarity
? faiss - A library for efficient similarity search and clustering of dense vectors
? Fuzzy string matching using cosine similarity
? Pointwise mutual information
? FIVE MOST POPULAR SIMILARITY MEASURES IMPLEMENTATION IN PYTHON
? MinHash Tutorial with Python Code
? NMF 알고리즘을 이용한 유사한 문서 검색과 구현(1/2) matrix factorization
? NMF 알고리즘을 이용한 유사 문서 검색과 구현(2/2) sklearn을 이용한 구현
? String Matching and Database Merging Machine Learning to compare and join heterogeneous data from heterogeneous sources
? Brain's Pick: 단어 간 유사도 파악 방법
? Siamese LSTM을 이용한 Quora 질문 유사도 판별
? 한글 데이터 머신러닝 및 word2vec을 이용한 유사도 분석
? EUCLIDEAN DISTANCE FOR FINDING SIMILARITY
? AWS 람다(Lambda)로 실시간 추천하기 – 로켓펀치의 전문기술 정보
? WMD 문서 유사도 구하기 (word mover's distance)
? Chapter 3 : 단어 임베딩을 사용하여 텍스트 유사성 계산하기
? 11. Deep Learning Cookbook/03. 단어 임베딩을 사용하여 텍스트 유사성 계산하기
? 엘라스틱서치의 벡터(Vector) 필드와 텐서플로우를 이용한 문서 유사도 검색 (1) >Similarity Search #elasticsearch
Summarize
? Text summarization with TensorFlow
? How to Run Text Summarization with TensorFlow
? Text summarization with TensorFlow
? github.com/tensorflow/models/textsum
? sequence-to-sequence Learning
? dataset A Neural Attention Model for Abstractive Sentence Summarization
? NDC 2017 마이크로토크 - 프로그래머가 뉴스 읽는 법
? tldr - Text summarization service
? 24 A Serious NLP Application Text Auto Summarization using Python
? Summarizing Tweets in a Disaster
? Unsupervised Text Summarization using Sentence Embeddings
? An Introduction to Text Summarization using the TextRank Algorithm (with Python implementation)
? Text Summarization on the Books of Harry Potter
Spark
? Natural Language Processing With Apache Spark
? Introducing the Natural Language Processing Library for Apache Spark
? spark-nlp - Natural Language Understanding Library for Apache Spark
? Deep learning text NLP and Spark Collaboration . 한글 딥러닝 Text NLP &Spark
? Deep Learning and NLP with Spark by Andy Petrella and Melanie Warrick
? Classifying Text in Money Transfers with Apache Spark - Jose A. Rodriguez-Serrano
? Deep Learning and NLP with Spark - by Andy Petrella
? SF Scala, David Hall, ScalaNLP Epic
? Natural Language Processing with CNTK and Apache Spark - Ali Zaidi
? Text By the Bay 2015: Marek Kolodziej, Unsupervised NLP Tutorial using Apache Spark
? TextMining과 NaiveBayes분류 알고리즘
Speller
? How to Write a Spelling Corrector
? 한글 검색 질의어 오타 패턴 분석과 사용자 로그를 이용한 질의어 오타 교정 시스템 구축
? 사쿠라 훈민정음
? Word Prediction using Convolutional Neural Networks
Text Mining
? Kaggle Solution: What’s Cooking ? (Text Mining Competition)
? How to create a text mining algorithm with Python
? Natural Language Processing (NLP) &Text Mining Tutorial Using NLTK | NLP Training | Edureka
TextRank
? NDC 2017 마이크로토크 - 프로그래머가 뉴스 읽는 법
? Text Summarization with Gensim gensim의 textrank
? textacy: higher-level NLP built on spaCy text analysis based on spaCy
? python-rake 키워드 추출 패키지
? 한국어 3줄 요약기 - TextRank 알고리즘을 사용한 3줄 요약기 크롬 확장 앱
TFIDF, TF-IDF
? The fastest way to identify keywords in news articles — TFIDF with Wikipedia (Python version)
? Machine Learning with Text - TFIDF Vectorizer MultinomialNB Sklearn (Spam Filtering example Part 2)
? What is TF-IDF? The 10 minute guide
? How I used text mining to decide which Ted Talk to watch
? Keyword Extraction with TF-IDF and scikit-learn – Full Working Example
Topic Modeling
? Topic Modeling with LDA Introduction
? Text Mining 101: Topic Modeling
? Topic Modeling in Multi-Aspect Reviews
? Topic Modeling of Twitter Followers
? Topic Modelling in Python with NLTK and Gensim
? Extracting Hidden Topics in a Corpus
? Topic Modeling with Scikit Learn
? An NLP Approach to Mining Online Reviews using Topic Modeling (with Python codes)
? Topic Modeling with LSA, PLSA, LDA &lda2Vec
? tomotopy - Python package of Tomoto, the Topic Modeling Tool
? Python용 토픽 모델링 패키지 - tomotopy 개발
Translation
? Introduction to Neural Machine Translation with GPUs (Part 1)
? Introduction to Neural Machine Translation with GPUs (Part 2)
? Introduction to Neural Machine Translation with GPUs (part 3)
? 문자 단위의 Neural Machine Translation
? Jointly Modeling Embedding and Translation to Bridge Video and Language
? Tips on Building Neural Machine Translation Systems
? Subword Neural Machine Translation
? Machine Learning is Fun Part 5: Language Translation with Deep Learning and the Magic of Sequences
? Google's Neural Machine Translation System
? Peeking into the neural network architecture used for Google's Neural Machine Translation
? Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation 여러 언어를 동시에 번역하도록 학습했더니 한번도 학습에 사용한 적이 없는 언어쌍에 대해서도 번역이 가능
? ZERO-SHOT LEARNING FOR VISION AND MULTIMEDIA
? gtbot - 구글 번역 API를 이용한 슬랙 번역 봇입니다
? Open-Source Neural Machine Translation in PyTorch http://opennmt.net
? OpenNMT-Colab-Tutorial OpenNMT Colab Tutorial Pytorch &&Tensorflow
? onlinedoctranslator.com 구글 api를 사용해 만든 번역 서비스
? Deep Learning Takes on Translation
? TensorFlow에서 나만의 신경 기계 번역 시스템 구축
? Learned in translation: contextualized word vectors
? 카카오번역기가 양질의 대규모 학습 데이터를 확보하는 방법
? Machine Translation Without the Data
? How to Configure an Encoder-Decoder Model for Neural Machine Translation
? Neural Korean to English Machine Translater with Gluon
? A history of machine translation from the Cold War to deep learning
? UNdreaMT: Unsupervised Neural Machine Translation pytorch
? Neural Machine Translation : Everything you need to know
? Neural Translation Model with Attention
? Word Piece Model (a.k.a sentencepiece) RNN
? Character Word LSTM Language Models paper review
? Neural Machine Translation with Attention
? Neural Machine Translation With Attention Mechanism
? word2word - Easy-to-use word-to-word translations for 3,564 language pairs
? Unsupervised Word Segmentation for Neural Machine Translation and Text Generation
Voice
? THE COMPUTERS ARE LISTENING HOW THE NSA CONVERTS SPOKEN WORDS INTO SEARCHABLE TEXT
? Google voice search: faster and more accurate
? Baidu Deep Voice explained Part 2 — Training
? Neural Voice Cloning with a Few Samples
? Tutorial: Asynchronous Speech Recognition in Python
? 책 읽어주는 딥러닝: 배우 유인나가 해리포터를 읽어준다면 DEVIEW 2017
? Multi-speaker Tacotron in TensorFlow. 오픈소스 딥러닝 다중 화자 음성 합성 엔진. http://carpedm20.github.io/tacotron
? SPEECH TO TEXT(STT) 라이브러리와 프로세싱을 이용하여 음성인식 테스트하기
? Getting robots to understand speech: Using Watson’s Natural Language Classifier service
? Mozilla, 음성데이터세트 ‘딥스피치(DeepSpeech)’ 공개
? How to build a simple speech recognition app
? 딥러닝 음성합성 multi-speaker-tacotron(tacotron+deepvoice) 설치 및 사용법
? 딥 러닝 음성 인식에 필요한 훈련 데이터를 직접 만들어보자
? Towards end-to-end speech recognition
? How to Make a Speech Emotion Recognizer Using Python And Scikit-learn Librosa, Numpy, Soundfile, Scikit-learn, PyAudio
? AudioSet - A massive dataset of manually annotated audio events
? g2pK: g2p module for Korean 발음 생성 모듈. TTS의 전처리 모듈로 흔히 사용
? Kaldi Speech Recognition Toolkit
? Kaldi asr(automatic speech recognition) 음성인식 오픈소스 라이브러리 사용법 및 예제 정리
? 음성인식모델로 음성합성 데이터 만들기 (kaldi 음성 인식 모델 환경 구현)
? KoG2P - Korean grapheme-to-phone conversion in Python python 발음 생성 모듈
? Tabletop Bringing Tabletop Audio to Actions on Google through media responses
? Tacotron, Wavenet-Vocoder, Koearn TTS
? 딥러닝 음성합성 multi-speaker-tacotron(tacotron+deepvoice)설치 및 사용법
? Toolkits for robust speech processing
? wav2letter++ Introducing Wav2letter++ - How Facebook Implements Speech Recognition Systems Completely Based on Convolutional Neural Networks
? WaveNet: A Generative Model for Raw Audio
? voice Common Voice Project
Wikipedia
? practice
? A Multilingual Corpus of Automatically Extracted Relations from Wikipedia
? Exploring Wikipedia with Gremlin Graph Traversals
? Fact Extraction from Wikipedia Text
? LSA-ing Wikipedia with Apache Spark
? wiki - Command line tool to fetch summaries from mediawiki wikis, like Wikipedia
? What are the ten most cited sources on Wikipedia? Let’s ask the data
? Transforming Wikipedia into an accurate cultural knowledge quiz
? Wikipedia Data Science: Working with the World’s Largest Encyclopedia
? 한국어 위키백과내 주요 문서 16만개에 포함된 지식을 추출하여 객체(entity), 속성(attribute), 값(value)을 갖는 트리플 형식의 데이터 75만개
Word2Vec
? awesome-sentence-embedding - A curated list of pretrained sentence(and word) embedding models
? An Idiot’s Guide to Word2vec Natural Language Processing
? Modern Methods for Sentiment Analysis
? Word vectors (word2vec) on named entities and phrases - I
? Five crazy abstractions my Deep Learning word2vec model just did
? Neural Language Model and Word2Vec
? 2015 py con word2vec이 추천시스템을 만났을 때
? Word2Vec Vector Algebra Comparison - Python(Gensim) VS Scala(Spark)
? FastText and Gensim word embeddings
? 단어 임베딩의 원리와 gensim.word2vec 사용법
? models.word2vec – Deep learning with word2vec
? Word2vec with Gensim - Python
? Getting started with Word2Vec in Gensim and making it work!
? Gensim Word2Vec Tutorial – Full Working Example
? Fast Sentence Embeddings is a Python library that serves as an addition to Gensim
? Word2Vec Tutorial
? Commented (but unaltered) version of original word2vec C implementation
? Demystifying Neural Network in Skip-Gram Language Modeling
? word2vec, LDA, and introducing a new hybrid algorithm: lda2vec
? Bag of Words Meets Bags of Popcorn
? Bag of Words Meets Bags of Popcorn
? An introduction to Bag of Words and how to code it in Python for NLP
? Vector Representations of Words
? 단어의 벡터 표현 (Vector Representations of Words)
? The Amazing Power of Word Vectors
? word2vec
? How to giving a specific word to word2vec model in tensorflow
? tag2vec - 인스타그램 태그를 Word2vec으로 학습시킨 태그 벡터 공간입니다. https://tag2vec.herokuapp.com
? Making Sense of Everything with words2map
? github.com/leeyonghwan92/news_clustering 동국대학교 4학년 학생 졸업 프로젝트
? 한글을 이용한 데이터마이닝및 word2vec이용한 유사도 분석
? 5-1. 텐서플로우(TensorFlow)를 이용해 자연어를 처리하기(NLP) – Word Embedding(Word2vec)
? On word embeddings - Part 3: The secret ingredients of word2vec
? Ali Ghodsi, Lec [3,1]: Deep Learning, Word2vec
? Pre-trained word vectors of 30+ languages
? Play with word embeddings in your browser
? Introduction to Natural Language Processing (NLP) and Bias in AI
? NLP Research part 1. Vector Representations of Words
? Word2Vec 그리고 추천 시스템의 Item2Vec
? 박근혜 탄핵 결정문 전문 Word2Vec Visualization w/Tensorflow
? code.google.com/archive/p/word2vec
? Word2GM (Word to Gaussian Mixture)
? Deep Learning #4: Why You Need to Start Using Embedding Layers
? A non-NLP application of Word2Vec
? PR-027:GloVe - Global vectors for word representation
? Lecture 2 | Word Vector Representations: word2vec
? 번역에서 배우기 : 문맥화된 단어 벡터(contextualized word vector)
? Word embeddings in 2017: Trends and future directions
? Aerin Kim - Phrase2Vec In Practice #AIWTB 2016
? Using Word2vec for Music Recommendations
? Use Neural Networks to Find the Best Words to Title Your eBook
? Learning meaningful location embeddings from unlabeled visits
? bilm-tf
? word2vec, glove 등의 lookup 기반 embedding 기법과는 다르게 context word embedding을 사용해서 downstream task의 성능 향상
1. 대용량 corpus를 이용해서 2-layer bilstm lm 모델을 만들고
1. 각 timestep에 있는 h값에 대한 linear combination 결과를 현재 timestep의 word embedding으로 사용
1. combination weight는 downstream task의 cost function을 통해서 조정
? Word2Bits - Quantized Word Vectors
? Text2Shape: Generating Shapes from Natural Language by Learning Joint Embeddings
? Word2Vec 모델 기초
? Text Embedding Models Contain Bias. Here's Why That Matters
? WEAT 테스트는 목표 단어 세트(예 : 아프리카계 미국인 이름, 유럽계 미국인 이름, 꽃, 곤충)와 속성 단어 세트 (예 : "안정", "즐거운"또는 "불쾌한")를 모델이 연관시키는 정도를 측정
? 두개의 주어진 단어 사이의 연관성은 단어에 대한 임베딩 벡터 사이의 코사인 유사성으로 정의
? An Intuitive Understanding of Word Embeddings: From Count Vectors to Word2Vec
? node2vec: Embeddings for Graph Data
? PyData Tel Aviv Meetup: Node2vec - Elior Cohen
? 700x faster node2vec models: fastest random walks on a graph
? Word2Vec
? Text Classification With Word2Vec
? K Means Clustering Example with Word2Vec in Data Mining or Machine Learning
? ELMO DEEP CONTEXTUALIZED WORD REPRESENTATIONS
? The Current Best of Universal Word Embeddings and Sentence Embeddings
? Word Embeddings and Document Vectors
? Various Optimisation Techniques and their Impact on Generation of Word Embeddings
? Word Vector Representation for Korean: Evaluation Set
? Word2Vec — a baby step in Deep Learning but a giant leap towards Natural Language Processing
? Neural Network Embeddings Explained
? Beyond Word Embeddings Part 1
? Word2Vec For Phrases — Learning Embeddings For More Than One Word
? How to incorporate phrases into Word2Vec – a text mining approach
? Magnitude: a fast, simple vector embedding utility library
? When and Why does King - Man + Woman = Queen? (ACL 2019)
? Word2vec: fish + music = bass
? role2vec - A scalable Gensim implementation of "Learning Role-based Graph Embeddings" (IJCAI 2018)
? 기계는 사람의 말을 어떻게 이해할까? 워드 임베딩(Word Embedding)
? 기초적이지만 꽤 재미있는 word embedding 놀이
? 성지석-Deep contextualized word representations
? KCharEmb - Tutorial for character-level embeddings in Korean sentence classification
[출처] https://github.com/hyunjun/bookmarks/blob/master/nlp.md
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.

