[pytorch] Using BERT with Pytorch

Using BERT with Pytorch

A super-easy practical guide to build you own fine tuned BERT based architecture using Pytorch.

Bert image — sesame street
BERT input presentation [1]
from pytorch_pretrained_bert.tokenization import BertTokenizertokenizer = BertTokenizer.from_pretrained(args.bert_model, do_lower_case=args.do_lower_case)def get_tokenized_samples(samples, max_seq_length, tokenizer):
    """
    we assume a function label_map that maps each label to an index or vector encoding. Could also be a dictionary.
    :param samples: we assume struct {.text, .label) 
    :param max_seq_length: the maximal sequence length
    :param tokenizer: BERT tokenizer
    :return: list of features
    """

    features = []
    for sample in samples:
        textlist = sample.text.split(' ')
        labellist = sample.label
        tokens = []
        labels = []
        for i, word in enumerate(textlist):
            token = tokenizer.tokenize(word) #tokenize word according to BERT
            tokens.extend(token)
            label = labellist[i]
            # fit labels to tokenized size of word
            for m in range(len(token)):
                if m == 0:
                    labels.append(label)
                else:
                    labels.append("X")
        # if we exceed max sequence length, cut sample
        if len(tokens) >= max_seq_length - 1:
            tokens = tokens[0:(max_seq_length - 2)]
            labels = labels[0:(max_seq_length - 2)]
            
        ntokens = []
        segment_ids = []
        label_ids = []
        # start with [CLS] token
        ntokens.append("[CLS]")
        segment_ids.append(0)
        label_ids.append(label_map(["[CLS]"]))
        for i, token in enumerate(tokens):
            # append tokens
            ntokens.append(token)
            segment_ids.append(0)
            label_ids.append(label_map(labels[i]))
        # end with [SEP] token
        ntokens.append("[SEP]")
        segment_ids.append(0)
        label_ids.append(label_map(["[SEP]"]))
        # convert tokens to IDs
        input_ids = tokenizer.convert_tokens_to_ids(ntokens)
        # build mask of tokens to be accounted for
        input_mask = [1] * len(input_ids) 
        while len(input_ids) < max_seq_length:
            # pad with zeros to maximal length
            input_ids.append(0)
            input_mask.append(0)
            segment_ids.append(0)
            label_ids.append([0] * (len(label_list) + 1))

        features.append((input_ids,
                              input_mask,
                              segment_ids,
                              label_id))
    return features
Fine Tune BERT pre-training to your task [1]
from pytorch_pretrained_bert.modeling import BertPreTrainedModel, BertModelclass MyBertBasedModel(BertPreTrainedModel):
    """
    MyBertBasedModel inherits from BertPreTrainedModel which is an abstract class to handle weights initialization and
        a simple interface for downloading and loading pre-trained models.
    """

    def __init__(self, config, num_labels):
        super(MyBertBasedModel, self).__init__(config)
        self.num_labels = num_labels
        self.bert = BertModel(config) # basic BERT model
        self.dropout = torch.nn.Dropout(config.hidden_dropout_prob)
        self.classifier = torch.nn.Linear(config.hidden_size, num_labels)
        self.apply(self.init_bert_weights)


    def forward(self, input_ids, token_type_ids=None, attention_mask=None, labels=None):
        sequence_output, _ = self.bert(input_ids, token_type_ids, attention_mask, output_all_encoded_layers=False)
        # now you can implement any architecture that receives bert sequence output
        sequence_output = self.dropout(sequence_output)
        logits = self.classifier(sequence_output)

        if labels is not None:
            loss_fct = MyLoss()
            # it is important to activate the loss only on un-padded inputs
            active_loss = attention_mask.view(-1) == 1
            active_logits = logits.view(-1, self.num_labels)[active_loss]
            active_labels = labels.view(-1, self.num_labels)[active_loss]
            loss = loss_fct(active_logits, active_labels)
            return loss
        else:
            return logits
경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

train_tokenized_samples = get_tokenized_samples(
    train_samples, args.max_seq_length, tokenizer)model = MyBertBasedModel.from_pretrained(args.bert_model,
          num_labels = num_labels)model.train()
for range(n_epochs):
    for sample in train_tokenized_samples:
        input_ids, input_mask, segment_ids, label_ids = sample
        loss = model(input_ids, segment_ids, input_mask, label_ids)
        loss.backward()
        optimizer.step()

 

 

 

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
공지 오라클 기본 샘플 데이터베이스 졸리운_곰 2014.01.02 86034
공지 [SQL컨셉] 서적 "SQL컨셉"의 샘플 데이타 베이스 SAMPLE DATABASE of ORACLE 가을의 곰을... 2013.02.10 78561
공지 [G_SQL] Sample Database 가을의 곰을... 2012.05.20 95291
101 [NoSQL] 성공적인 NoSQL 도입을 위한 키포인트 : NoSQL 데이터 모델링 file 졸리운_곰 2024.08.10 1080
100 [NoSQL] [MongoDB] NOSQL 데이터 모델링 기법 살펴보기 file 졸리운_곰 2024.08.09 1002
99 [NoSQL][MongoDB] Truncate a collection 졸리운_곰 2023.06.04 1271
98 [NoSQL] MongoDB 인증 모드 (password) 설정 졸리운_곰 2023.03.26 1417
97 [NoSQL] [Redis] Redis Persistence(영속성) 졸리운_곰 2021.04.11 1644
96 [NoSQL] [Cloud] Redis 설치, 사용 방법, 데이터 백업을 위한 RDB & AOF 개념 및 간단한 Redis 사용 사례 연구 file 졸리운_곰 2021.04.11 1716
95 [mongodb , 몽고디비] How to Use MongoDB Comparison Query Operators in Java 졸리운_곰 2021.02.19 1508
94 [mongodb, 몽고디비] How do you query for “is not null” in Mongo? 졸리운_곰 2021.02.19 1268
93 [MongoDB] 확장 검색 쿼리 - Aggregation 파이프라인 스테이지(Pipline Stage)와 표현식 졸리운_곰 2021.02.16 1192
92 [MongoDB] 확장 검색 쿼리 - 범용 Aggregation 졸리운_곰 2021.02.16 1347
91 [MongoDB] 확장 검색 쿼리 - Aggregation의 목적 및 작동방식 file 졸리운_곰 2021.02.16 1067
90 [MongoDB] 확장 검색 쿼리 - 집계 파이프 라인연산자 종류 졸리운_곰 2021.02.16 1140
89 [MongoDB] 확장 검색 쿼리 - Aggregation 파이프라인 스테이지(Pipline Stage)와 표현식 졸리운_곰 2021.02.16 1299
88 [mongodb, 몽고디비 쿼리] Mongodb query on substring of a field 졸리운_곰 2021.02.16 1305
87 [mongodb]Spring boot와 mongoDB 연동하기 그리고 REST API 졸리운_곰 2021.01.21 1704
86 [WebApp / Express] 간단한 MongoDB Middleware 만들기 졸리운_곰 2020.12.29 1152
85 [mongodb] Error: network error while attempting to run command 'isMaster' on host '127.0.0.1:27017' file 졸리운_곰 2020.12.16 1450
84 [mongodb] How to Update a Document in MongoDB using Java 졸리운_곰 2020.12.14 2327
83 [Java, MongoDB] Mapping a BSON MongoDB document to a MyClass.class object? 졸리운_곰 2020.12.09 1820
82 [mongodb] MongoDB CRUD 동작의 이해 졸리운_곰 2020.12.07 1284
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED