Tag: rnn
All the articles with the tag "rnn".
| 종류 | 점수 함수 | Q의 출처 | K, V의 출처 | 출처 논문 |
|---|---|---|---|---|
| Bahdanau (additive) | decoder 직전 상태 | encoder 전체 | Bahdanau et al. (2015) | |
| Luong dot | decoder 현재 상태 | encoder 전체 | Luong et al. (2015) | |
| Luong general | decoder 현재 상태 | encoder 전체 | Luong et al. (2015) | |
| Luong concat | decoder 현재 상태 | encoder 전체 | Luong et al. (2015) | |
| Encoder self-attention | encoder 이전 층 | encoder 이전 층 | Vaswani et al. (2017) | |
| Masked self-attention | (뒤쪽 ) | decoder 이전 층 | decoder 이전 층 | Vaswani et al. (2017) |
| Encoder-decoder attention | decoder 이전 층 | encoder 출력 | Vaswani et al. (2017) |
Attention 메커니즘 정리 - Seq2Seq에서 Transformer까지
Lab11-5에서 Seq2Seq model을 공부하면서 입력 문장 전체를 vector 하나로 압축한다는 점이 계속 걸렸다.
def evaluate(pairs, source_vocab, target_vocab, encoder, decoder, target_max_length):
for pair in pairs:
print(">", pair[0])
print("=", pair[1])
source_tensor = tensorize(source_vocab, pair[0])
source_length = source_tensor.size()[0]
encoder_hidden = torch.zeros([1, 1, encoder.hidden_size]).to(device)
for ei in range(source_length):
_, encoder_hidden = encoder(source_tensor[ei], encoder_hidden)
decoder_input = torch.Tensor([[SOS_token]]).long().to(device) # 수정해야 작동
decoder_hidden = encoder_hidden
decoded_words = []
for di in range(target_max_length):
decoder_output, decoder_hidden = decoder(decoder_input, decoder_hidden)
_, top_index = decoder_output.data.topk(1) # 1개의 가장 큰 요소를 반환
if top_index.item() == EOS_token:
decoded_words.append("<EOS>")
break
else:
decoded_words.append(target_vocab.index2vocab[top_index.item()])
decoder_input = top_index.squeeze().detach()
predict_words = decoded_words
predict_sentence = " ".join(predict_words)
print("<", predict_sentence)
print("")Seq2Seq
Seq2Seq model은 아래와 같은 구조를 가지고 있다. 일종의 Encoder-Decoder 구조라고도 할 수 있는데 모든 입력을 다 받은 후에 출력을 생성하는 구조이다.
# load data
xy = np.loadtxt("data-02-stock_daily.csv", delimiter=",")
xy = xy[::-1] # reverse order
# split train-test set
train_size = int(len(xy) * 0.7)
train_set = xy[0:train_size]
test_set = xy[train_size - seq_length:]
앞서 언급한대로 scaling을 하고 학습하기 좋은 형태로 data를 가공해야 한다.
def minmax_scaler(data):
numerator = data - np.min(data, 0)
denominator = np.max(data, 0) - np.min(data, 0)
return numerator / (denominator + 1e-7)
train_set = minmax_scaler(train_set)
test_set = minmax_scaler(test_set)Timeseries
timeseries(시게열) data는 일정 시간 간격으로 배치된 data를 말한다. 매장의 시간별 매출, 요일별 주식 시가/종가 등이 여기에 속할 수 있다.
# data setting
x_data = []
y_data = []
# window를 오른쪽으로 움직이면서 자름
for i in range(0, len(sentence) - sequence_length):
x_str = sentence[i:i + sequence_length]
y_str = sentence[i + 1: i + sequence_length + 1]
print(i, x_str, '->', y_str)
x_data.append([char_dic[c] for c in x_str]) # x str to index (dict 사용)
y_data.append([char_dic[c] for c in y_str]) # y str to index
x_one_hot = [np.eye(dic_size)[x] for x in x_data]
X = torch.FloatTensor(x_one_hot)
Y = torch.LongTensor(y_data)
'''output
0 if you wan -> f you want
1 f you want -> you want
2 you want -> you want t
3 you want t -> ou want to
4 ou want to -> u want to
...
166 ty of the -> y of the s
167 y of the s -> of the se
168 of the se -> of the sea
169 of the sea -> f the sea.
'''RNN - longseq
앞서 살펴보았던 RNN 예제들은 모두 한 단어나 짧은 문장에 대해 RNN을 학습시키는 내용들이었다. 하지만 우리가 다루고 싶은 데이터는 더 긴 문장이거나 내용을 가질 가능성이 높다.
char_set = ['h', 'i', 'e', 'l', 'o']
# hyper parameters
input_size = len(char_set)
hidden_size = len(char_set)
learning_rate = 0.1
# data setting
x_data = [[0, 1, 0, 2, 3, 3]]
x_one_hot = [[[1, 0, 0, 0, 0],
[0, 1, 0, 0, 0],
[1, 0, 0, 0, 0],
[0, 0, 1, 0, 0],
[0, 0, 0, 1, 0],
[0, 0, 0, 1, 0]]]
y_data = [[1, 0, 2, 3, 3, 4]]
X = torch.FloatTensor(x_one_hot)
Y = torch.LongTensor(y_data)
마찬가지로 one-hot encoding하여 Tensor로 바꾼다. 다만 각 알파벳 변수에 배열을 저장하는 방식이 아니라 char_set에 저장된 알파벳을 x_data의 값을 인덱스로 불러오는 방식이다.
one-hot encoding은 x_data에 적용하여 학습한다.
RNN - hihello / charseq
hihello 문제는 같은 문자들이 다음 문자가 다른 경우 이를 예측하는 문제를 말한다. hihello에서 'h'와 'l'은 2번씩 등장하지만 어디에 문자가 위치하느냐에 따라 다음에 올 문자가 달라진다.
rnn = torch.nn.RNN(input_size, hidden_size)
outputs, _status = rnn(input_data)
print(outputs)
print(outputs.size())
'''output
tensor([[[-0.7497, -0.6135],
[-0.5282, -0.2473],
[-0.9136, -0.4269],
[-0.9136, -0.4269],
[-0.9028, 0.1180]],
[[-0.5753, -0.0070],
[-0.9052, 0.2597],
[-0.9173, -0.1989],
[-0.9173, -0.1989],
[-0.8996, -0.2725]],
[[-0.9077, -0.3205],
[-0.8944, -0.2902],
[-0.5134, -0.0288],
[-0.5134, -0.0288],
[-0.9127, -0.2222]]], grad_fn=<StackBackward>)
torch.Size([3, 5, 2])
'''RNN Basics
PyTorch에서 RNN은 in/output size만 잘 맞춰주면 바로 사용이 가능하다. "h, e, l, o" 4개의 알파벳으로 이루어진 데이터셋을 통해 2차원의 output(class가 2개)을 내는 RNN을 만들어볼 것이다.
activation과 weight를 명시하여 표현하면 다음과 같다.
이런 RNN의 구조를 응용하여 다음과 같은 구조들로 사용할 수 있다.

-
one to many : 하나의 입력을 받아 여러 출력을 내는 구조이다. 하나의 이미지를 받아 그에 대한 설명을 문장(여러개의 단어)으로 출력하는 것을 예로 들 수 있다.
-
many to one : 여러 입력을 받아 하나의 출럭을 내는 구조이다. 문장을 입력받아 그 문장이 나타내는 감정의 label을 출력하는 것을 예로 들 수 있다.
-
many to many : 2가지의 구조가 있는 것을 볼 수 있다.
- 입력이 다 끝나는 지점부터 여러 출력을 내는 구조로, 문장을 입력받아 번역하는 모델을 예로 들 수 있다. 이 경우 문장의 중간에 번역을 진행하면 다 끝나고 나서 문장의 의미가 달라질 수 있기 때문에 먼저 입력 문장을 다 듣고 번역을 진행하게 된다.
- 입력 하나하나를 받으면서 그때마다 모델의 출력을 내는 구조이다. 영상을 처리할 때 frame 단위의 이미지로 나눠 입력을 받은 후 각 frame을 입력 받을 때마다 처리하는 것을 예로 들 수 있다.
RNN intro
RNN은 sequential data를 잘 학습하기 위해 고안된 모델이다. Sequential data란 단어, 문장이나 시게열 데이터와 같이 데이터의 순서도 데이터의 일부인 데이터들을 말한다.