selfAttentionLayer can't process sequence-to-label problem?

Question

cui,xingxing 2024-1-5

1
链接

此问题的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/2066601-selfattentionlayer-can-t-process-sequence-to-label-problem

编辑： cui,xingxing 2024-4-27

selfAttentionLayer why can't handle the following simple sequence classification problem, already through the flattenLayer into one-dimensional data, on the contrary, lstm specify "outputMode" as "last" will pass.

% Here use simple data, for demonstration purposes only
XTrain = rand(3,200,1000); % dims "CTB"
TTrain = categorical(randi(4,1000,1));
% define my layers
numClasses = numel(categories(TTrain));
layers = [inputLayer(size(XTrain),"CTB");
    flattenLayer;
    selfAttentionLayer(6,48);
    % lstmLayer(20,OutputMode="last"); % use lstmLayer is ok!
    layerNormalizationLayer;
    fullyConnectedLayer(numClasses);
    softmaxLayer];
net = dlnetwork(layers);
% train network
lossFcn = "crossentropy";
options = trainingOptions("adam", ...
    MaxEpochs=1, ...
    InitialLearnRate=0.01,...
    Shuffle="every-epoch", ...
    GradientThreshold=1, ...
    Verbose=true);
netTrained = trainnet(XTrain,TTrain,net,lossFcn,options);
Error using trainnet
Number of observations in predictors (1000) and targets (1) must match. Check that the data and network are consistent.

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

请先登录，再进行评论。

请先登录，再回答此问题。

Answer 1

cui,xingxing 2024-1-7

0
链接

此回答的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/2066601-selfattentionlayer-can-t-process-sequence-to-label-problem#answer_1384691

编辑：cui,xingxing 2024-4-27

在 MATLAB Online 中打开

In terms of the output feature map dimensions, there is a time "T" dimension that has to be eliminated in order to match the output dimensions, which can usually be done by indexing1dLayer. So the layers array is added before the fullyConnectedLayer.

% Here use simple data, for demonstration purposes only
XTrain = rand(3,200,1000); % dims "CTB"
TTrain = categorical(randi(4,1000,1));
% define my layers
numClasses = numel(categories(TTrain));
layers = [inputLayer(size(XTrain),"CTB");
    flattenLayer;
    selfAttentionLayer(6,48);
    % lstmLayer(20,OutputMode="last"); % use lstmLayer is ok!
    layerNormalizationLayer;
    
    indexing1dLayer; % Add this!!!
    fullyConnectedLayer(numClasses);
    softmaxLayer];
net = dlnetwork(layers);
% train network
lossFcn = "crossentropy";
options = trainingOptions("adam", ...
    MaxEpochs=1, ...
    InitialLearnRate=0.01,...
    Shuffle="every-epoch", ...
    GradientThreshold=1, ...
    Verbose=true);
netTrained = trainnet(XTrain,TTrain,net,lossFcn,options);
    Iteration    Epoch    TimeElapsed    LearnRate    TrainingLoss
    _________    _____    ___________    _________    ____________
            1        1       00:00:02         0.01          1.5374
            7        1       00:00:06         0.01          1.5272
Training stopped: Max epochs completed

-------------------------Off-topic interlude-------------------------------

I am currently looking for a job in the field of CV algorithm development, based in Shenzhen, Guangdong, China. I would be very grateful if anyone is willing to offer me a job or make a recommendation. My preliminary resume can be found at: https://cuixing158.github.io/about/ . Thank you!

Email: cuixingxing150@gmail.com

5 个评论
显示 3更早的评论隐藏 3更早的评论

DGM 2024-3-5

Posted as a comment-as-flag by chang gao:

Useful answer.

jingwen 2024-4-15

Your answer helps me! Thank you

请先登录，再进行评论。

selfAttentionLayer can't process sequence-to-label problem?

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

采纳的回答

5 个评论
显示 3更早的评论隐藏 3更早的评论

更多回答（0 个）

另请参阅

类别

标签

产品

版本

Community Treasure Hunt

selfAttentionLayer can't process sequence-to-label problem?

0 个评论 显示 -2更早的评论隐藏 -2更早的评论

采纳的回答

5 个评论 显示 3更早的评论隐藏 3更早的评论

更多回答（0 个）

另请参阅

类别

标签

产品

版本

Community Treasure Hunt

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

5 个评论
显示 3更早的评论隐藏 3更早的评论