Proper use of ClassificationTree.fit for categorical variables?

Question

the cyclist 2013-11-8

0
链接

此问题的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/105398-proper-use-of-classificationtree-fit-for-categorical-variables

评论： the cyclist 2013-11-8

The documentation for fitting classification trees states that X needs to be a floating point array, but also indicates that X can represent categorical variables (using the 'CategoricalPredictors' Name-Value argument).

Is the proper way to handle this to

(1) take the categorical variable, e.g.

category1 = {'duck','duck','goose','squash','quartz'}';
category2 = {'animal','animal','animal','vegetable','mineral'}';

(2) run those through grp2idx()

numcat1 = grp2idx(category1);
numcat2 = grp2idx(category2);

(3) Embed those in my X:

X = [numcat1 numcat2 otherTrulyNumericalVariables]

(4) Identify those as categorical

tree = ClassificationTree.fit(X,Y,'CategoricalPredictors',[1 2])

Seems like that's probably right, but I'd love an expert to vet that idea. The documentation doesn't have a categorical example.

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

请先登录，再进行评论。

请先登录，再回答此问题。

Answer 1

Ilya 2013-11-8

0
链接

此回答的直接链接

https://ww2.mathworks.cn/matlabcentral/answers/105398-proper-use-of-classificationtree-fit-for-categorical-variables#answer_114600

在 MATLAB Online 中打开

Yes, this would be one way to accomplish this. You'd have to be careful when you convert new data to numeric for prediction. If the new data are missing a level (for example, 'goose' does not appear in the value set), grp2idx can return different indices for the same categorical values. One way to avoid this pitfall would be by using the nominal type and specifying the level order explicitly, for example:

category1 = nominal({'duck','duck','goose','squash','quartz'},...
      [],{'goose','squash','quartz' 'duck'})
numcat1 = double(category1)

Depending on how you get your data, you might find it easier to put your entire data (numeric and categorical variables) into a table or, if you are not in R2013b yet, into a dataset object and then extract numeric and categorical variables from that object.

1 个评论
显示 -1更早的评论隐藏 -1更早的评论

the cyclist 2013-11-8

Thanks! I didn't know about the nominal() command. I have typically solved the missing-level problem you describe by assigning the categories of the test data using the ismember() command against the unique categories of the training data (with the second argument of grp2idx).

请先登录，再进行评论。

Proper use of ClassificationTree.fit for categorical variables?

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

采纳的回答

1 个评论
显示 -1更早的评论隐藏 -1更早的评论

更多回答（0 个）

另请参阅

类别

标签

产品

Community Treasure Hunt

Proper use of ClassificationTree.fit for categorical variables?

0 个评论 显示 -2更早的评论隐藏 -2更早的评论

采纳的回答

1 个评论 显示 -1更早的评论隐藏 -1更早的评论

更多回答（0 个）

另请参阅

类别

标签

产品

Community Treasure Hunt

0 个评论
显示 -2更早的评论隐藏 -2更早的评论

1 个评论
显示 -1更早的评论隐藏 -1更早的评论