How does tree bagger handle NaN values

1 view (last 30 days)
Jason Summers
Jason Summers on 7 Feb 2020
Answered: Puru Kathuria on 27 Dec 2020
In building a random forest classifier I have some features with a large amount of NaN values, but it is not clear to me how Tree Bagger handles these NaNs. I've seen quite a bit of documentation of how that is handled in other high level programming languages, but I don't see explicitly how this is done in Matlab. Can anyone point me in the right direction so I can understand the default settings for this or user specified settings?

Answers (1)

Puru Kathuria
Puru Kathuria on 27 Dec 2020
General rules that are followed while NaN or missing values are encountered:
  • Rule1: The algorithm simply discards the data points where all the features have NaN values and does not use them while training.
  • Rule 2: If a data point have a few NaN feature values then the algorithm will find the split on the basis of valid values first.

Categories

Find more on Statistics and Machine Learning Toolbox in Help Center and File Exchange

Products


Release

R2017b

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!