confused about random_state in decision tree of scikit learn

Tags:

Confused about random_state parameter, not sure why decision tree training needs some randomness. My thoughts, (1) is it related to random forest? (2) is it related to split training testing data set? If so, why not use training testing split method directly (http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.train_test_split.html)?

http://scikit-learn.org/stable/modules/generated/sklearn.tree.DecisionTreeClassifier.html

>>> from sklearn.datasets import load_iris >>> from sklearn.cross_validation import cross_val_score >>> from sklearn.tree import DecisionTreeClassifier >>> clf = DecisionTreeClassifier(random_state=0) >>> iris = load_iris() >>> cross_val_score(clf, iris.data, iris.target, cv=10) ...                              ... array([ 1.     ,  0.93...,  0.86...,  0.93...,  0.93...,         0.93...,  0.93...,  1.     ,  0.93...,  1.      ])

regards, Lin

312

asked Aug 26 '16 03:08

Lin Ma

1 Answers

This is explained in the documentation

The problem of learning an optimal decision tree is known to be NP-complete under several aspects of optimality and even for simple concepts. Consequently, practical decision-tree learning algorithms are based on heuristic algorithms such as the greedy algorithm where locally optimal decisions are made at each node. Such algorithms cannot guarantee to return the globally optimal decision tree. This can be mitigated by training multiple trees in an ensemble learner, where the features and samples are randomly sampled with replacement.

So, basically, a sub-optimal greedy algorithm is repeated a number of times using random selections of features and samples (a similar technique used in random forests). The random_state parameter allows controlling these random choices.

The interface documentation specifically states:

If int, random_state is the seed used by the random number generator; If RandomState instance, random_state is the random number generator; If None, the random number generator is the RandomState instance used by np.random.

So, the random algorithm will be used in any case. Passing any value (whether a specific int, e.g., 0, or a RandomState instance), will not change that. The only rationale for passing in an int value (0 or otherwise) is to make the outcome consistent across calls: if you call this with random_state=0 (or any other value), then each and every time, you'll get the same result.

answered Oct 02 '22 14:10

Ami Tavory

Related questions
                            
                                How to force PyYAML to load strings as unicode objects?
                            
                                SHA-256 implementation in Python
                            
                                How to multiply functions in python?
                            
                                Django - exception handling best practice and sending customized error message
                            
                                Best video manipulation library for Python? [closed]
                            
                                Is there a library function in Python to turn a generator-function into a function returning a list?
                            
                                What is the fastest way to output large DataFrame into a CSV file?
                            
                                How do I run doctests with PyCharm?
                            
                                Most Pythonic way to declare an abstract class property
                            
                                Customize module search path (PYTHONPATH) via pipenv
                            
                                Resampling Minute data
                            
                                python: iterating through a dictionary with list values
                            
                                Fastest way to parse large CSV files in Pandas
                            
                                What is the difference between random.normalvariate() and random.gauss() in python?
                            
                                What's the difference between dtype and converters in pandas.read_csv?
                            
                                Python - TypeError - TypeError: '<' not supported between instances of 'NoneType' and 'int'
                            
                                Installing pygraphviz on Windows 10 64-bit, Python 3.6
                            
                                How do I get PyCharm to show entire error diffs from pytest?
                            
                                Why are main runnable Python scripts not compiled to pyc files like modules? [duplicate]
                            
                                How can I access s3 files in Python using urls?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

confused about random_state in decision tree of scikit learn

Tags:

python

machine-learning

python-2.7

scikit-learn

decision-tree

Lin Ma

People also ask

1 Answers

Ami Tavory

Recent Activity

Donate For Us