# Decision tree (ML) theory

A tree like graph which uses branching method to show the output of every decision.

* Nodes - each attribute / feature of a dataset, which asks a question. Then, the edges from these nodes will be represent all the possible questions for it and reaches the end when it reaches the class label.
    
* Node types -&gt;
    
    * Root node - the first node from the where the branching starts.
        
    * Leaf nodes - sub nodes after the root node.
        
    * Pure nodes - nodes with the class label or once which has reached the conclusion.
        
* The entire process is recursive, it keeps on repeating for sub nodes. Terminates once the leaf node reaches a class label.
    

![](https://cdn-images-1.medium.com/max/824/0*J2l5dvJ2jqRwGDfG.png align="left")

*Source:* [https://cdn-images-1.medium.com/max/824/0\*J2l5dvJ2jqRwGDfG.png](https://cdn-images-1.medium.com/max/824/0*J2l5dvJ2jqRwGDfG.png)

* Decision tree stopping conditions:-
    
    1. We obtain a pure node and no further decision is to be made for it.
        
    2. When all the given attributes have considered and no more attributes are present for splitting.
        

Algorithms for decision tree:-

1. ID3 (Iterative Dichotomiser 3) - creates a multiway tree (tree with more than 2 children). Finds maximum information (information gain) for each node i.e., splitting of each node as much as possible. Afterwards, pruning (removing) is done to improve accuracy or generalize the tree further according to the given data. Can only handle categorical attributes.
    
2. C4.5 - successor to ID3, chooses the best attribute with which to do the splitting based on finding the information gain for each attribute at each step.
    
    Can handle continuous attributes.
    
3. CART (Classification and regression tasks) - Similar to C4.5, it can handle both categorical (classification, uses Gini index for splitting) and continuous (regression, uses mean square error for splitting).
    

Terms related to decision tree:

1. Entropy - amount of impurity in a dataset.
    
2. Information gain - amount of information an attribute can give. Higher the value better the split.
    
3. Gini index - similar to information gain except that a lower gini index value indicates a pure subset.
