Feature scaling
Greetings! I'm Vishnu Vinay, a Computer Science and Engineering graduate holding a B. Tech degree. Currently immersed in the captivating world of Artificial Intelligence, I am on a quest for knowledge while pursuing a graduate certificate in Artificial Intelligence with Machine Learning. My passion lies in sharing insights and discoveries in the fields of AI, Machine Learning, Artificial General Intelligence, and Robotics through engaging blog posts. Proficient in Python, ML libraries, and algorithms, I find joy in developing and deploying ML models, with a focus on leveraging AWS Sagemaker. Join me on this exciting journey of unraveling the mysteries of AI through the lens of coding, exploration, and the ever-evolving landscape of machine learning. Let's embark on this knowledge-sharing adventure together!
Normalization - Changing the values of features in dataset to a range of between 0 and 1.
Standardization - Changing mean to 0 and standard deviation to 1 for all features of a given dataset. Range of values can be any lower and upper limit (not necessary to be between 0 and 1).
Scaling is required in case of algorithms which requires gradient descent calculations.
Normalization code:
import pandas as pd
import numpy as np
from sklearn import datasets
# getting the dataset
wine = datasets.load_wine()
# converting data to dataframe
wine_df = pd.DataFrame(wine.data)
# setting the column names
wine_df.columns = wine.feature_names
# By running this code, we can the see the min and max is not ranging between 0 and 1.
wine_df.describe()
# getting only the input columns in numpy array format
X = wine.data
# Applying normalization
from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
scaler.fit(X)
wine_scaled = scaler.transform(X)
wine_scaled # final scaled data
# Converting it to dataframe
wine_scaled_df = pd.DataFrame(wine_scaled)
# setting the column names
wine_scaled_df.columns = wine.feature_names
# After the normalization, after running this code we can that the values are ranging between 0 and 1
wine_scaled_df.describe()
Before normalization -

After normalization -

Standardization code:
import pandas as pd
import numpy as np
from sklearn import datasets
# getting the dataset
wine = datasets.load_wine()
# converting data into dataframe
wine_df = pd.DataFrame(wine.data)
# setting up the column names
wine_df.columns = wine.feature_names
# After running this code we will see that the mean and standard deviation (std) values are not 0 and 1 respectively.
wine_df.describe()
# getting only the input columns in numpy array format
X = wine.data
# Applying standardization
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
wine_scaled = scaler.fit_transform(X)
# Converting it to dataframe
wine_scaled_df = pd.DataFrame(wine_scaled)
# Setting up the column names
wine_scaled_df.columns = wine.feature_names
# After running this code, we can see that the mean is 0 and std is 0 for all features.
wine_scaled_df.describe().round(2)
Before standardization -

After standardization -
