Convolution neural networks (CNN)
Greetings! I'm Vishnu Vinay, a Computer Science and Engineering graduate holding a B. Tech degree. Currently immersed in the captivating world of Artificial Intelligence, I am on a quest for knowledge while pursuing a graduate certificate in Artificial Intelligence with Machine Learning. My passion lies in sharing insights and discoveries in the fields of AI, Machine Learning, Artificial General Intelligence, and Robotics through engaging blog posts. Proficient in Python, ML libraries, and algorithms, I find joy in developing and deploying ML models, with a focus on leveraging AWS Sagemaker. Join me on this exciting journey of unraveling the mysteries of AI through the lens of coding, exploration, and the ever-evolving landscape of machine learning. Let's embark on this knowledge-sharing adventure together!
CNN is a type of neural network which helps and performs well in image processing tasks.
In the case of neural networks (NN), it takes each pixel as a feature. However, we can simplify it by performing some manual preprocessing like combining a set of pixels and considering that set as a feature.
The problem is what set of pixels should we combine. Thus, we assign some weights to each pixel and later on adjust the weights to our requirements.
These above steps are complicated as they will require a lot of work and time.
We have the option of performing these steps automatically meaning, the network gets an image dataset as input and it will generate the required features by recognizing the patterns from it, which can be done by the techniques provided by CNN.
Eg:- We have an image dataset containing images of some people wearing spectacles and some not. CNN will try to capture the spectacle portion from an image, learns from it and then by using adequate weights with the pixels, it will combine that part into a single region or feature.
Working of CNN:-
First, let's talk about 1 CNN unit.
An image will be passed through this CNN unit, it will do some changes or generate a new image.
Captures a combination of nearby pixel values.
The CNN unit will have a filter. The filter is a k*k matrix, which consists of weights.
The filter is applied over the entire image.
The following equation is applied - {ΣjΣipijwij (0<=i<=k-1, 0<=j<=k-1)}. Here, w - weights, p - pixels, and (i, j) - matrix index values.
These values will give rise to a pattern.
Now, CNN will have multiple units. Each unit learns a different pattern from the image.
Initially, the weights will be random values. Later on, the CNN will decide the optimal weights automatically by using the patterns generated in each iteration.
Important factors in CNN:-
Stride - refers to the step size with which the convolutional filter/kernel moves across the input image.
When applying the convolution operation, the filter slides over the input with a certain stride value, moving horizontally and vertically.
The stride determines how much the filter shifts after each application.
A larger stride value leads to a smaller output size, as the filter covers a larger area in each step.
Similarly, in the case of stride = 1, the input size and output size will be the same.
A smaller stride value isn't considered because it leads to a larger output size.

Padding - involves adding extra layers of pixels around the input image before applying the convolution operation.
It helps to retain spatial information and mitigate the size reduction that occurs during convolutional operations.
Padding is applied by adding zeros (zero-padding) around the edges of the input.
Types of padding:-
Valid (No Padding): When no padding is applied, the convolutional filter is only applied to positions where it completely overlaps with the input.
Same (Zero Padding): is used to preserve the spatial dimensions of the input to match the output dimensions. It achieves this by adding zeros evenly around the input
If stride = 1 and the same padding is applied, then the input and output sizes will be the same.
If stride != 1 and the same padding is applied, then the output size is nearly equal to n / s. Where n is input image size or dimensions and s is stride.

Channels:-
A black-and-white image will have 2 dimensions.
Similarly, colored images have 3 dimensions i.e., RGB (red, green, blue) values which are packed on top of one other. These are called channels.
Handling channels in CNN -
Applying 2D filters over the image. Eg:- image dimension = 28*28*3, filter size = 4*4
Applying filters with depth i.e., 3D filters which fill up the entire portion. Eg:- image dimension = 28*28*3, filter = 4*4*3

Both layers are CNN layers. First one with 5 units and the second with 4 units. Also, padding = same.
Our single image gets through the first layer's units and forms 5 different images (all will have some changes and some can be completely new images) from each unit respectively.
The output from these 5 units gets passed through each unit in the second layer. Thereby, each unit has more data compared to the first layer. Thus, it can generate better results.
However, the final size will not be 28*28*4 because in between these layers, a pooling layer exists (which is discussed in detail below).
Pooling layer:-
is used to progressively reduce the spatial size (width and height) of the input feature while retaining the most important information.
Added after convolution layer.
Reduces the image size by combining pixel values.
Types:-
Average pooling - calculates the average value of the region.

Max pooling - selects the maximum value within each pooling region.

Data flow in CNN:-


