Skip to main content

Command Palette

Search for a command to run...

Gradient descent

Updated
•2 min read•View as Markdown
V

Greetings! I'm Vishnu Vinay, a Computer Science and Engineering graduate holding a B. Tech degree. Currently immersed in the captivating world of Artificial Intelligence, I am on a quest for knowledge while pursuing a graduate certificate in Artificial Intelligence with Machine Learning. My passion lies in sharing insights and discoveries in the fields of AI, Machine Learning, Artificial General Intelligence, and Robotics through engaging blog posts. Proficient in Python, ML libraries, and algorithms, I find joy in developing and deploying ML models, with a focus on leveraging AWS Sagemaker. Join me on this exciting journey of unraveling the mysteries of AI through the lens of coding, exploration, and the ever-evolving landscape of machine learning. Let's embark on this knowledge-sharing adventure together!

It is an optimization algorithm used to minimize the cost function in various machine learning algorithms.

  • The idea is to find the optimal 'm' value.

  • Initially, we choose a random value of m and c. Afterward, increase or decrease the m and c values by moving toward the optimal m.

  • The increasing or decreasing value of m and c will depend on the positive and negative slope.

  • Equations:

    new m, m' = m - slope_m

    new c, c' = c - slope_c

    In these cases, sometimes the slope can be very high or very low, thus will lead to an overshoot in the case of m and c values.

    Therefore, we use a constant known as learning rate (α).

  • Final equations:

  • Stopping condition for gradient descent -

    1. Run for a maximum number of iterations.

    2. The cost function from the previous iteration and the current iteration doesn't change much.

Learning rate (α):-

It is a constant that controls the rate at which we are moving (to find the optimal m).

Selecting lr is a crucial part -

  1. Case 1 - α is too high -> overshoot

  2. Case 2 - α is too low -> moves very slowly, may not even reach the optimal m.

  3. Case 3 - Adaptive α -> initially selecting a high value and then reducing it on the go.

More from this blog

Untitled Publication

31 posts