Atlas · skill

Calculus for Machine Learning

Calculus for machine learning explains how a model's output and loss change when its inputs or parameters change. The competence connects derivatives, gradients and the chain rule to optimization, so practitioners can understand training behavior, build differentiable objectives and diagnose incorrect or unstable updates.

conceptCalculus

What it is

A derivative describes local change; a gradient collects partial derivatives of a scalar function with respect to several variables. Machine learning combines many such functions into a computational graph. Applying the chain rule through that graph makes it possible to calculate how each parameter contributes to the final loss. Calculus also describes curvature, local approximations and the distinction between a stationary point and a minimum. The practical scope includes multivariable differentiation, vector-valued transformations and reasoning about smoothness. Automatic differentiation performs the bookkeeping, but does not choose a suitable objective or establish that the result represents the intended problem.

What the work involves

A practitioner should be able to derive the gradient of a simple objective, follow how tensor operations compose and identify where a transformation prevents gradient flow. Important decisions include reduction over examples, treatment of constants and the scale of competing loss terms. For a custom operation, compare analytical or automatic gradients with a small numerical check away from discontinuities. The result is an objective whose updates can be explained, together with evidence that the implementation differentiates the intended quantity rather than an accidental reshaping or detached intermediate.

Illustrative example

Consider an illustrative regression model that predicts delivery duration. Its loss combines prediction error with a penalty on large coefficients. Increasing the penalty changes the gradient as well as the loss value, so the analyst works through both terms before selecting an optimizer. A small synthetic dataset provides a check: perturb one coefficient, compare the observed loss change with the gradient's prediction and verify that an update in the opposite direction lowers the loss locally.

Limits and common mistakes

A gradient gives local information and does not guarantee a globally best model. Numerical differences can be unreliable with poorly chosen step sizes, floating-point noise or nonsmooth operations. Functions such as maximum and absolute value need careful treatment at their corners, and discrete choices cannot generally be differentiated directly. Backpropagation is an application of the chain rule, not a separate justification for the objective. Check finite values, gradient magnitude and dependency paths before interpreting a stalled training run as a lack of useful data.

Prerequisites

  • Gradients live in vector spaces alongside linear algebra.

Sources and further reading

Last updated: 2026-10-10