Convexity
The property that a function is a single bowl with no false bottoms — the case where gradient descent is guaranteed to find the one true minimum.
Convexity
A function is convex if it curves the same way everywhere — like a bowl. Pin down two points on its graph, draw the straight chord between them, and the function never pokes above that chord. Formally, for any two inputs a, b and any blend t \in [0,1],
This one inequality has an enormous consequence: a convex function has no local minima other than the global one. There is a single bottom, and every downhill direction eventually leads there. For an optimizer this is paradise — gradient descent cannot get stuck, cannot be fooled by a false valley, and (with a sensible step size) is guaranteed to converge to the best possible answer. This is why a whole field, convex optimization, exists: when you can phrase a problem convexly, you can essentially declare it solved.
One bowl versus many
Contrast the two functions below. The smooth parabola is convex: a single basin, and a ball released anywhere slides to the same point. The wiggly one is non-convex: same overall upward sweep, but riddled with local dips where descent can come to rest short of the true minimum. Real learning problems — especially neural networks — live in the second picture.
When the world is not convex
Most interesting models are non-convex, and that is the source of nearly every training headache: sensitivity to initialization, the need for momentum and random restarts, and the gnawing doubt about whether you found the answer or just an answer. The practical wisdom is encouraging, though — in very high dimensions the many minima of a large network tend to be nearly as good as one another, so descent usually lands somewhere fine even without convexity's guarantee.