Visualizing high-dimensional optimization terrain
A loss landscape is a slice through the high-dimensional parameter space, showing loss values along two random directions from a trained checkpoint. The result is a 2D contour map: valleys are good minima, ridges are bad, and saddle points are tricky local plateaus.
This visualization is not the true landscape (true loss space has millions of dimensions), but it reveals structural insights. A narrow basin with steep walls requires careful learning rates; a wide plateau is hard to optimize out of.
What landscapes reveal about optimization difficulty
Neural networks often converge to saddle points, not true minima. Saddles look locally flat in some directions but have curvature (negative or positive Hessian) in others. Optimization algorithms like SGD with momentum naturally escape saddles because their inertia carries them through flat regions.
Landscapes also vary by architecture: wider networks tend to have flatter minima and more connected basins, while narrow networks have sharper, more isolated minima. Flatter minima often generalize better because they're less sensitive to weight perturbations.
Practical implications for training
Understanding the landscape explains why batch size, learning rate, and initialization matter. A large batch gives high-variance gradients near a saddle, helping escape; a small batch provides smoothing. The landscape also shows why some initialization schemes work better: they start closer to well-connected basins rather than isolated islands.