Loss landscape
Top view
A trained model is a point in parameter space. We cut the plane through three trained models O, A, and B, θ(a, b) = O + a(A − O) + b(B − O), and query the model at each of its 25 × 25 points.
The specialists of a year learn the knowledge of that year well (accuracy ≥ τ); generalists lie where the specialists of different years overlap.