[
    {
        "id": "osp-16982",
        "type": "article-journal",
        "title": "How Inefficient Is Natural Gradient Descent? From Exact Optimality to Θ(\\sqrt{ \\log d }) Divergence",
        "author": [
            {
                "family": "Sharon",
                "given": "Guni"
            },
            {
                "family": "Kuhnle",
                "given": "Alan"
            }
        ],
        "URL": "https://omanscience.com/en/articles/how-inefficient-is-natural-gradient-descent-from-exact-optimality-to-sqrt-log-d-divergence",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this overhead by the inefficiency ratio \\(R \\ge 1\\), the Fisher length of the mixture geodesic divided by the Fisher--Rao distance, and bound its supremum over endpoint pairs as a function of the parameter dimension \\(d\\). A tensor criterion identifies the regime (I) families, with \\(R=1\\) everywhere: exactly those with quadratic potential or dimension one, such as fixed-covariance Gaussians. For non-quadratic families, we prove two further regimes: (II) bounded third-order skewness plus finite Fisher--Rao diameter yields a dimension-independent bound; and (III) for products of scale families---including Gaussian covariances and Gamma rates---\\(R\\) grows as \\(Θ(\\sqrt{\\log d})\\), unbounded in \\(d\\). Under a per-step Fisher-chord budget, \\(R\\) translates to a practical computational cost: NGD requires asymptotically at least \\(R\\) times as many steps as an optimizer following the Fisher--Rao geodesic. Experiments confirm all three regimes: \\(R=1\\) to machine precision for quadratic-potential families (I), the categorical bound \\(π/(2\\sqrt{2})\\) is approached but not attained (II), and sampled scale-product \\(R\\) grows with \\(d\\), reaching \\(R \\approx 1.5\\) for long, high-dimensional moves (III)."
    }
]