Decentralized SGD with Controlled Disagreement Finds Flatter Minima
- lab arXiv
- lab arXivLabs
- location Feb
- location Jun
- location Mon
- location Tue
- location UTC
- person Zesen Wang
A new preprint challenges the long-held view that decentralized machine-learning training is inferior to centralized methods, showing that a technique called Decentralized SGD with Adaptive Consensus (DSGD-AC) can find flatter minima and achieve higher test accuracy than both standard decentralized and centralized approaches [1][2]. The paper, titled "Decentralized SGD with Controlled Disagreement Finds Flatter Minima," was submitted to the arXiv preprint repository on 2 February 2026 and revised on 23 June 2026 [1]. The authors, including Zesen Wang, introduce DSGD-AC, which uses a time-dependent scaling mechanism to maintain consensus errors between workers throughout training [1][2]. The work argues that these errors, typically seen as undermining convergence, act as a useful implicit regularizer [2]. The core mechanism of DSGD-AC changes the stationary variance of disagreement modes by balancing two effects. It preserves consensus-error magnitude through weaker graph damping while allowing curvature-dependent damping to shape the disagreement directions [2]. This balance can produce a stronger Hessian-weighted loss-envelope penalty around the deployed model, even when normalized Hessian alignment is weaker than in standard decentralized SGD [2]. Empirical results on image classification tasks showed that DSGD-AC reached flatter solutions and higher test accuracy than standard decentralized SGD and even centralized SGD [1][2]. The findings open a new perspective on the design of decentralized learning algorithms, suggesting that controlled disagreement can be beneficial rather than detrimental [2]. The paper appears on arXiv, an open-access repository for electronic preprints that is moderated but not peer-reviewed [6]. As of November 2024, arXiv receives about 24,000 submissions per month and hosts over two million articles [6]. The platform also features arXivLabs, a framework for community-developed tools that appear on article pages, though new proposals for such projects are currently paused while the development team modernizes arXiv's infrastructure [3][4].
research-papersafety-researchcommentary
Background sources we checked (7)
- arxiv.org ↗ Decentralized training is often regarded as inferior to centralized training because the consensus errors between workers are thought to undermine convergence and generalization. This work challenges this view by introducing decentralized SGD with Adaptive Consensus (DSGD-AC), wh…
- info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
- blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
- info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
- en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
- en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
- en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…
Sources
- export.arxiv.org — Decentralized SGD with Controlled Disagreement Finds Flatter Minima ↗