Transformer Geometry Observatory TGO-I: Spectral Geometry Observatory

83d ago · Global · primary source: export.arxiv.org

A new analytical framework called the Transformer Geometry Observatory (TGO) has been introduced to systematically probe the internal representational geometry of Vision Transformers, an area its creators describe as relatively underexplored compared to the models' widespread application success [1]. The framework's first installment, designated TGO-I, concentrates on the spectral geometry of Vision Transformer representations [1]. Researchers deployed a ViT-Small/16 architecture trained on the ImageNet-100 dataset to track a suite of metrics throughout the training process, including Effective Rank, Stable Rank, Participation Ratio, Spectral Entropy, Spectral Flatness, and Spectral Anisotropy [2]. The analysis also examined covariance structure, eigenspectra, and singular value spectra across all layers and training epochs [3]. The results documented a consistent increase in dimensional utilization as training progressed, accompanied by decreasing anisotropy, increasing spectral entropy, increasing participation ratio, and progressively flatter eigenspectra [1]. Covariance matrices became increasingly diagonal, indicating progressively weaker feature correlations [4]. This pattern runs counter to the common intuition that training should concentrate information into a small number of dominant directions; instead, the team observed a progressive redistribution of variance across representational dimensions [2]. The effect was most pronounced in the final CLS token representation, which consistently exhibited the highest effective dimensionality and lowest anisotropy within the network [1]. In contrast, the Patch Embedding and Positional Embedding layers remained comparatively stable throughout training, suggesting that the majority of geometric evolution occurs within the Transformer processing blocks and the final global representation [5]. The research was posted on the arXiv preprint server, an open-access repository that hosts scientific papers in fields including computer science and has grown to a submission rate of about 24,000 articles per month as of late 2024 [9]. The TGO framework is presented as a surgical set of experiments and analysis pipelines built to analyze the representational geometry and dynamics of Vision Transformers, with TGO-I serving as the inaugural observatory within the broader project [3].

tool-releaseresearch-papercommentary

Background sources we checked (10)
  • arxiv.org ↗ Despite the widespread adoption of Vision Transformers (ViTs) and their success across numerous computer vision applications, the fundamental understanding of their dimensional and representational geometry remains relatively underexplored. To address this gap, we introduce Trans…
  • arxiv.org ↗ # TGO-I: Spectral Geometry Observatory ... Despite the widespread adoption of ViT and its utilization in multiple computer vision applications, the fundamental understanding of their dimensional and representational geometry remains less explored as compared to their application …
  • arxiv.org ↗ # TGO-I: Spectral Geometry Observatory ... Despite the widespread adoption of ViT and its utilization in multiple computer vision applications, the fundamental understanding of their dimensional and representational geometry remains less explored as compared to their application …
  • arxiv.org ↗ # TGO-I: Spectral Geometry Observatory ... Despite the widespread adoption of ViT and its utilization in multiple computer vision applications, the fundamental understanding of their dimensional and representational geometry remains less explored as compared to their application …
  • info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
  • blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
  • info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
  • en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
  • en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…

Sources

Spot something wrong? Report an issue