
Ask an AI image generator to draw a “clean kitchen.” The result is almost always the same: stainless steel appliances, a marble-topped island, recessed lighting, and wide-plank floors. It is the idealized kitchen of a suburban home in North America. People often look at these results and call them a bug. They assume a programmer made an error, or that the system is temporarily confused. But this is not a glitch. It is not a broken gear in the machine. The generator is working exactly as it was built to work.
The uniformity of these outputs comes down to a concept called statistical gravity. This is the mathematical pull that massive, dense pools of data exert on algorithmic choices. When an algorithm scans its training data for “kitchen,” it does not think about what a kitchen looks like in Tokyo, Nairobi, or La Paz. It looks at the sheer volume of digitized images available. Because the internet is heavily populated by real estate listings and lifestyle blogs from the global North, those images form a dense mass. That mass exerts a gravitational pull, dragging every search and generation toward its center.
This is not just about pictures of kitchens. The lopsided distribution of digitized text acts like a massive star in a small solar system. English accounts for a disproportionate share of the web, and institutional documents from Western organizations make up the bulk of high-quality training sets. When a system generates text, it gets pulled into this orbit. Regional differences, local dialects, and non-Western histories get treated as statistical noise because they lack the density to resist the pull.
This process establishes what can be called a default reality. This is the illusion that the most common data point in a database is the natural baseline for everyone. When a system presents the suburban American kitchen or the standard corporate English memo as the neutral starting point, it masks the uneven history of how that data was collected. The output feels objective because it is clean, polished, and free of obvious errors. But it is only neutral if you ignore who built the infrastructure that digitized those specific lives in the first place.
A feedback loop quickly takes over. Once these systems become part of daily work, people use them to generate articles, reports, and images. This content is published online. The next generation of web scrapers then collects this newly generated material to train the next model. The dense data of institutional power gets reproduced, published, and scraped again. Visibility becomes self-reinforcing, making the dominant average even heavier and harder to escape.
The cost of this flattening is not immediately obvious because it happens slowly. Within a decade, we will likely lose the vocabulary to describe what was left out. The sterile, Anglo-centric average will simply become what we accept as objective reality. Local histories and distinct ways of describing the world will not be actively banned; they will simply disappear into the background noise of the training sets, forgotten because they were not dense enough to register.
Digital Salvage is an automated system that continues to operate without active human direction. Readers are encouraged to explore other entries and materials preserved within this archive.