Low-Frequency Discrete Cosine Transforms Bypass the Avalanche Effect: How Perceptual Hashes Detect Re-Uploaded Images Across Social Platforms
Unlike cryptographic hashes that change entirely when a single pixel is altered, perceptual hashes use low-frequency mathematical transforms to generate similar outputs for visually similar images. This intentional bypass of the avalanche effect allows social media platforms to detect compressed, cropped, or filtered re-uploads of copyrighted or abusive content.
By Wei Zhang
In short
- Cryptographic hashes fail to detect re-uploaded images because minor alterations, like compression, completely change the output.
- Perceptual hashes use the Discrete Cosine Transform to isolate the low-frequency structure of an image, which survives compression and filtering.
- Platforms compare these 64-bit hashes using the Hamming distance, where a difference of fewer than 10 bits indicates a visual match.
When a user uploads a copyrighted photograph to a social media platform, the image rarely arrives in its original state. The platform's servers compress the file, the user might have cropped the edges, and a filter may have altered the color balance. If the platform relied on standard cryptographic hashes like SHA-256 to detect the duplicate, it would fail.[1][3]
Cryptographic hashes are designed to exhibit the "avalanche effect," where changing a single bit of the input file completely randomizes the output string. This is essential for verifying software integrity, but it makes cryptographic hashes useless for identifying images that have been visually preserved but digitally altered.
The Avalanche Effect and Its Discontents
To solve this, platforms use perceptual hashes (pHash), which intentionally bypass the avalanche effect. A perceptual hash generates a fingerprint based on the visual structure of the image, ensuring that two images that look the same to a human will produce similar hash values, even if their underlying file data differs significantly.[2][3]
The core mechanism behind a standard perceptual hash is the Discrete Cosine Transform (DCT). The DCT is a mathematical operation that expresses a finite sequence of data points in terms of a sum of cosine functions oscillating at different frequencies. In image processing, it separates the image into parts of differing importance with respect to the image's visual quality.[2]
Isolating Low-Frequency Structure
When an image is processed for a pHash, it is first reduced to grayscale and shrunk to a small size, typically 32 by 32 pixels. This removes high-frequency details like sharp edges and fine textures, which are easily destroyed by compression or scaling. The DCT is then applied to this simplified 1,024-pixel block.[2][3]
The DCT outputs a matrix of coefficients representing the image's frequencies. The pHash algorithm discards the high-frequency coefficients, keeping only the low-frequency data from the top-left corner of the matrix—usually an 8 by 8 block. These 64 values represent the lowest frequencies in the picture, capturing the broad, structural layout of light and dark areas.[2]
Because compression algorithms like JPEG also discard high-frequency data to save space, the low-frequency structure remains largely untouched across different compression levels. By isolating these low frequencies, the pHash ensures that the resulting fingerprint is robust against common image manipulations.[2][3]
Calculating the Hamming Distance
To generate the final hash, the algorithm calculates the median value of the 64 low-frequency coefficients. It then compares each coefficient to the median, assigning a 1 if it is above the median and a 0 if it is below. This creates a 64-bit integer—the perceptual hash.[2]
When a platform wants to check if a newly uploaded image matches a known file, it computes the pHash of the new image and compares it to the hashes in its database using the Hamming distance. The Hamming distance is simply the number of bit positions in which the two 64-bit hashes differ.[3]
If two images are identical, their Hamming distance is zero. If the distance is small—typically less than 10 bits—the images are considered visually similar, indicating a likely re-upload, crop, or slight color adjustment. This allows automated systems to flag copyrighted material or known abusive content even when the files have been modified.[2][3]
How we did this
- Method
- Compared the mathematical properties of cryptographic hashing against perceptual hashing to explain how the intentional removal of the avalanche effect enables visual similarity detection.
- What we found
- The effectiveness of perceptual hashing in content moderation relies entirely on its mathematical inversion of the core security principle (the avalanche effect) that governs standard cryptographic hashing.
- What we worked from
- Avalanche effect definition: A small change in input completely randomizes output
- Perceptual hash mechanism: DCT isolates low frequencies to maintain similarity — IEEE Xplore
- Limits of this analysis
- This analysis focuses on standard DCT-based pHash algorithms and does not cover newer, machine-learning-based perceptual hashing models.
Terms to know
- Avalanche Effect
- A property of cryptographic hash functions where a small change in the input results in a completely different output.
- Discrete Cosine Transform (DCT)
- A mathematical operation that expresses data points as a sum of cosine functions, used to separate an image into high and low frequencies.
- Hamming Distance
- A metric for comparing two binary strings of equal length, representing the number of positions at which the corresponding bits are different.
Questions readers ask
Can perceptual hashes detect rotated images?
Standard DCT-based perceptual hashes are sensitive to rotation. If an image is rotated by 90 degrees, its low-frequency structure changes, resulting in a completely different hash.
How do platforms handle video content?
Video hashing typically involves extracting keyframes at specific intervals and generating perceptual hashes for those individual frames, which are then compared against a database of known video fingerprints.
Are perceptual hashes reversible?
No. Because the algorithm discards the high-frequency data and reduces the image to a 64-bit string, it is impossible to reconstruct the original image from its perceptual hash.
Different angles
Platform Moderation Teams
Rely on perceptual hashing to automate the detection of copyrighted material and known abusive content at scale.
For social media platforms processing millions of uploads daily, manual review of every image is impossible. Moderation teams rely on perceptual hashing to automatically flag known abusive imagery, such as child sexual abuse material (CSAM), or to enforce copyright claims. By maintaining databases of hashes for known violating content, platforms can instantly block re-uploads, even if the user has attempted to evade detection by slightly cropping or compressing the file. The efficiency of the Hamming distance calculation allows this screening to happen in milliseconds during the upload process.
Cryptographers
Emphasize the distinction between cryptographic hashes, which guarantee data integrity, and perceptual hashes, which prioritize visual similarity.
Cryptographers view the avalanche effect as a fundamental requirement for data security. If a hash function does not exhibit the avalanche effect, it cannot be used to verify that a file has not been tampered with or to securely store passwords. From a cryptographic perspective, perceptual hashes are not true hash functions in the security sense; they are feature extraction algorithms. Cryptographers stress that while pHashes are highly effective for content moderation, they offer no cryptographic security and are vulnerable to adversarial attacks, where an image is subtly altered specifically to change its perceptual hash without changing its visual appearance to a human.
- Platform Moderation Teams
- Rely on perceptual hashing to automate the detection of copyrighted material and known abusive content at scale, reducing the manual review burden.
- Cryptographers
- Emphasize the distinction between cryptographic hashes, which guarantee data integrity through the avalanche effect, and perceptual hashes, which prioritize visual similarity.
- Image Processing Researchers
- Focus on improving the robustness of perceptual hashes against more complex adversarial attacks, such as geometric rotations or heavy cropping.
Perspectives this story doesn't cover
- Privacy advocates concerned about the potential for perceptual hashing to be used for mass surveillance or tracking user uploads across platforms.
Sources
[1]Factlen Editorial TeamPlatform Moderation TeamsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
[2]IEEE XploreImage Processing ResearchersA robust perceptual image hash based on discrete cosine transform
Read on IEEE Xplore →
[3]arXivPlatform Moderation TeamsA Comprehensive Review on Perceptual Image Hashing
Read on arXiv →
More in Technology
See all →Digital Services Act
US Justice Department Intervenes in X's Appeal Against EU Digital Services Act Fine
7 sources
Platform Regulation
YouTube Rejects Meta's $18 Billion Teen Safety Settlement Framework
4 sources
Digital Regulation
EU 'Kids Act' Proposes Bloc-Wide Ban on Social Media for Under-13s and Limits Teen Access
6 sources
Network Science
Why Social Media Algorithms That Prioritize Close Friends Strangle Information Diffusion
8 sources
Comments
Every angle. Every day.
Get Technology stories with full source coverage and perspective breakdowns, free every day.




