Skip to main content
Database ArchitecturePostgreSQL· 8 min read· in Technology

How Sequential Timestamps in UUIDv7 Prevent B-Tree Page Splits and Buffer Bloat

The UUIDv7 standard replaces random identifiers with time-ordered keys, eliminating the index fragmentation that degrades database performance. The 48-bit timestamp allows inserts to pack cleanly into B-Tree leaf pages, drastically reducing memory bloat and write-ahead log overhead.

By Sergei Orlov

In August 2026, database engineers migrating a heavily queried PostgreSQL 18.4 cluster swapped their primary keys from random UUIDv4 strings to the new UUIDv7 standard. The result was a 23-fold increase in multi-row insert performance. The massive speedup did not come from changing the database engine, but from rearranging the bits inside the identifier.[4]

For years, developers have relied on Universally Unique Identifiers to generate database keys across distributed systems without central coordination. The dominant standard, UUIDv4, provides 122 bits of pure randomness, ensuring that a sample of 32 quadrillion values has a 99.99 percent chance of containing zero duplicates. However, that exact randomness acts as a silent performance killer inside relational databases.[6]

Because UUIDv4 values lack sequential order, they wreak havoc on the B-Tree indexes that databases like PostgreSQL and MySQL use to organize data. The chaotic insertion pattern forces the engine to constantly split index pages, inflating storage requirements. The Internet Engineering Task Force solved this in May 2024 with RFC 9562, introducing a time-ordered alternative.[2][6]

The anatomy of a B-Tree page split

To understand why random identifiers degrade performance, one must look at how a relational database stores an index. PostgreSQL organizes primary keys into a B-Tree, a hierarchical structure where the bottom layer consists of 8-kilobyte leaf pages. Each leaf page holds an ordered array of key values and pointers to the actual table rows.[3]

When an application inserts a sequential key, such as a traditional 64-bit integer, the database engine knows exactly where it belongs. Because every new value is mathematically larger than the last, it always lands on the rightmost leaf page of the B-Tree. That single active page stays pinned in the server's shared memory, absorbing thousands of inserts before filling up.[2][3]

A random UUIDv4 does the exact opposite. Every generated value is equally likely to sort anywhere in the entire index, meaning every insert dives into a different, unpredictable leaf page. The database engine must constantly fetch cold pages from the filesystem cache or the physical disk just to write a single row.[2]

Random inserts force the database to split full pages, leaving the index fragmented and bloated.

The structural damage occurs when a random insert targets a leaf page that is already full. Because the new key must be placed in strict sorted order, the database cannot simply append it. Instead, PostgreSQL must perform a page split, moving roughly half of the existing entries to a brand new 8-kilobyte page to make room.[1][3]

"Random inserts split pages down the middle all over the index, leaving them half-full, a larger index that holds the same rows," notes database architecture blog DBGorilla. This mechanical reality dictates how the storage engine behaves under load. Over time, a B-Tree fed with random UUIDs settles at an average leaf density of just 69 percent.[1][5]

This fragmentation means a UUIDv4 index consumes roughly twice as much disk space as a sequential integer index holding the exact same number of rows. The database must traverse a physically larger tree to find any given record. This bloat measurably slows down point lookups and range scans across the entire application.[1][6]

The buffer bloat cascade

The wasted space on disk is only the beginning of the performance degradation. The more severe consequence is buffer bloat, which chokes the database's available random-access memory. PostgreSQL relies on a designated memory area, called shared buffers, to keep frequently accessed index pages ready for immediate reads and writes.[1][3]

With a sequential key, the hot working set is tiny, requiring only a handful of rightmost pages to remain cached. With random UUIDv4 keys, the hot working set becomes the entire index, because the very next insert could target any page. Once the fragmented index outgrows the allocated shared buffers, cache hit rates collapse.[1][2]

The database is then forced into a cycle of cache thrashing, evicting useful pages to load cold ones for a single insert, only to need the evicted pages moments later. This constant shuffling burns CPU cycles and saturates the storage hardware with random input and output operations. The entire server slows down to accommodate the churn.[6]

UUIDv4 generates exponentially more write-ahead log volume by dirtying pages across the entire index.

The chaos also multiplies the write-ahead log volume, which PostgreSQL uses to ensure crash recovery. The engine's safety mechanics dictate that the first modification to any page after a system checkpoint must log the entire 8-kilobyte block. Sequential inserts dirty only a few pages, but random inserts dirty pages across the entire tree.[1][5]

"UUIDv4 keys scatter across the index, cause 50/50 page splits, and bloat WAL," confirms database optimization firm MyDBA.dev. This write amplification forces the database to write gigabytes of redundant page images to the log. The overhead accelerates disk wear and bogs down replication to standby servers.[3]

On high-throughput systems, this logging overhead alone can become the primary bottleneck for the entire application cluster. The storage controller struggles to keep up with the massive influx of page images, causing insert latency to spike unpredictably. The database spends more time managing its own safety mechanisms than actually storing user data.[1]

The evolution from version 1 to version 7

The concept of a time-based UUID is not entirely new. The original UUIDv1, designed in the 1990s, also incorporated a timestamp to ensure uniqueness. However, it placed the time fields in reverse order and embedded the generating machine's physical MAC address, creating severe privacy vulnerabilities that led the industry to abandon it.[4]

UUIDv4 solved the privacy issues by relying entirely on random number generators, but it sacrificed all mechanical sympathy with database storage. For years, engineering teams attempted to bridge the gap by writing custom identifier generators. These bespoke sequential extensions fragmented the ecosystem and required custom database plugins to function correctly.[1][6]

RFC 9562 unifies these disparate efforts into a single, language-agnostic standard. The specification replaces the pure randomness of version 4 with a structured, time-ordered layout. The total length remains 128 bits, requiring no changes to the underlying database column types.[2][5]

The UUIDv7 standard places a 48-bit timestamp in the most significant bits, ensuring chronological sorting.

The critical innovation lies in the first 48 bits, which encode a Unix timestamp in milliseconds since the epoch. Because the timestamp occupies the most significant bits, any UUIDv7 generated later in time will automatically sort after one generated earlier. This guarantees that standard lexicographical sorting will arrange the keys in chronological order.[2][6]

How the timestamp fixes the tree

"New keys land at the right edge of the tree, exactly where the ordered BIGINT used to put them," explains engineering outlet pgEdge. "That one change restores everything UUID v4 broke: sequential locality, the hot working-set, and clean page splits." The database engine can finally append records efficiently again.[2]

Following the timestamp, a four-bit version field identifies the format, leaving 74 bits of cryptographically secure randomness. This random tail ensures that if dozens of application servers generate identifiers during the exact same millisecond, the resulting keys remain globally unique. The system achieves this without requiring any network coordination between nodes.[2]

Following the timestamp, a four-bit version field identifies the format, leaving 74 bits of cryptographically secure randomness.

The 74 bits of random data provide a massive cryptographic space, yielding over 18 sextillion possible combinations per millisecond. This ensures that even if a global fleet of servers generates millions of keys simultaneously, the probability of a collision remains infinitesimally small. The format retains the decentralized generation benefits of its predecessor.[2][6]

Furthermore, the RFC permits implementations to use 12 of those random bits as a dedicated monotonic counter. If a single machine generates multiple keys within the exact same millisecond, the counter increments sequentially. This optional construct guarantees strict ordering even at microsecond resolutions, keeping the B-Tree perfectly packed.[2]

Because the values are time-clustered, inserts naturally target the rightmost pages of the B-Tree. The leaf pages fill up completely and split cleanly, pushing the index fill factor back up to between 90 and 95 percent. The working set shrinks back to a manageable size, completely eliminating the buffer bloat.[5]

Insert performance on UUIDv4 degrades as the index outgrows available memory, while UUIDv7 remains stable.

The reduction in touched pages also slashes the write-ahead log overhead. Benchmarks published by database engineers show that switching to time-ordered UUIDs can reduce full-page image writes to the log by up to 99.97 percent during heavy bulk loads. The storage hardware is finally freed to focus on actual throughput.[5]

The migration to PostgreSQL 18

The database industry is rapidly standardizing around the new format. PostgreSQL 18, entering broad deployment in late 2026, ships with a native generation function built directly into the core engine. This allows developers to define time-ordered primary keys natively, without relying on third-party extensions or custom SQL functions.[1][3]

For applications using older database versions, the transition is still straightforward. Because UUIDv7 is an application-level standard, client libraries in Java, Python, and Go can generate the time-ordered identifiers before sending the insert command. The database simply stores the 16-byte binary value, completely unaware of how it was minted.[1]

Migrating an existing table from UUIDv4 to UUIDv7 does not require a costly data type rewrite. Both versions share the exact same binary footprint, meaning no table locks are required to alter the column. Developers can simply update the default generation function for new rows, allowing the index to gradually stabilize.[1][5]

However, the legacy UUIDv4 values will remain scattered throughout the older sections of the B-Tree. To reclaim the wasted space and eliminate the historical fragmentation, database administrators must eventually rebuild the index. This is typically accomplished using a concurrent reindex command during a scheduled maintenance window.[1]

While developer tooling companies market UUIDv7 as a magical performance cure-all, the reality is strictly mechanical. The format does not make the database engine inherently faster; it simply stops the application from actively sabotaging the memory management. The 16-byte footprint still remains twice as large as a traditional 8-byte integer.[1][3]

Illustration: By keeping the active index pages in memory, UUIDv7 frees storage hardware to handle actual application throughput.

For systems that do not require decentralized key generation, traditional sequential integers remain the most efficient choice. But for distributed architectures, UUIDv7 offers a near-perfect compromise. By aligning the structure of the identifier with the physical realities of B-Tree storage, the standard eliminates a decade-old performance trap.[3][6]

Key points

  • UUIDv4's pure randomness forces database B-Trees to constantly split pages, leading to severe index fragmentation and memory bloat.
  • RFC 9562 introduces UUIDv7, which embeds a 48-bit Unix timestamp to ensure new identifiers sort chronologically and append cleanly to the index.
  • Migrating to time-ordered identifiers can reduce write-ahead log overhead by up to 99.97 percent and increase insert performance without altering database column types.
Storage Optimization 50%Systems Architecture 50%
Storage Optimization
Advocates for mechanical sympathy between data structures and storage hardware.
Systems Architecture
Focuses on the practical implementation of decentralized identifiers in microservices.

Perspectives this story doesn't cover

  • Legacy System Maintainers

Sources

Source coverage

6 outlets

2 viewpoints surfaced

Storage Optimization 50%Systems Architecture 50%
  1. [1]DBGorillaStorage Optimization

    The fourth, subtler effect: B-tree page splits

    Read on DBGorilla →
  2. [2]pgEdgeSystems Architecture

    UUID version 7 comes in (as standardized in RFC 9562)

    Read on pgEdge →
  3. [3]MyDBA.devStorage Optimization

    UUID vs BIGINT Primary Key in Postgres: Which Wins?

    Read on MyDBA.dev →
  4. [4]Daily.devSystems Architecture

    A production migration from UUID v1/v4 to UUID v7 primary keys

    Read on Daily.dev →
  5. [5]MediumStorage Optimization

    Time clustering keeps recent values in the same region of the index

    Read on Medium →
  6. [6]Dev.toSystems Architecture

    Benchmarking UUIDv4 and UUIDv7 in PostgreSQL

    Read on Dev.to →

Comments

Stay informed

Every angle. Every day.

Get Technology stories with full source coverage and perspective breakdowns, free every day.