One Build Server's Clock Drift Caused Three Teams to Cache Invalid Artifacts

Jul 18, 2026 By Deepa Iyer

At 03:14 UTC on a Tuesday in late March, the on-call engineer at Wavelength received a page. Three separate CI pipelines had been failing for roughly six hours with test errors referencing stale dependencies. The failures looked unrelated at first—Team A's integration tests were pulling a library version that didn't match their source, Team B's build was producing artifacts with timestamps in the future, and Team C's downstream consumers were seeing mismatched method signatures. By dawn, the engineering director had declared an incident. The root cause, uncovered after twelve hours of digging, was a single build server whose clock had drifted 47 milliseconds relative to NTP. That 47-millisecond rift had broken three CI pipelines, poisoned artifact caches, and cost roughly twenty engineer-hours of debugging. The story of how a sub-second timing error cascaded through a distributed build system is a cautionary tale about the hidden assumptions in our infrastructure.

The 47-Millisecond Rift That Broke Three CI Pipelines

The build server in question lived in rack 7 of Wavelength's primary colocation facility. It was a standard-issue machine running Ubuntu 22.04, tasked with compiling microservice artifacts for the company's core platform. Its clock had drifted by 47 milliseconds relative to the NTP pool, a skew that was invisible to the monitoring stack because no alert fired for offsets under 100 milliseconds.

When the server's clock drifted forward, any artifact it produced received a timestamp roughly 47 milliseconds ahead of the true time. The artifact cache—a shared Redis-backed store keyed on a combination of content hash and modification time (mtime)—interpreted the future timestamp as a newer version. Subsequent builds on other nodes saw the artifact as fresher than the actual latest build, so they pulled the stale artifact instead of rebuilding.

Team A's CI pushed an artifact with the future timestamp at roughly 21:00 UTC. Team B's build pulled that artifact at 21:15, treating it as newer than its declared dependency. Team C's integration test, running at 22:00, consumed a mismatched library that caused a cascade of method-not-found errors. Each team spent hours blaming their own code, running git bisect, and restarting builds. The pattern only emerged when Sasha Chen, a lead engineer, correlated the build logs across all three teams and noticed that every affected build had passed through the same build cluster rack.

How Clock Drift Escapes Monitoring at Wavelength

Wavelength's observability stack was built to catch large anomalies. CPU spikes, memory leaks, disk I/O bottlenecks—those triggered alerts within minutes. But sub-second clock drift was a blind spot. The NTP daemon on each build node was configured to sync every 3600 seconds, a typical setting for servers where time precision wasn't considered critical. The monitoring system checked for clock skew but only raised an alert when the offset exceeded 100 milliseconds.

The artifact cache had a more subtle vulnerability. It keyed entries on a hash of the artifact's content combined with the file's mtime. The assumption was that mtime would always be monotonic—that a newer build would have a later mtime. But when a node's clock drifted forward, an artifact from an older build could appear newer than a subsequent build on a different node. The hash collision was improbable; the mtime logic broke ordering.

The design made sense at the time. Content-addressed caches are standard practice, and mtime is a cheap way to avoid storing redundant metadata. But the combination assumed clock synchronization across all build nodes, an assumption that was never explicitly documented or tested. As one engineer later noted in the post-mortem, "We had no test that said 'what if this machine's clock is wrong by 50 milliseconds?'"

The monitoring gap was not unique to Wavelength. A survey of post-mortems from other companies reveals that sub-second clock drift is a recurring failure mode in distributed build systems. At a 2023 SREcon talk, an engineer from a major cloud provider described a similar incident where 30 milliseconds of drift caused a cascade of cache invalidations across a fleet of 200 build agents. The talk's key takeaway: most teams configure NTP once and forget about it.

Cache Invalidation: The Silent Poisoning Cascade

The cascade began with Team A's build server in rack 7. At 20:58 UTC, the server's clock was 47 milliseconds ahead of the true time. A developer pushed a commit to a shared library repository. The CI pipeline on that server compiled the library and wrote the artifact to the cache with a timestamp of 20:58:00.047. The correct time was 20:57:59.000. The artifact was now stamped 1.047 seconds into the future relative to the NTP-synchronized nodes.

At 21:15, Team B's build server—on a different rack with a properly synchronized clock—requested the latest version of that library. The cache returned the artifact with the future timestamp, which was greater than the timestamp of any other version. Team B's build pulled it without recompilation, assuming it was the freshest. But the artifact had been built from a slightly older commit—a race condition in the CI pipeline's checkout step had pulled the previous commit instead of the latest.

Team C's integration test, running at 22:00, consumed a service that depended on the library. The service had been built by Team B's pipeline and included the stale library. When the integration test called a method that had been renamed in the latest source, it received a NoMethodError. The test failed. The on-call engineer for Team C spent two hours examining their own code, convinced the bug was a recent change. It was not.

The cache eviction policy made things worse. The LRU (least recently used) algorithm kept stale entries because they were being accessed by the failing builds. Each failed build read the stale artifact, updating its access time, so the cache never evicted it. The system was effectively poisoning itself with its own failure traffic.

This kind of cascading failure is a classic pattern in distributed systems. A single subtle fault—here, a 47-millisecond clock drift—produces a symptom that looks like a code bug. Each team isolated themselves, assuming the problem was local. Without centralized log correlation, the pattern might have taken even longer to spot. The incident underscores the need for cross-team observability: dashboards that show artifact timestamps across all caches, alerts for non-monotonic mtime sequences, and a shared understanding of how the build pipeline depends on time.

Detective Work: Tracing the Drift to a Single Rack

Sasha Chen, a lead infrastructure engineer, was paged at 03:14 UTC after the third team reported a similar failure pattern. He started by collecting build logs from all three teams over the preceding 12 hours. The logs showed that every failed build had at least one artifact with a timestamp that was slightly ahead of the build's start time. The timestamps weren't wildly off—just 40 to 50 milliseconds in the future—but they were consistently ahead.

Chen then mapped the build nodes used by each team. All three teams had used nodes in rack 7 for the affected builds. Rack 7 contained 12 build servers, but only one—node 7-4—had the clock drift. Chen confirmed this by running an NTP query on each node. Node 7-4 reported an offset of +47 milliseconds. The other nodes in the rack were within 2 milliseconds of the NTP pool.

The root cause was a failing crystal oscillator on the node's baseboard management controller (BMC). The BMC's clock had been drifting for roughly two weeks, but the drift had been gradual—a few milliseconds per day—until it crossed the 40-millisecond threshold. The NTP daemon on the host OS was supposed to correct for BMC drift, but it was configured to sync only once per hour. Between syncs, the drift accumulated.

Chen replaced the BMC and the clock stabilized within 2 milliseconds of NTP within 10 minutes. He then invalidated all cached artifacts from node 7-4 and triggered clean rebuilds for the three teams. The pipelines passed within an hour. The total time from first failure to resolution was roughly 18 hours.

The detective work highlights a broader lesson: in distributed systems, correlation often precedes causation. Chen didn't need to understand the oscillator failure immediately; he only needed to see that artifacts from one rack had skewed timestamps. That pattern—consistent timestamp offset across multiple teams—was the key signal. Many post-mortems emphasize that the hardest part is not fixing the bug but finding the pattern in the noise.

Remediation: From Clock Monitoring to Cache Design

The post-mortem produced four concrete changes. First, Wavelength deployed chrony on all build nodes with a sync interval of 64 seconds, keeping clock skew below 10 milliseconds. Second, they added a monitoring alert for any clock skew exceeding 5 milliseconds, with a pager duty escalation for skew over 50 milliseconds. Third, they redesigned the artifact cache key to include a monotonically increasing build ID generated by the CI orchestrator, removing the mtime from the key entirely. The build ID was a simple integer counter distributed via a Redis atomic increment, ensuring total ordering regardless of clock skew.

Fourth, they added a cache TTL enforced by wall-clock time, not mtime. Each cached artifact now had an expiration timestamp set by the cache server's own clock. If the artifact's wall-clock TTL expired, the cache would treat it as a miss and force a rebuild, even if the mtime appeared recent. This was a safety net against future clock anomalies.

The company also published internal guidelines for cache hygiene: always use monotonic identifiers when ordering matters, never rely on timestamps from untrusted sources, and test cache behavior under clock skew conditions. The guidelines were adopted by at least five other teams within a month.

Not everyone agreed on the scope of the fixes. Some engineers argued that the cache redesign was overkill—that tightening NTP sync and adding alerts would prevent a recurrence. Others pointed out that similar incidents had happened at other companies and that a monotonic build ID was a cheap insurance policy. The debate mirrored a broader tension in distributed systems: whether to invest in stronger time synchronization or to design systems that tolerate clock skew.

This debate is not new. In the early 2000s, the designers of Google's Chubby lock service chose to rely on accurate clocks for lease expiration, arguing that clock skew could be kept small enough through careful monitoring. More recently, Amazon's DynamoDB adopted a fully clock-free approach for its distributed transactions, using logical clocks instead. Wavelength's incident suggests that for build caches—which are less latency-sensitive than lock services—a hybrid approach makes sense: use physical clocks for coarse-grained TTLs and monotonic counters for ordering. The cost of a counter is negligible; the cost of a clock-drift-induced outage is not.

Another design alternative that was considered but rejected was to make the artifact cache entirely content-addressed, with no timestamps at all. In that approach, the cache key would be a hash of the artifact's contents, and the build system would always rebuild if the hash changed. The downside is that rebuilds become more frequent, increasing build times and resource usage. For Wavelength's scale—roughly 10,000 builds per day—the increase would have been significant, roughly 15–20% more CPU time. The team decided that the monotonic build ID offered a better trade-off: it preserved the performance benefits of content-addressed caching while eliminating the clock-skew vulnerability.

The incident also prompted a review of NTP infrastructure. Wavelength's colocation facility had a single stratum-1 NTP server, which was itself synchronized via GPS. The team added a second stratum-1 server as a backup and configured all build nodes to use both servers, with a failover if the offset between them exceeded 10 milliseconds. This redundancy made it less likely that a single BMC failure could cause widespread clock drift.

The Hidden Danger of Time in Distributed Build Systems

Distributed systems assume clock monotonicity. Build caches, distributed locks, lease mechanisms, and even log ordering all depend on the idea that time moves forward at roughly the same rate on every node. When that assumption fails, the failure modes are subtle and hard to diagnose. A 47-millisecond drift is not a crash. It is a slow poison that manifests as mysterious test failures, intermittent build issues, and teams wasting hours on false leads.

Build caches are especially sensitive because they combine two things that are hard to get right: content addressing and time-based ordering. Content hashing ensures correctness in the face of bit flips, but it does not help with ordering. When the ordering is wrong, the cache returns the wrong artifact, and the build fails in a way that looks like a code bug, not an infrastructure problem.

Most teams overlook NTP configuration in CI. They assume that cloud providers or colo vendors handle time synchronization, or that a once-per-hour sync is good enough. But as build fleets grow and artifact caches become shared, the tolerance for clock skew shrinks. A drift of 50 milliseconds in a single node can poison a cache that serves hundreds of builds per day.

Wavelength's incident mirrors similar outages at other firms. In 2022, a major e-commerce company reported a two-day outage caused by a 200-millisecond clock drift on a build server that invalidated every artifact in its cache. In 2024, a CI vendor's post-mortem described a multi-hour outage where a 30-millisecond drift caused their distributed test scheduler to assign tests to the wrong shards. The pattern is consistent: clock drift is a hidden failure mode that grows more dangerous as systems become more distributed.

The fix is not to eliminate clock drift—that is impossible. The fix is to design systems that expect drift and handle it gracefully. Monotonic identifiers, wall-clock TTLs, and explicit clock monitoring are three tools that every team operating a distributed build system should consider. The alternative is a 47-millisecond rift that breaks three CI pipelines and steals a day of engineering time.

As one of the engineers on the incident later wrote in the post-mortem: "We spent 18 hours debugging a problem that could have been prevented by a 5-line configuration change and a monotonic counter. The lesson is not that time is hard. The lesson is that we forgot we were building a distributed system."

Recommend Posts
Tech

One Build Engineer Trades a Safer Package Registry for a Two-Minute Install Lag

By Sara Park/Jul 18, 2026

A build engineer adopts a signed package registry for security, trading two minutes per install for verifiable provenance. The cost in developer hours and the industry's next steps.
Tech

One Maintainer's Twelve-Hour Firewall Patch Left a TLS Handshake Dead for Three Years

By Deepa Iyer/Jul 18, 2026

A single firewall patch by one OpenSSL maintainer silently broke TLS 1.3 resumption for three years, costing retransmission and developer hours. The story exposes the bus factor and funding gaps in critical infrastructure.
Tech

One Distributed Query’s Storage Layer Bill Exceeded Its Feature Budget by Five Figures

By Lucas Mendes/Jul 18, 2026

How a single distributed join triggered a five-figure cloud bill, and why storage economics must be a first-class query constraint for engineering teams.
Tech

A Distributed Systems Role Pays Less Than Monolith Work at Equivalent Scale

By Lucas Mendes/Jul 18, 2026

Engineers working on distributed systems often earn 10–15% less than peers on monoliths at similar scale. The article examines why and how to navigate the gap.
Tech

One Monorepo’s Shared Schema Enum Forced Thirty Teams Into a Single Error String

By Yusuke Tanaka/Jul 18, 2026

How a single protobuf enum in a monorepo root forced thirty teams to standardize error strings, increased build times, and led to workarounds that defeated schema enforcement. Lessons from Google’s error model and a pragmatic shard fix.
Tech

Operating Cost Drives an LLM Provider's API Price to Ten Times the Inference

By Lucas Mendes/Jul 18, 2026

LLM API prices can exceed inference costs by 10x. This article breaks down the operating expenses, contract lock-ins, and what procurement teams can do about it.
Tech

One Platform Team’s Private API Cost Ten Engineers a Week of Manual Sync

By Yusuke Tanaka/Jul 18, 2026

A platform team's undocumented endpoint forced ten engineers into a week of manual reconciliation. Here's how contract-first development and shared tooling eliminated the waste.
Tech

One Auth Engineer Replaced Eight Vendor SDKs With a Single LDAP Config File

By Lucas Mendes/Jul 18, 2026

How one engineer replaced eight authentication SDKs with a single LDAP config, cutting attack surface and maintenance overhead. A deep dive into the trade-offs and operational reality.
Tech

One Maintainer’s Two-Line CSS Fix Cut Load Times by Forty Percent

By Yusuke Tanaka/Jul 18, 2026

A single maintainer cut LCP by 40% with a two-line CSS change. This article breaks down the fix, why modern bundlers miss it, and how to apply it without new tooling.
Tech

One Cloud Database Vendor’s Write Path Locked Nine Clients Into a Single SLA Clock

By Yusuke Tanaka/Jul 18, 2026

How a shared consensus group and single clock source penalize fast writers in multi-tenant databases, and what engineering teams can do about it.
Tech

One Build System’s Config Parser Swallowed Three Teams’ Deployment Scripts

By Yusuke Tanaka/Jul 18, 2026

Monzo's custom TOML parser silently dropped unknown keys for years. When a strict mode update shipped, three teams' deployment scripts broke. A post-mortem reveals the root cause and lessons for build system maintainers.
Tech

One Supply Chain Engineer's Two-Week Patch Audit Found Dormant Signing Keys Across Seven SDKs

By Yusuke Tanaka/Jul 18, 2026

A routine audit by a supply chain engineer uncovered dormant signing keys in seven SDKs, exposing hundreds of apps to potential supply chain attacks. Here's how they did it and what teams can learn.
Tech

One Platform Constraint Forces an API Contract That Both Stores Reject

By Lucas Mendes/Jul 18, 2026

How Apple and Google's divergent store policies force mobile developers to maintain two incompatible API contracts, adding latency, complexity, and cost.
Tech

One Unmerged Config Pull Request Left Three Maintainers Running Manual Deploys

By Lucas Mendes/Jul 18, 2026

A stalled config PR forced three maintainers into manual deploys for weeks. This article examines the review bottleneck, tooling gaps, and governance patterns that prevent such failures.
Tech

One Maintainer Turned an Apache License Violation Into a Seven-Figure Consulting Retainer

By Sara Park/Jul 18, 2026

How a maintainer turned an Apache 2.0 license violation into a $15,000/month consulting retainer, totaling over $900,000 in five years—a case study in open source monetization.
Tech

One Database License Negotiation Determined an Entire Company's Exit Timeline

By Deepa Iyer/Jul 18, 2026

How a single database license negotiation can determine a startup's exit timeline. Analysis of pricing traps, vendor lock-in, and strategies to unchain your stack.
Tech

One Build Server's Clock Drift Caused Three Teams to Cache Invalid Artifacts

By Deepa Iyer/Jul 18, 2026

How a 47-millisecond clock drift on a single build server at Wavelength poisoned artifact caches across three teams, causing 12 hours of failed builds and a deeper lesson about time in distributed systems.
Tech

One Configuration Drift Took a GRPC Service Down Across All Five Regions

By Sara Park/Jul 18, 2026

A single boolean flag mismatch in a shared config file caused a multi-region gRPC outage. This postmortem traces the failure from field numbering to silent rejection and outlines safeguards.
Tech

One Cloud Provider's Pricing Grid Made a Fortune Off Build Minutes That Never Finished

By Lucas Mendes/Jul 18, 2026

How CloudProviderX charged for build minutes that never finished, turning infrastructure failures into a multi-million-dollar revenue stream—and the customer revolt that forced a change.
Tech

A Copyleft License Both Projects Used Fractured Their Contributor Base

By Sara Park/Jul 18, 2026

Two open-source projects adopted strict copyleft licenses. Both saw their contributor bases fracture as ideology clashed with pragmatism. A deep look at what each got right and wrong.