One Unmerged Config Pull Request Left Three Maintainers Running Manual Deploys

Jul 18, 2026 By Lucas Mendes

In early 2025, a mid-sized open-source project with roughly a dozen active contributors hit a wall that wasn't a code bug or a security vulnerability. It was a configuration pull request that stayed in draft for eleven days. That one unmerged change cascaded into a situation where three core maintainers spent the better part of two weeks running manual deployment cycles, each taking roughly 40 minutes, with no automated rollback path. The postmortem, shared internally and later summarized in a public blog post, called it a "failure of review rotation." This article walks through how that happened, what it cost, and what governance patterns could prevent it.

The Pull Request That Stayed Draft

The change itself was unremarkable: a configuration file that adjusted a rate-limiting threshold for an upstream API integration. The project used GitHub with CODEOWNERS to require approval from at least one of three listed maintainers. One of those maintainers was on vacation, another was overwhelmed with triage across multiple repositories, and the third—the one who normally handled config changes—was waiting for a second opinion on a subtle dependency ordering issue.

Days passed. The PR accumulated no comments beyond an initial "Looks okay, but let's wait for [maintainer B]." Meanwhile, the CI pipeline continued to run against the old configuration, which was causing intermittent failures in production. The team had disabled automatic deploys to avoid shipping a known-bad state, but that decision inadvertently locked out the fix.

After roughly a week, the maintainers realized the situation was untenable. They decided to manually build and deploy the artifact from the PR branch, bypassing the usual pipeline. The first manual deploy took about 45 minutes, including shell work, secret rotation, and a sanity check. Over the next few days, they repeated this process roughly a dozen times, each cycle taking 30–50 minutes depending on context switching.

The postmortem, published roughly three weeks later, listed three root causes: no automated rollback path, the inability to merge without the vacationing maintainer's sign-off, and the lack of a time-based escalation for stalled PRs. A similar incident had been noted in a previous postmortem about platform team sync, but the fix had not been prioritized.

Why Single-Point Review Bottlenecks Form

Open-source maintainers face a well-documented triage problem. The Linux kernel, for example, sees thousands of patches per release cycle, but even smaller projects can suffer from notification overload. In this case, the three maintainers were subscribed to multiple repositories and issue trackers. One reported receiving roughly 200 GitHub notifications per day, making it easy to miss a draft PR that didn't trigger a review request.

GitHub's CODEOWNERS feature can require reviews from specific individuals or teams, but it cannot enforce a time-to-review. A PR can sit for weeks if the designated owner is unavailable. The project had no fallback rule—no secondary reviewer list, no automatic reassignment after 48 hours. This is a common pattern in small teams, where the bus factor is already high. As noted in a related article about a firmware maintainer's bus factor, single points of failure in review are often accepted as unavoidable until they cause an incident.

Review latency is amplified by the size of the team. With only three maintainers, a single vacation reduces the available reviewer pool by a third. If another is focused on a different subsystem, the effective pool for config changes can shrink to one. This is not a failure of individual effort but a structural gap in how review capacity is allocated.

Some projects mitigate this by using a rotating reviewer schedule or requiring at least two approvals from a larger set. But those patterns require a critical mass of contributors and a culture of shared responsibility. In this project, the maintainers had not yet adopted such practices, partly because the team had been stable for years and the problem had never surfaced so acutely.

Manual Deploy Days: The Operational Cost

Each manual deploy in this incident involved roughly 15–20 minutes of shell work: pulling the branch, building the artifact, copying it to the production server, updating configuration files, and restarting services. Then came the manual smoke test—running a few curl commands and checking logs. The entire cycle, from start to finish, averaged about 40 minutes, but interruptions and context switches often stretched it to an hour.

The error rate was significant. The team estimated that roughly 1 in 6 manual runs introduced a mistake: a typo in a file path, a forgotten secret rotation, or a service restart in the wrong order. One such error caused a five-minute production outage that affected a subset of users. The team caught it quickly, but the incident eroded confidence in the manual process.

There was no audit trail for configuration drift. Each manual deploy modified production state without leaving a clear record in a version-controlled system. The maintainers kept notes in a shared document, but those notes were not always updated. After the third incident, one maintainer wrote a script to log each manual action, but by then the damage to team morale was done.

Hotfix urgency overrode normal checks. In one instance, a maintainer skipped the smoke test because the fix was "trivial"—a single line change—and ended up deploying a stale artifact that didn't include a previous hotfix. The team had to roll back and redeploy, adding another 40 minutes. The cumulative time spent on manual deploys over the two-week period was roughly 12–15 hours, not counting the mental overhead of context switching.

Tooling Gaps That Enable the Scenario

The project lacked several pieces of infrastructure that would have prevented or mitigated the incident. There was no merge queue—no mechanism to enforce that all required status checks pass before merging. The branch protection rules were minimal: they required one approval but did not require updated branches or passing CI. This meant that even if the PR had been approved, it might have merged without the latest CI run.

CI secrets were rotated manually on shared machines. The maintainers used a shared development server where secrets were stored in plain-text environment files. This is not uncommon in small teams, but it creates a security risk and makes it harder to audit who accessed what. The postmortem recommended moving to a secrets manager, but as of the writing of this article, that migration had not been completed.

There was no pre-merge integration testing for configuration changes. The CI pipeline ran unit tests and linting, but it did not deploy to a staging environment that mirrored production. The team had a staging server, but it was not kept in sync with production—it ran a different database version and had fewer replicas. As a result, config changes that worked in staging sometimes failed in production due to environment-specific variables.

The accumulated tech debt from these gaps was estimated at roughly 2–3 weeks of engineering time, spread across multiple incidents over the previous six months. This is a hedged figure, as the team did not track tech debt formally, but the postmortem noted that addressing the root causes would require at least two sprints of focused work.

Governance Patterns That Prevent Repeat Failures

The most immediate fix was to adopt a rotating reviewer schedule. The project now assigns two maintainers to each week's review duty, with a third as backup. If a PR is not reviewed within 24 hours, it is automatically reassigned to the backup. This is a simple policy change that required no tooling investment, and it has already reduced the average review time from roughly 4 days to under 8 hours.

Setting a maximum merge-wait time of 24 hours is another pattern that works well in practice. Some projects use a bot that pings reviewers after 12 hours and escalates after 24. This does not guarantee a review, but it surfaces stalled PRs before they become critical. The project in question now uses a GitHub Action that sends a Slack reminder after 12 hours and reassigns after 24.

Automating rollback on failed health checks is a higher-effort change but a critical one. The team now runs a health check script after every deploy, and if the check fails, the deploy is automatically rolled back to the previous known-good state. This required adding a few hundred lines of infrastructure code, but it eliminated the manual rollback step that had caused the five-minute outage.

Merge trains with required status checks are another safeguard. By requiring that all commits in a train pass CI before any merge, the team ensures that no broken commit enters the main branch. This pattern is common in projects that use GitHub's merge queue feature, which was introduced in 2023. The project adopted it after the incident, and it has prevented at least two similar config-related stalls.

When Automation Can't Replace Human Judgment

Not every review bottleneck can be solved with tooling. Configuration changes often require domain-specific reasoning that automated tests cannot capture. In this incident, the subtle dependency ordering issue that delayed the PR was a legitimate concern: the new rate limit interacted poorly with a retry mechanism in a downstream service. A human reviewer needed to understand that interaction.

Automated tests missed environment-specific variables. The CI pipeline ran in a containerized environment that did not have the same network latency or service discovery as production. A mock test passed, but the real system would have behaved differently. The team later added integration tests that ran against a staging environment, but even those cannot cover every edge case.

Human review caught a subtle bug in a different PR during the same period: a change that would have caused a circular dependency in the startup sequence. The reviewer noticed it because they had recently worked on a similar issue. That kind of pattern recognition is hard to automate. The postmortem emphasized that the goal is not to eliminate human review but to ensure it happens in a timely manner.

Balancing automation and human judgment is an ongoing challenge. Over-automating rare paths can introduce complexity that itself becomes a maintenance burden. Some projects have adopted a policy of automating only the most common failure modes and leaving edge cases to human review. This is a pragmatic approach, but it requires regularly revisiting which paths are "common" as the project evolves.

The project's post-incident review improved the situation, but it did not eliminate risk. A similar incident could still occur if a different combination of circumstances aligns—say, two maintainers on leave and a config change that touches a new subsystem. The team now documents each incident and reviews the governance patterns every quarter, but they acknowledge that some degree of risk is inherent in any human-run system.

Trade-offs in Automation Investment

One common counter-argument is that the cost of automation can outweigh the benefits for infrequent incidents. In this project, the manual deploy period lasted only two weeks, and the team spent roughly 12–15 hours on it. Building a full merge queue with automated rollback required an estimated two sprints of development time—roughly 3–4 weeks of engineering effort for a team of two. At a typical fully-loaded engineering cost of around US$ 100–150 per hour, that investment is on the order of US$ 12,000–18,000. The direct labor cost of the manual deploys was only about US$ 1,500–2,000. However, the outage caused by the manual error affected a subset of users for five minutes, and the team's morale took a hit. Quantifying the full cost is difficult, but the postmortem argued that the automation investment was justified by the risk of a larger outage and the ongoing maintenance burden of manual processes.

Another trade-off is the complexity of the merge queue itself. GitHub's merge queue feature works well for projects with a linear commit history, but it can introduce delays when multiple PRs are queued. The project found that the average time from merge request to merge increased from roughly 10 minutes to about 20 minutes, because each PR must wait for CI to pass on a combined branch. For urgent hotfixes, the team now uses a bypass mechanism that requires two approvals and a manual override. This adds a small amount of friction, but it prevents the queue from being a bottleneck during emergencies.

Some projects argue that a simpler alternative is to expand the reviewer pool rather than investing in automation. Recruiting and onboarding new maintainers takes time and effort, but it directly addresses the bus factor. The project in question attempted to recruit two additional maintainers after the incident, but only one was successfully onboarded within three months. The other candidate dropped out due to time constraints. This highlights that human solutions are not always faster or easier than technical ones.

Finally, there is the question of whether the incident was actually a one-off. The team had experienced similar stalls in the past, but they had always been resolved by a maintainer returning from leave or a quick Slack ping. The postmortem noted that the incident was the first time the stall lasted more than a week, and it was the first time manual deploys were required. This suggests that the risk was latent but not fully appreciated. A cost-benefit analysis done before the incident would likely have concluded that automation was not worth it. After the incident, the calculus changed. This is a common pattern in incident-driven investment: the cost of inaction is only fully understood after a failure.

The project's experience is not unique. A survey of open-source maintainers conducted in 2024 found that roughly 40% of respondents had experienced a deployment delay of more than 24 hours due to a review bottleneck. Of those, about 15% resorted to manual deploys. The survey, which included responses from roughly 200 projects, also found that projects with a merge queue were half as likely to report manual deploys. While the survey's sample is not necessarily representative of all open-source projects, it suggests that the patterns discussed here are widespread.

Recommend Posts
Tech

One Build Engineer Trades a Safer Package Registry for a Two-Minute Install Lag

By Sara Park/Jul 18, 2026

A build engineer adopts a signed package registry for security, trading two minutes per install for verifiable provenance. The cost in developer hours and the industry's next steps.
Tech

One Maintainer's Twelve-Hour Firewall Patch Left a TLS Handshake Dead for Three Years

By Deepa Iyer/Jul 18, 2026

A single firewall patch by one OpenSSL maintainer silently broke TLS 1.3 resumption for three years, costing retransmission and developer hours. The story exposes the bus factor and funding gaps in critical infrastructure.
Tech

One Distributed Query’s Storage Layer Bill Exceeded Its Feature Budget by Five Figures

By Lucas Mendes/Jul 18, 2026

How a single distributed join triggered a five-figure cloud bill, and why storage economics must be a first-class query constraint for engineering teams.
Tech

A Distributed Systems Role Pays Less Than Monolith Work at Equivalent Scale

By Lucas Mendes/Jul 18, 2026

Engineers working on distributed systems often earn 10–15% less than peers on monoliths at similar scale. The article examines why and how to navigate the gap.
Tech

One Monorepo’s Shared Schema Enum Forced Thirty Teams Into a Single Error String

By Yusuke Tanaka/Jul 18, 2026

How a single protobuf enum in a monorepo root forced thirty teams to standardize error strings, increased build times, and led to workarounds that defeated schema enforcement. Lessons from Google’s error model and a pragmatic shard fix.
Tech

Operating Cost Drives an LLM Provider's API Price to Ten Times the Inference

By Lucas Mendes/Jul 18, 2026

LLM API prices can exceed inference costs by 10x. This article breaks down the operating expenses, contract lock-ins, and what procurement teams can do about it.
Tech

One Platform Team’s Private API Cost Ten Engineers a Week of Manual Sync

By Yusuke Tanaka/Jul 18, 2026

A platform team's undocumented endpoint forced ten engineers into a week of manual reconciliation. Here's how contract-first development and shared tooling eliminated the waste.
Tech

One Auth Engineer Replaced Eight Vendor SDKs With a Single LDAP Config File

By Lucas Mendes/Jul 18, 2026

How one engineer replaced eight authentication SDKs with a single LDAP config, cutting attack surface and maintenance overhead. A deep dive into the trade-offs and operational reality.
Tech

One Maintainer’s Two-Line CSS Fix Cut Load Times by Forty Percent

By Yusuke Tanaka/Jul 18, 2026

A single maintainer cut LCP by 40% with a two-line CSS change. This article breaks down the fix, why modern bundlers miss it, and how to apply it without new tooling.
Tech

One Cloud Database Vendor’s Write Path Locked Nine Clients Into a Single SLA Clock

By Yusuke Tanaka/Jul 18, 2026

How a shared consensus group and single clock source penalize fast writers in multi-tenant databases, and what engineering teams can do about it.
Tech

One Build System’s Config Parser Swallowed Three Teams’ Deployment Scripts

By Yusuke Tanaka/Jul 18, 2026

Monzo's custom TOML parser silently dropped unknown keys for years. When a strict mode update shipped, three teams' deployment scripts broke. A post-mortem reveals the root cause and lessons for build system maintainers.
Tech

One Supply Chain Engineer's Two-Week Patch Audit Found Dormant Signing Keys Across Seven SDKs

By Yusuke Tanaka/Jul 18, 2026

A routine audit by a supply chain engineer uncovered dormant signing keys in seven SDKs, exposing hundreds of apps to potential supply chain attacks. Here's how they did it and what teams can learn.
Tech

One Platform Constraint Forces an API Contract That Both Stores Reject

By Lucas Mendes/Jul 18, 2026

How Apple and Google's divergent store policies force mobile developers to maintain two incompatible API contracts, adding latency, complexity, and cost.
Tech

One Unmerged Config Pull Request Left Three Maintainers Running Manual Deploys

By Lucas Mendes/Jul 18, 2026

A stalled config PR forced three maintainers into manual deploys for weeks. This article examines the review bottleneck, tooling gaps, and governance patterns that prevent such failures.
Tech

One Maintainer Turned an Apache License Violation Into a Seven-Figure Consulting Retainer

By Sara Park/Jul 18, 2026

How a maintainer turned an Apache 2.0 license violation into a $15,000/month consulting retainer, totaling over $900,000 in five years—a case study in open source monetization.
Tech

One Database License Negotiation Determined an Entire Company's Exit Timeline

By Deepa Iyer/Jul 18, 2026

How a single database license negotiation can determine a startup's exit timeline. Analysis of pricing traps, vendor lock-in, and strategies to unchain your stack.
Tech

One Build Server's Clock Drift Caused Three Teams to Cache Invalid Artifacts

By Deepa Iyer/Jul 18, 2026

How a 47-millisecond clock drift on a single build server at Wavelength poisoned artifact caches across three teams, causing 12 hours of failed builds and a deeper lesson about time in distributed systems.
Tech

One Configuration Drift Took a GRPC Service Down Across All Five Regions

By Sara Park/Jul 18, 2026

A single boolean flag mismatch in a shared config file caused a multi-region gRPC outage. This postmortem traces the failure from field numbering to silent rejection and outlines safeguards.
Tech

One Cloud Provider's Pricing Grid Made a Fortune Off Build Minutes That Never Finished

By Lucas Mendes/Jul 18, 2026

How CloudProviderX charged for build minutes that never finished, turning infrastructure failures into a multi-million-dollar revenue stream—and the customer revolt that forced a change.
Tech

A Copyleft License Both Projects Used Fractured Their Contributor Base

By Sara Park/Jul 18, 2026

Two open-source projects adopted strict copyleft licenses. Both saw their contributor bases fracture as ideology clashed with pragmatism. A deep look at what each got right and wrong.