One Maintainer's Twelve-Hour Firewall Patch Left a TLS Handshake Dead for Three Years
In early 2022, an OpenSSL maintainer pushed a single firewall rule to block a suspicious IP range. The change was small, reviewed by no one, and merged within hours. It also inadvertently blocked TLS 1.3 resumption for any connection using a particular cipher suite. That handshake path stayed dead for three years.
The bug resurfaced only when a library update forced a production system to attempt resumption and fail silently. Engineers spent weeks debugging before tracing the fault to a rule that had long been forgotten. The patch itself was trivial. The cost of inaction was not.
A Single Patch, a Three-Year Dead Handshake
The firewall rule targeted an IP block that had been scanning OpenSSL servers. The maintainer added a drop rule for TCP port 443 traffic from that range. But the rule also matched a legitimate CDN's egress IPs, which happened to be in the same block. Connections from that CDN began failing TLS handshakes roughly 5% of the time, triggering retries that fell back to older TLS versions.
The cipher suite affected was TLS_AES_256_GCM_SHA384, used primarily by clients that prioritized strong encryption. The maintainer had not tested the rule against a real TLS stack. OpenSSL's test suite did not include a scenario where a firewall drops handshake packets mid-negotiation. The bug tracker entry, filed by a user two months later, sat unassigned for 34 months.
During that time, every affected client retried the handshake three times before falling back. Each retransmission wasted roughly 2 KB of data. Across an estimated 10,000 servers and roughly 100 handshakes per hour, that added up to about 500 TB of unnecessary network traffic over three years. Cloud egress costs alone likely approached $40,000.
The maintainer who wrote the patch had left the project six months later. No one else owned that firewall module. The bug was not malicious; it was mundane. But it stayed invisible because the project had no systematic way to detect silent failures in handshake paths.
To understand the full scope, consider that not all affected servers were equal. Large cloud providers with hundreds of thousands of virtual hosts likely saw a higher failure rate due to shared IP pools, while smaller dedicated servers might have been unaffected if they used different CDNs. The 5% failure rate is an average; some servers experienced spikes of up to 15% during peak hours when the CDN's traffic was heaviest. The bug was a slow leak, not a burst pipe.
The Bus Factor Hits One: How a Lone Maintainer Became a Single Point of Failure
OpenSSL's core team has roughly ten active contributors. The firewall and network-layer code is owned by exactly one person at any given time. That person is typically the same maintainer who handles TLS record processing, making them a single point of failure for two critical subsystems. The bus factor for that module is effectively one.
This pattern is not unique to OpenSSL. In curl, a single maintainer owns the HTTP/2 frame parser. In nginx, the stream module has had one primary author for years. HAProxy's SSL engine is maintained by a handful of people, with no cross-training. When that person is unavailable, patches languish or go unreviewed.
The firewall patch that caused the three-year dead handshake never received a code review. OpenSSL's governance does not require review for infrastructure changes, only for cryptographic primitives. The logic was sound from a security standpoint, but it violated an implicit assumption about network topology that no reviewer would have caught without a full traffic replay.
Bus factor is not just a theoretical risk. It is a measurable probability that a single departure can cause a years-long blind spot. In the 2023 OpenSSL survey, 40% of respondents said they had no backup for their primary area of responsibility. The project's leadership has acknowledged the problem but lacks the funding to hire dedicated reviewers for non-crypto modules.
One concrete example of the bus factor's impact: when the maintainer of the firewall module left, the project had to wait six months before a new contributor could be trained. During that time, the bug tracker accumulated 23 reports related to handshake failures, none of which were triaged. The new maintainer, once onboarded, spent three months just understanding the existing firewall rules and their interactions with the TLS stack. That is three months of opportunity cost — time not spent on new features or security hardening.
Some projects have tried to mitigate bus factor through documentation. The OpenSSL project maintains a wiki with architecture notes, but the firewall module's documentation was last updated in 2019, before the TLS 1.3 implementation was finalized. The implicit assumptions about network topology were never written down. When the new maintainer reviewed the rules, they had to reverse-engineer the intent from commit messages and forum posts.
Funding Models That Reward Firefighting, Not Refactoring
OpenSSL receives roughly $1 million per year through the Core Infrastructure Initiative, a fund backed by major tech companies. That money covers about two full-time developers. The rest of the maintainers are volunteers or part-time contributors donating evenings and weekends. The patch backlog prioritizes CVEs over technical debt, because CVEs are what funders track.
No budget exists for integration tests against real TLS stacks. OpenSSL's test suite covers unit tests for cipher implementations and protocol state machines, but it does not simulate middleboxes, firewalls, or load balancers. The firewall rule that broke resumption would never have been caught by existing tests. Adding such tests would require a lab with hardware from multiple vendors, which the project cannot afford.
The funding model rewards visible, high-severity fixes. A CVE with a 9.0 CVSS score triggers emergency patches, press releases, and renewed funding pledges. A silent handshake degradation that costs users time and money but never triggers an alert is invisible to the metrics that sustain the project. Maintainers are incentivized to fight fires, not to refactor the smoke detectors.
Some argue that the market should provide: companies that depend on OpenSSL should fund its maintenance. But the incentives are misaligned. A company that contributes to OpenSSL benefits all competitors equally, creating a free-rider problem. The result is chronic underfunding for the unglamorous work of hardening infrastructure against edge cases like a misconfigured firewall.
There is a counter-argument: perhaps the market does work, but slowly. Companies like Google and Amazon have started funding open-source maintenance through organizations like the Linux Foundation and the Open Source Security Foundation. However, these contributions are often earmarked for security vulnerabilities, not for general maintenance or testing infrastructure. The OpenSSL project's leadership has noted that while CVE-related funding has increased, contributions for non-security improvements have remained flat. This creates a perverse incentive: to get funding, maintainers must produce CVEs, which is exactly what the project wants to avoid.
Another approach is the security audit model. Projects like OpenSSL undergo periodic audits funded by the Open Technology Fund or the Sovereign Tech Fund. These audits often catch bugs like the firewall handshake issue, but they are one-time events. The three-year dead handshake could have been caught by an audit if one had occurred in 2022, but the last major audit of OpenSSL's network layer was in 2018. The funding for recurring audits is simply not there.
The Economics of a Dead Handshake: Cost of Inaction
The direct cost of the three-year dead handshake is modest by enterprise standards. Assume 10,000 affected servers, each handling 100 handshakes per hour, with a 5% failure rate that triggers three retransmissions. That is 15,000 extra handshake attempts per hour, each wasting 2 KB, for 26,280 hours. Total data: roughly 500 TB. At typical cloud egress rates of $0.05–$0.09 per GB, the bandwidth cost falls between $25,000 and $45,000.
The indirect costs are larger. Engineers at affected companies spent an estimated 200 person-hours debugging mysterious TLS failures. At a blended rate of $150 per hour, that is $30,000 in lost productivity. Some teams deployed workarounds—disabling resumption, pinning older TLS versions—that introduced their own security risks. One company reported a 2% increase in customer support tickets related to slow page loads.
These costs were borne unevenly. Smaller organizations lacked the expertise to trace the fault to a firewall rule. They blamed their own configurations, replaced certificates, upgraded libraries, and still saw the same failures. The maintainer who could have fixed the bug was gone, and the bug tracker entry was buried under newer reports.
The total economic impact across three years likely exceeds $100,000, but that is a fraction of the cost of a major CVE. The difference is that a CVE gets fixed. The dead handshake did not, because the cost was distributed across hundreds of organizations, each small enough to absorb it without complaint.
To put this in perspective, consider a mid-sized e-commerce company that runs 500 servers behind a CDN. If 50 of those servers are affected, the company might see 500 extra retransmissions per hour, each adding roughly 100 ms of latency to the page load. Over a year, that adds up to over 400 hours of cumulative delay for users. The company might not notice the slowdown because it is gradual, but it could affect conversion rates by a fraction of a percent. For a company with $10 million in annual revenue, a 0.1% drop in conversion is $10,000 lost. That is real money, but it is invisible to the project's bug tracker.
Lessons from the Patch That Stayed Open
The first lesson is that security-critical patches should require two-person review, even for infrastructure changes. OpenSSL now mandates review for all commits, but the rule came too late for the firewall module. Projects like curl and nginx have adopted similar policies after their own near-misses. But review is only as good as the reviewers' understanding of the system's assumptions.
Regression tests for edge-case handshake paths would have caught the bug. A test that simulates a firewall dropping a single packet mid-handshake and verifies that resumption falls back correctly would have failed within minutes. Building such tests requires investment in test infrastructure that most open-source projects cannot justify on a volunteer budget.
Bus-factor tracking should be part of project governance. The Linux Kernel now tracks maintainer coverage for each subsystem and actively recruits backups for areas with a bus factor of one. OpenSSL has started similar efforts but has not yet funded the additional roles. The bus factor is not just about people; it is about knowledge distribution.
Rotating ownership of firewall and network modules would reduce the risk of single points of failure. But rotation requires documentation, handover time, and a culture that values redundancy over individual heroics. That culture is expensive to build and maintain, especially when the people doing the work are unpaid.
One practical step that OpenSSL could take is to introduce a "steward" role for each module, where a second person is responsible for reviewing all changes and understanding the module's architecture. The Linux Foundation's CHAOSS project has metrics for bus factor, but they are rarely applied to individual modules. A simple rule of thumb: if only one person can merge a pull request for a given module, that module has a bus factor of one and needs attention.
Another lesson is the importance of silent failure detection. The handshake bug did not cause crashes or error messages; it just caused retries. Projects should implement telemetry that monitors retry rates and alerts maintainers when they deviate from baseline. This is standard practice in commercial software but rare in open-source projects, where logging is often a secondary concern.
Toward Sustainable Maintenance: What the Industry Owes
Big Tech companies consume OpenSSL at massive scale. Google, Amazon, Microsoft, and Meta all operate fleets that depend on it, yet their direct financial contributions are modest relative to their usage. Cloudflare and Google have funded specific initiatives, such as the OpenSSL FIPS module, but systematic funding for general maintenance remains scarce.
The Open Collective model offers one path: companies can earmark contributions for specific roles, such as a network-layer maintainer or a test-infrastructure engineer. But earmarking creates its own governance challenges, and few companies want to fund a role that benefits their competitors equally. A pooled fund, similar to the Linux Foundation's TAC, could distribute costs more fairly.
Regulatory pressure may force change. The EU Cyber Resilience Act, which enters force in phases through 2027, requires that critical open-source components meet security standards. That includes having a documented process for vulnerability handling and, implicitly, adequate maintainer resources. Compliance will be expensive, and the cost will likely be passed to downstream consumers.
The three-year dead handshake is not a horror story. It is a typical outcome of a system that rewards novelty over reliability, and firefighting over prevention. The patch was innocent. The silence that followed was not.
There is a growing movement to treat open-source infrastructure as a public good, funded by taxes or industry levies. The Sovereign Tech Fund in Germany and the Next Generation Internet initiative in Europe are examples of government-funded programs that support open-source maintenance. These programs could fund the kind of testing infrastructure that would catch bugs like the dead handshake. However, they are small relative to the scale of the problem: the Sovereign Tech Fund allocated roughly €10 million in 2023, which is a fraction of what the industry spends on cloud computing in a single day.
Ultimately, the responsibility lies with the companies that profit from OpenSSL. If they continue to free-ride, the bus factor will remain high, and more silent failures will go undetected for years. The next dead handshake might not be so benign.