One Unmerged Config Pull Request Left Three Maintainers Running Manual Deploys
In early 2025, a mid-sized open-source project with roughly a dozen active contributors hit a wall that wasn't a code bug or a security vulnerability. It was a configuration pull request that stayed in draft for eleven days. That one unmerged change cascaded into a situation where three core maintainers spent the better part of two weeks running manual deployment cycles, each taking roughly 40 minutes, with no automated rollback path. The postmortem, shared internally and later summarized in a public blog post, called it a "failure of review rotation." This article walks through how that happened, what it cost, and what governance patterns could prevent it.
The Pull Request That Stayed Draft
The change itself was unremarkable: a configuration file that adjusted a rate-limiting threshold for an upstream API integration. The project used GitHub with CODEOWNERS to require approval from at least one of three listed maintainers. One of those maintainers was on vacation, another was overwhelmed with triage across multiple repositories, and the third—the one who normally handled config changes—was waiting for a second opinion on a subtle dependency ordering issue.
Days passed. The PR accumulated no comments beyond an initial "Looks okay, but let's wait for [maintainer B]." Meanwhile, the CI pipeline continued to run against the old configuration, which was causing intermittent failures in production. The team had disabled automatic deploys to avoid shipping a known-bad state, but that decision inadvertently locked out the fix.
After roughly a week, the maintainers realized the situation was untenable. They decided to manually build and deploy the artifact from the PR branch, bypassing the usual pipeline. The first manual deploy took about 45 minutes, including shell work, secret rotation, and a sanity check. Over the next few days, they repeated this process roughly a dozen times, each cycle taking 30–50 minutes depending on context switching.
The postmortem, published roughly three weeks later, listed three root causes: no automated rollback path, the inability to merge without the vacationing maintainer's sign-off, and the lack of a time-based escalation for stalled PRs. A similar incident had been noted in a previous postmortem about platform team sync, but the fix had not been prioritized.
Why Single-Point Review Bottlenecks Form
Open-source maintainers face a well-documented triage problem. The Linux kernel, for example, sees thousands of patches per release cycle, but even smaller projects can suffer from notification overload. In this case, the three maintainers were subscribed to multiple repositories and issue trackers. One reported receiving roughly 200 GitHub notifications per day, making it easy to miss a draft PR that didn't trigger a review request.
GitHub's CODEOWNERS feature can require reviews from specific individuals or teams, but it cannot enforce a time-to-review. A PR can sit for weeks if the designated owner is unavailable. The project had no fallback rule—no secondary reviewer list, no automatic reassignment after 48 hours. This is a common pattern in small teams, where the bus factor is already high. As noted in a related article about a firmware maintainer's bus factor, single points of failure in review are often accepted as unavoidable until they cause an incident.
Review latency is amplified by the size of the team. With only three maintainers, a single vacation reduces the available reviewer pool by a third. If another is focused on a different subsystem, the effective pool for config changes can shrink to one. This is not a failure of individual effort but a structural gap in how review capacity is allocated.
Some projects mitigate this by using a rotating reviewer schedule or requiring at least two approvals from a larger set. But those patterns require a critical mass of contributors and a culture of shared responsibility. In this project, the maintainers had not yet adopted such practices, partly because the team had been stable for years and the problem had never surfaced so acutely.
Manual Deploy Days: The Operational Cost
Each manual deploy in this incident involved roughly 15–20 minutes of shell work: pulling the branch, building the artifact, copying it to the production server, updating configuration files, and restarting services. Then came the manual smoke test—running a few curl commands and checking logs. The entire cycle, from start to finish, averaged about 40 minutes, but interruptions and context switches often stretched it to an hour.
The error rate was significant. The team estimated that roughly 1 in 6 manual runs introduced a mistake: a typo in a file path, a forgotten secret rotation, or a service restart in the wrong order. One such error caused a five-minute production outage that affected a subset of users. The team caught it quickly, but the incident eroded confidence in the manual process.
There was no audit trail for configuration drift. Each manual deploy modified production state without leaving a clear record in a version-controlled system. The maintainers kept notes in a shared document, but those notes were not always updated. After the third incident, one maintainer wrote a script to log each manual action, but by then the damage to team morale was done.
Hotfix urgency overrode normal checks. In one instance, a maintainer skipped the smoke test because the fix was "trivial"—a single line change—and ended up deploying a stale artifact that didn't include a previous hotfix. The team had to roll back and redeploy, adding another 40 minutes. The cumulative time spent on manual deploys over the two-week period was roughly 12–15 hours, not counting the mental overhead of context switching.
Tooling Gaps That Enable the Scenario
The project lacked several pieces of infrastructure that would have prevented or mitigated the incident. There was no merge queue—no mechanism to enforce that all required status checks pass before merging. The branch protection rules were minimal: they required one approval but did not require updated branches or passing CI. This meant that even if the PR had been approved, it might have merged without the latest CI run.
CI secrets were rotated manually on shared machines. The maintainers used a shared development server where secrets were stored in plain-text environment files. This is not uncommon in small teams, but it creates a security risk and makes it harder to audit who accessed what. The postmortem recommended moving to a secrets manager, but as of the writing of this article, that migration had not been completed.
There was no pre-merge integration testing for configuration changes. The CI pipeline ran unit tests and linting, but it did not deploy to a staging environment that mirrored production. The team had a staging server, but it was not kept in sync with production—it ran a different database version and had fewer replicas. As a result, config changes that worked in staging sometimes failed in production due to environment-specific variables.
The accumulated tech debt from these gaps was estimated at roughly 2–3 weeks of engineering time, spread across multiple incidents over the previous six months. This is a hedged figure, as the team did not track tech debt formally, but the postmortem noted that addressing the root causes would require at least two sprints of focused work.
Governance Patterns That Prevent Repeat Failures
The most immediate fix was to adopt a rotating reviewer schedule. The project now assigns two maintainers to each week's review duty, with a third as backup. If a PR is not reviewed within 24 hours, it is automatically reassigned to the backup. This is a simple policy change that required no tooling investment, and it has already reduced the average review time from roughly 4 days to under 8 hours.
Setting a maximum merge-wait time of 24 hours is another pattern that works well in practice. Some projects use a bot that pings reviewers after 12 hours and escalates after 24. This does not guarantee a review, but it surfaces stalled PRs before they become critical. The project in question now uses a GitHub Action that sends a Slack reminder after 12 hours and reassigns after 24.
Automating rollback on failed health checks is a higher-effort change but a critical one. The team now runs a health check script after every deploy, and if the check fails, the deploy is automatically rolled back to the previous known-good state. This required adding a few hundred lines of infrastructure code, but it eliminated the manual rollback step that had caused the five-minute outage.
Merge trains with required status checks are another safeguard. By requiring that all commits in a train pass CI before any merge, the team ensures that no broken commit enters the main branch. This pattern is common in projects that use GitHub's merge queue feature, which was introduced in 2023. The project adopted it after the incident, and it has prevented at least two similar config-related stalls.
When Automation Can't Replace Human Judgment
Not every review bottleneck can be solved with tooling. Configuration changes often require domain-specific reasoning that automated tests cannot capture. In this incident, the subtle dependency ordering issue that delayed the PR was a legitimate concern: the new rate limit interacted poorly with a retry mechanism in a downstream service. A human reviewer needed to understand that interaction.
Automated tests missed environment-specific variables. The CI pipeline ran in a containerized environment that did not have the same network latency or service discovery as production. A mock test passed, but the real system would have behaved differently. The team later added integration tests that ran against a staging environment, but even those cannot cover every edge case.
Human review caught a subtle bug in a different PR during the same period: a change that would have caused a circular dependency in the startup sequence. The reviewer noticed it because they had recently worked on a similar issue. That kind of pattern recognition is hard to automate. The postmortem emphasized that the goal is not to eliminate human review but to ensure it happens in a timely manner.
Balancing automation and human judgment is an ongoing challenge. Over-automating rare paths can introduce complexity that itself becomes a maintenance burden. Some projects have adopted a policy of automating only the most common failure modes and leaving edge cases to human review. This is a pragmatic approach, but it requires regularly revisiting which paths are "common" as the project evolves.
The project's post-incident review improved the situation, but it did not eliminate risk. A similar incident could still occur if a different combination of circumstances aligns—say, two maintainers on leave and a config change that touches a new subsystem. The team now documents each incident and reviews the governance patterns every quarter, but they acknowledge that some degree of risk is inherent in any human-run system.
Trade-offs in Automation Investment
One common counter-argument is that the cost of automation can outweigh the benefits for infrequent incidents. In this project, the manual deploy period lasted only two weeks, and the team spent roughly 12–15 hours on it. Building a full merge queue with automated rollback required an estimated two sprints of development time—roughly 3–4 weeks of engineering effort for a team of two. At a typical fully-loaded engineering cost of around US$ 100–150 per hour, that investment is on the order of US$ 12,000–18,000. The direct labor cost of the manual deploys was only about US$ 1,500–2,000. However, the outage caused by the manual error affected a subset of users for five minutes, and the team's morale took a hit. Quantifying the full cost is difficult, but the postmortem argued that the automation investment was justified by the risk of a larger outage and the ongoing maintenance burden of manual processes.
Another trade-off is the complexity of the merge queue itself. GitHub's merge queue feature works well for projects with a linear commit history, but it can introduce delays when multiple PRs are queued. The project found that the average time from merge request to merge increased from roughly 10 minutes to about 20 minutes, because each PR must wait for CI to pass on a combined branch. For urgent hotfixes, the team now uses a bypass mechanism that requires two approvals and a manual override. This adds a small amount of friction, but it prevents the queue from being a bottleneck during emergencies.
Some projects argue that a simpler alternative is to expand the reviewer pool rather than investing in automation. Recruiting and onboarding new maintainers takes time and effort, but it directly addresses the bus factor. The project in question attempted to recruit two additional maintainers after the incident, but only one was successfully onboarded within three months. The other candidate dropped out due to time constraints. This highlights that human solutions are not always faster or easier than technical ones.
Finally, there is the question of whether the incident was actually a one-off. The team had experienced similar stalls in the past, but they had always been resolved by a maintainer returning from leave or a quick Slack ping. The postmortem noted that the incident was the first time the stall lasted more than a week, and it was the first time manual deploys were required. This suggests that the risk was latent but not fully appreciated. A cost-benefit analysis done before the incident would likely have concluded that automation was not worth it. After the incident, the calculus changed. This is a common pattern in incident-driven investment: the cost of inaction is only fully understood after a failure.
The project's experience is not unique. A survey of open-source maintainers conducted in 2024 found that roughly 40% of respondents had experienced a deployment delay of more than 24 hours due to a review bottleneck. Of those, about 15% resorted to manual deploys. The survey, which included responses from roughly 200 projects, also found that projects with a merge queue were half as likely to report manual deploys. While the survey's sample is not necessarily representative of all open-source projects, it suggests that the patterns discussed here are widespread.