One Cloud Database Vendor’s Write Path Locked Nine Clients Into a Single SLA Clock
In a cloud database, the write path is the spine of every transaction. When nine clients discovered that their vendor's shared infrastructure locked them into a single SLA clock, their P99 latencies tripled during a neighbor's batch insert. The culprit wasn't a bug—it was the consensus protocol itself, serializing commits across tenants that should have been isolated.
The Write Path Bottleneck That No One Talked About
Every write in a distributed database passes through a write-ahead log (WAL) before it is applied to storage. In single-tenant deployments, that log is dedicated to one workload. But in multi-tenant setups, a shared WAL means all tenants march to the same drumbeat. The vendor in question used a single Raft group to order all commits across nine customers. The result: a commit latency floor set by the slowest participant.
Latency spikes ripple across tenants. A bursty client running a bulk data load can flood the leader with entries, causing followers to lag. Meanwhile, a latency-sensitive client performing point writes sees its commits queued behind the bulk load's entries. The shared log becomes a bottleneck that penalizes fast writers, who must wait for the cluster-wide clock to tick before their commit is acknowledged.
No isolation between workload profiles means a batch analytics job can degrade a real-time payment system. The vendor's SLA guaranteed a 99th percentile write latency of 10 milliseconds—but that was measured cluster-wide, not per tenant. When one tenant's workload spiked, the P99 for all nine clients jumped to 30 milliseconds or more. Fast writers paid the price for slow neighbors, and there was no escape.
To understand why this happens, consider the anatomy of a Raft commit. The leader receives a write, appends it to its log, then sends AppendEntries RPCs to followers. It must wait for a quorum (majority) of followers to acknowledge before it can apply the entry and respond to the client. In a multi-tenant cluster, all writes—regardless of tenant—go through the same leader and the same log. The leader processes entries in order, so a large batch insert from one tenant can block a single-row insert from another. This is not a scheduler issue; it is a fundamental property of the consensus protocol when applied to a shared log.
One might ask: why not use multiple Raft groups, one per tenant? That would require partitioning the cluster, which increases operational complexity and reduces resource utilization. The vendor chose a single group for simplicity and cost efficiency, but the trade-off is that tenants are not isolated at the write path level. The result is a system where the write latency of each tenant is a function of the aggregate workload, not its own.
How a Single SLA Clock Emerges from Shared Infrastructure
The root cause is architectural. Multi-tenant databases often share a single consensus group—Raft or Paxos—to order writes across all tenants. This group elects a leader that receives all write requests, appends them to a log, and replicates them to followers. The pace of commits is governed by the leader's clock, which must wait for a quorum of followers to acknowledge each entry.
Clock skew between nodes forces artificial waits. Even if all tenants send writes at the same rate, the leader must apply a global ordering. A fast writer's commit cannot proceed until the leader has finished replicating the previous entry—which may belong to a slow writer. The SLA guarantees set a worst-case upper bound, but the actual latency for any given write is a function of the cluster's aggregate load, not the tenant's own workload.
Fast committers pay for slow neighbors' delays. In one observed case, a tenant performing single-row inserts saw consistent 5-millisecond latencies during quiet hours. When a neighbor launched a batch insert of 10,000 rows, the same tenant's P99 shot to 22 milliseconds. The shared log queue had grown, and the leader spent more time replicating large entries. The fast writer's writes were stuck behind the slow writer's bulk load.
But the problem is not just about batch inserts. Even a steady stream of small writes from multiple tenants can cause contention. The leader's CPU and network bandwidth are finite resources. When ten tenants each send 1,000 writes per second, the leader must process 10,000 writes per second. The latency of each write includes queuing delay at the leader, which grows non-linearly with load. In a single-tenant system, that queuing delay is a function of the tenant's own throughput. In a shared system, it is a function of the sum of all tenants' throughputs. This is the hidden tax: a tenant that uses 10% of the leader's capacity still experiences the queuing delay of 100% utilization when other tenants are active.
Another subtle effect is the impact of network latency. Followers may be geographically distributed, and the leader must wait for the slowest follower in the quorum. If one tenant's writes are replicated to a follower in a different region, the commit latency for all tenants increases. The vendor may optimize by placing followers close to the leader, but that reduces fault tolerance. These trade-offs are rarely visible to the customer.
Nine Clients, One Throttle: The Contractual Trap
Each client signed the same write latency SLA, but the fine print tied performance to a cluster-wide clock. The vendor's terms specified a 99th percentile latency of 10 milliseconds measured across all writes in the cluster. This meant the vendor could meet its SLA even if individual tenants experienced spikes, as long as the aggregate average stayed within bounds. For the nine clients, the guarantee was effectively meaningless.
No per-tenant throughput guarantees existed in the contracts. A client could see its write throughput drop by half during a neighbor's peak, yet the vendor would claim the SLA was satisfied. The contractual language defined “write latency” as the time from request receipt to acknowledgment, but it did not specify how that latency was measured or what constituted a breach. Clients assumed they were isolated; they were not.
Bursty workloads degraded all nine simultaneously. One tenant's periodic batch jobs—running every hour for five minutes—would spike the cluster's P99 to 40 milliseconds. The other eight tenants had no recourse: they could not request a dedicated log shard without renegotiating their contract at a premium. Renegotiation leverage was zero for individual clients, as the vendor held the only copy of the data and migration costs were prohibitive.
This contractual trap is not unique to this vendor. Many cloud database providers use cluster-level SLAs for multi-tenant offerings. The rationale is that per-tenant SLAs are harder to monitor and enforce, and they reduce the provider's ability to oversubscribe. But the effect is that the customer bears the risk of interference. A savvy procurement team should ask for a per-tenant SLA, but most teams do not know to ask.
Moreover, the vendor's monitoring dashboards often show cluster-level metrics, masking per-tenant variance. A tenant might see a steady 5 ms average but a P99 of 50 ms, with no visibility into the cause. The vendor's support team may attribute the spikes to "normal multi-tenant behavior" and refuse to escalate. In one case, a client spent three months debugging their application code before discovering that a neighbor's ETL job was the root cause.
Measuring the Hidden Tax on Write Throughput
To quantify the impact, a team instrumented their application with per-tenant latency histograms. During a neighbor's batch insert, the P99 latency jumped from 8 milliseconds to 31 milliseconds—a 3.9x increase. The cluster-level P99, however, only rose from 9 to 14 milliseconds, because the batch insert's writes were faster on average. The cluster metric masked the tenant's pain.
Idle cluster still showed a 10-millisecond floor due to clock sync. Even with no load, the leader and followers exchanged heartbeat and clock synchronization messages that added a baseline latency. This floor was baked into the shared log's design: every commit required a round-trip to a quorum, and clock skew between nodes added a small but consistent delay. For a tenant doing single-millisecond writes, the floor was a 10x penalty.
Benchmarks mask cross-tenant interference. Vendor-published benchmarks typically measure a single tenant running on a dedicated cluster. In multi-tenant setups, performance degrades non-linearly. A benchmark showing 5-millisecond P99 for a single tenant becomes 25 milliseconds when four tenants share the same Raft group. The shared log bottleneck is invisible in standard tests, yet it dominates real-world workloads.
To illustrate, consider a simple queuing model. The leader processes writes at a rate of R writes per second. Each write has a service time S, which includes network I/O and log append. The queue length Q is proportional to utilization U = λ / R, where λ is the aggregate arrival rate. The average waiting time in the queue is (U * S) / (1 - U). In a single-tenant system, λ is the tenant's own rate. In a shared system, λ is the sum of all tenants' rates. If each tenant uses 10% of the leader's capacity, then with nine tenants, U = 0.9, and the waiting time is (0.9 * S) / 0.1 = 9S. That is a 9x increase in latency compared to a single tenant at 10% utilization, where waiting time is (0.1 * S) / 0.9 ≈ 0.11S. The math is simple, but the effect is dramatic.
Another measurement pitfall is that client-side retries can inflate latency. When a write times out, the client retries, and the retry adds to the load. In a shared system, retries from one tenant can exacerbate contention for others. A tenant with a poorly configured retry policy can cause a cascade of failures. The vendor may not detect this because retries are client-side.
Workarounds That Engineering Teams Actually Use
Client-side batching amortizes the latency penalty. By grouping multiple writes into a single commit request, a tenant can reduce the number of round-trips and improve throughput. The downside: each write's latency increases by the batching interval. For latency-sensitive operations, this is a non-starter. Teams often use two client configurations—one for batch writes, another for low-latency writes—but the shared log still serializes them.
Separate cluster for latency-critical writes is the most common escape. Teams spin up a dedicated cluster for their high-priority workloads, paying for the full infrastructure. This eliminates cross-tenant interference but doubles cost. For startups or mid-size companies, the expense is often unjustifiable. Some vendors offer “dedicated log shards” at a premium, which effectively creates a single-tenant Raft group within the shared cluster.
Asynchronous fallback paths for non-SLA ops let teams bypass the shared log for writes that do not require immediate consistency. A team might use a local queue to acknowledge writes quickly, then asynchronously commit them to the database. This works for analytics or logging workloads but fails for transactional systems that require linearizability. Custom proxies that reorder commits locally have been built, but they add complexity and risk.
Another workaround is to use a different database for latency-sensitive writes. Some teams run a separate, single-tenant database (like PostgreSQL) for critical transactions and use the multi-tenant database for bulk analytics. This introduces data consistency challenges and operational overhead, but it can be effective if the two systems are synchronized asynchronously.
Rate limiting at the application layer can also help. By throttling a tenant's own write rate, a team can reduce the impact on others, but this is a self-imposed constraint that defeats the purpose of multi-tenancy. Some vendors offer rate limiting as a feature, but it is often disabled by default.
One team we spoke to built a custom middleware that routes writes to different Raft groups based on a tenant ID. They used a proxy that partitioned the key space into shards, each with its own consensus group. This required significant engineering effort, but it gave them per-tenant isolation. The downside: they had to manage the shard topology themselves, and rebalancing was complex. They eventually migrated to a vendor that supported per-tenant sharding natively.
What the Next Generation of Write Paths Must Change
Per-tenant logical clocks decouple commit pacing. Instead of a single global clock, each tenant gets its own logical timestamp, allowing the leader to order writes independently. This requires changes to the consensus protocol: the log must be sharded by tenant, with each shard maintaining its own Raft group. Early prototypes in open-source databases like CockroachDB and FoundationDB show that shard-level consensus can reduce interference.
Shard-level consensus groups replace the global log. Instead of one Raft group for the entire cluster, each tenant’s data lives in its own shard with its own replication group. Writes to different shards proceed in parallel, and clock skew between shards does not affect commit latency. The trade-off is increased metadata overhead and more complex rebalancing. But for workloads with strong isolation requirements, the benefit outweighs the cost.
SLA guarantees scoped to tenant, not cluster. Next-generation contracts should specify per-tenant latency percentiles, measured independently. Vendors who offer this today—like Google Cloud Spanner with per-table splits—charge a premium, but the market is moving toward finer granularity. Open-source alternatives like TiDB and YugabyteDB already prototype per-tenant isolation in their write paths, though production readiness varies.
Another promising approach is the use of epoch-based commit protocols, where each tenant's writes are grouped into epochs that are committed independently. This reduces the coupling between tenants and allows the system to parallelize commits. However, epoch-based protocols introduce complexity in failure handling and garbage collection.
Vendor lock-in dissolves when the clock is per-writer. Once a tenant can control its own commit pacing, the incentive to stay with a single vendor weakens. Data migration becomes cheaper when isolation is built in, not bolted on. The nine clients in our story could have avoided the trap by asking four questions before signing—questions that every database buyer should ask today.
But there is a counter-argument: per-tenant isolation increases cost and complexity. The vendor must manage many Raft groups, which means more network connections, more elections, and more metadata. The overhead may be justified for high-value tenants, but for low-value tenants, the shared log is cheaper. The market may segment into "premium" and "standard" tiers, with isolation as a priced feature. This is already happening: AWS Aurora Serverless v2 offers per-workload scaling, and Azure Cosmos DB offers dedicated throughput per partition.
Another concern is that per-tenant isolation can lead to resource fragmentation. If each tenant has its own Raft group, the cluster may have many small groups, each with its own leader and followers. This can increase the load on the cluster's metadata service and make rebalancing harder. Some vendors mitigate this by using a shared metadata layer with tenant-specific logs, but that adds latency.
Ultimately, the right approach depends on the workload. For a SaaS provider with many small tenants, a shared log with rate limiting may be sufficient. For a financial services firm with a few large tenants, per-tenant isolation is essential. The key is to understand the trade-offs and ask the right questions.
Four Questions to Ask Before Signing the Next SLA
First: Is the write path shared across tenants? If the vendor uses a single consensus group for all customers, you will share the clock. Demand a diagram of the replication architecture. Second: What clock source governs commit timestamps? A global wall clock or a single Raft leader's clock creates a shared bottleneck. Ask if each tenant gets its own logical clock or shard.
Third: Can we get a dedicated log shard? Even if the vendor offers multi-tenancy, some allow you to purchase a private log shard that isolates your writes. The cost may be high, but it is cheaper than a dedicated cluster. Fourth: Are latency SLAs measured per-tenant or cluster-wide? A cluster-wide SLA is worthless for a tenant with bursty neighbors. Insist on per-tenant percentiles in the contract, with independent monitoring and penalties for breaches.
These questions would have saved the nine clients months of debugging and renegotiation. The shared SLA clock is not a bug—it is a design choice that benefits the vendor, not the customer. As the database market matures, the winners will be those who decouple the write path, one tenant at a time.
Conclusion: The Path Forward
The story of nine clients locked into a single SLA clock is a cautionary tale for any team evaluating cloud databases. The write path is the backbone of transactional systems, and sharing it across tenants without isolation is a recipe for unpredictable latency. The industry is moving toward per-tenant consensus, but adoption is slow. In the meantime, engineering teams must be proactive: instrument per-tenant metrics, negotiate per-tenant SLAs, and consider workarounds like dedicated shards or separate clusters.
The hidden tax of shared write paths is real, and it is not going away until vendors change their architectures. But customers have power: by demanding isolation, they can accelerate the shift. The next time you sign a database SLA, remember the nine clients. Ask the four questions. Do not let your writes be slaves to someone else's clock.