One Monorepo’s Shared Schema Enum Forced Thirty Teams Into a Single Error String
When thirty teams share a single codebase, every decision about shared abstractions becomes a negotiation. At a large technology company, the decision to centralize error reporting through a single protobuf enum in the monorepo root seemed reasonable at first. It promised consistent error handling across services, simplified client parsing, and aligned with the monorepo's philosophy of shared ownership. But over time, that single enum became a bottleneck, a source of friction, and a lesson in how schema centralization can constrain more than it enables.
A Single Enum Tied Every Error to One String
The monorepo housed roughly thirty teams, each responsible for distinct services—ranging from authentication to billing to content delivery. Early on, the platform team introduced a shared protobuf definition for error codes. The enum lived in a single file under the monorepo root, and a lint rule enforced that every gRPC response used only that canonical enum. No team could define its own error type without modifying the shared file.
The rationale was straightforward: uniformity would simplify client-side error handling. A mobile app could expect the same enum values regardless of which backend service it called. But the reality was that each team had unique error conditions. Billing needed codes for insufficient funds and payment gateway timeouts; content delivery needed codes for cache misses and origin failures. All of them had to fit into the same flat enum namespace.
Debugging became a hunt for a single needle in a haystack. When an error surfaced in production, engineers would grep through logs for the canonical error string, but the context—the specific service, the request parameters, the internal state—was lost. The enum value alone rarely told the full story. Teams started appending free-form strings in the error message field, defeating the purpose of the enum.
No team could extend the enum locally without a cross-team code review and approval from a committee that met biweekly. The process slowed development velocity. A simple addition like a new error for a third-party API timeout could take weeks to land, and teams often resorted to reusing existing codes with ambiguous meanings.
How the Monorepo’s Bazel Graph Enforced Uniformity
The monorepo used Bazel as its build system, and the shared enum file sat near the root of the dependency graph. Every service that depended on the shared proto—and almost all did—had to rebuild when the enum changed. Over time, the build graph expanded as more services added dependencies on the shared proto for other reasons, like common message types.
Changes to the enum required a cross-team code review from at least two representatives of the platform team. The review board met every two weeks, and changes were batched. A team needing a new error code for a critical release often had to wait, or worse, ship with a generic error that masked the real issue. The process created an artificial scarcity of error codes.
Build times grew as the graph expanded. Rough estimates put the overhead at 20–40% per build when the shared proto was touched, because Bazel invalidated caches for all downstream targets. Teams avoided touching the enum unless absolutely necessary, which meant stale or incorrect error codes persisted. The build graph coupling became a self-reinforcing cycle: the more teams depended on the shared proto, the harder it was to change.
Some teams experimented with conditional compilation and build flags to include team-specific error codes, but Bazel's strict visibility rules made it difficult. The shared proto was visible to all targets, but any deviation from the canonical file required either forking the proto (defeating the purpose) or adding a new field to the shared message, which again required committee approval.
The Wire Protocol Became a Single Point of Failure
All gRPC services in the monorepo used the same error proto definition. The enum was embedded in the response message, and clients—both internal and external—parsed the enum to determine error type. A single malformed enum value could break client parsing across the entire ecosystem. For example, if a deprecated value was accidentally reused for a new error, old clients would misinterpret the error.
Backward compatibility mandated that no enum value could be deleted. Deprecated values accumulated over years, cluttering the namespace and making it harder for developers to choose the correct code. Some deprecated codes were still used by legacy clients, so they could not be removed. The enum grew to over 200 values, many with overlapping semantics.
Clients had to handle stale error codes. A mobile app released a year ago might receive an error code that had been deprecated but still sent by an older service. The client code grew branches for every possible code, leading to technical debt. The platform team maintained a mapping of deprecated codes to current ones, but the mapping itself became a source of bugs.
One incident involved a service that mistakenly returned a deprecated enum value intended for internal testing. The client, expecting the new value, fell into an unhandled case and crashed. The fix required a coordinated rollout across all clients, a process that took weeks. The single enum had become a single point of failure.
Team Workarounds: Local Wrappers and String Hacks
As the enum's limitations became apparent, teams began to build workarounds. The most common approach was to wrap the shared enum in a team-specific error struct that added context. The wrapper would include the canonical enum value plus a free-form string and a map of metadata. The gRPC response would still carry the canonical enum, but the team's logging pipeline would parse the metadata for debugging.
Some teams embedded the original error string in the metadata field, effectively duplicating the information. Others used HTTP status codes as an escape hatch, mapping internal errors to standard HTTP codes and relying on clients to interpret the status code instead of the enum. This defeated the purpose of having a canonical enum in the first place.
Logging pipelines became complex. Each team had its own convention for metadata keys: some used JSON, others used key-value pairs separated by pipes. Aggregating error logs across services required custom parsers for each team's format. The platform team tried to enforce a standard metadata schema, but that required another shared proto, leading to the same problems.
Ultimately, the workarounds defeated the schema enforcement that the enum was meant to provide. The uniformity goal was lost, and the cost of maintaining the shared enum was borne by all teams, while the benefits accrued only to the clients that could tolerate the limited error information.
Lessons from Google’s Error Model and gRPC Status
Google’s canonical error model, as used in gRPC, provides a structured approach: a code (enum) plus a human-readable message and optional details. The gRPC status proto allows per-service error details via the details field, which can carry typed protobuf messages. This model acknowledges that a single enum is insufficient for rich error reporting.
Stripe’s API, for example, uses typed error objects with a type enum, a code, a detail message, and additional parameters. The type enum is broad (e.g., card_error, invalid_request_error), while the code provides specific granularity (e.g., card_declined). This two-level approach avoids a flat namespace while still allowing clients to handle errors generically.
The monorepo could have adopted a richer error envelope from the start. Instead of a single enum, the shared proto could have defined a base message with a required code and an optional details field that each team could extend with their own protobuf type. This would have preserved cross-service uniformity while allowing team-specific context.
Migrating to such a model would take roughly 3–6 months, including updating all services and clients. The platform team would need to define the base message, provide migration tooling, and coordinate a phased rollout. The effort is significant, but the alternative—continuing with a constrained enum—incurs ongoing friction and technical debt.
Pragmatic Fix: Split Enum Per Service Domain
One pragmatic fix that emerged was to shard the enum by service ownership. Instead of one flat enum, the monorepo defined multiple enums, each owned by a team or domain. A shared base enum remained for truly cross-cutting errors (e.g., authentication failures, rate limiting), but each service owned its own error space.
The sharding was implemented using Bazel visibility rules. Each team’s enum was placed in a package with restricted visibility, so only the owning service and its immediate clients could depend on it. The shared base enum remained in the root, but its growth was contained. The build graph coupling decreased because changes to a team-specific enum no longer invalidated all services.
Teams could now add error codes without waiting for committee approval. The base enum still required cross-team review, but it changed infrequently. The result was a balance between uniformity and autonomy. Clients that needed to handle errors generically could rely on the base enum; clients that needed specific context could depend on the service-specific enum.
The migration to sharded enums took about four months. The platform team provided a script to extract team-specific codes from the flat enum and generate new proto files. Each team reviewed its own codes and deprecated unused ones. The flat enum was eventually removed, and the monorepo’s error handling became more scalable.
Trade-offs and Counter-arguments: Why Centralization Still Appeals
Despite the pitfalls, centralizing error strings is not without merit. Proponents argue that a single enum forces teams to think about error semantics holistically, preventing a proliferation of nearly identical codes. For example, before centralization, two teams might independently define INVALID_INPUT with subtly different meanings. A shared enum eliminates that duplication and ensures consistent client handling.
Another advantage is discoverability. When all error codes live in one file, engineers can glance at the full set and understand the system's failure modes. With sharded enums, a developer might not know that the billing service has a PAYMENT_TIMEOUT code unless they explicitly look at that team's proto. Documentation becomes more fragmented.
Centralization also simplifies compliance. If the platform team needs to audit error codes for security or regulatory reasons, a single file is easier to inspect than dozens of team-owned files. The trade-off is agility: the shared enum becomes a bottleneck for changes, but for stable, cross-cutting errors, the overhead is minimal.
Some teams argued that the real issue was not centralization per se, but the lack of a mechanism for team-specific extensions. If the shared proto had included a details field from the start, the enum could have remained flat while teams added context. The problem was the design of the error model, not the fact that it was shared.
Yet the experience showed that even with a details field, teams would still need to extend the enum for new error categories. A flat enum cannot capture all dimensions of error semantics. The sharding approach acknowledges that error types are inherently tied to service boundaries and that a one-size-fits-all enum is an illusion.
Alternative Approaches: Error Codes as an API Contract
Another perspective is to treat error codes as part of the service contract, not as a shared schema. In this model, each service defines its own error codes and documents them via protobuf or OpenAPI specifications. Clients are expected to handle errors per service, which may increase client complexity but reduces coupling between services.
This approach is common in microservice architectures where services are independently deployable and owned by separate teams. For example, Amazon’s internal service-oriented architecture encourages each service to define its own error codes, with a common base for generic errors like authentication. The trade-off is that client code must be aware of multiple error namespaces, but the benefit is that teams can evolve their error codes without cross-team coordination.
In the monorepo context, this approach can be implemented by placing each service's error proto in its own package with restricted visibility. Clients that depend on multiple services must import multiple error protos, but Bazel’s dependency analysis ensures that only necessary protos are built. The shared base enum remains for truly cross-cutting errors, but each service owns its specific codes.
Real-World Example: Stripe’s Two-Level Error Model
Stripe’s API error model provides a concrete example of a two-level approach that avoids the pitfalls of a flat enum. Stripe uses a type field (e.g., api_error, card_error, invalid_request_error) and a code field (e.g., card_declined, expired_card). The type is broad enough for client-side generic handling, while the code provides specific details for debugging.
This model allows Stripe to add new codes without breaking existing clients, because clients can fall back to handling the type if they don’t recognize the code. The monorepo could adopt a similar pattern: a shared base enum for broad categories (e.g., AUTH_ERROR, RATE_LIMITED), and a service-specific enum for detailed codes. The base enum would change rarely, while service enums could evolve independently.
Implementing this in a monorepo with Bazel requires careful package organization. The base enum would live in a package with wide visibility, while service enums would be in packages with restricted visibility. Clients that need only the base enum can depend on it alone; clients that need service-specific codes add the relevant package dependency. This reduces build graph coupling while preserving uniformity where it matters most.
Takeaway: Schema Centralization Has a Cost
Shared enums simplify but constrain. The monorepo’s experience shows that a single error string for all services is an anti-pattern. Error strings are API contracts, and centralizing them without room for extension creates friction that teams will work around, often in ways that defeat the original purpose.
Design for per-service extensibility early. A richer error model, like gRPC’s status details or Stripe’s typed errors, allows teams to add context without breaking clients. Monorepos do not mean a single namespace for everything; they can support multiple namespaces with clear ownership boundaries.
The fix—sharding enums by domain—is not a silver bullet. It requires discipline to maintain the base enum and clear communication about which enum to use. But it restores team autonomy and reduces build graph coupling. The lesson is that schema centralization has a cost, and that cost should be weighed against the benefits of uniformity from the start.
For more on how centralized decisions can backfire, see this article on build minute pricing and this one on LDAP config consolidation.