TL;DR
- Once an AWS serverless estate outgrows a single CloudFormation or SAM stack — usually somewhere around fifty resources — the split creates a problem nobody warns you about: some resources can only exist once, and they need exactly one owner.
- Two different kinds of "once" are in play. Physically singular things — a DNS record, a custom domain — fail loudly on the second deploy. Logically singular things — an authorizer, a shared layer — duplicate silently and rot for eighteen months. The second kind is the expensive one.
- Make the owner a dedicated stack with a single job: construct the shared resources and hand them out. No business logic, no service functions. It should change about four times a year.
- Choose how services receive those shared values deliberately. Nested stacks give you correct ordering and one blast radius. Exports give independence and a coupling you can't easily undo. SSM Parameter Store gives the loosest coupling and the least protection — usually the right answer past a handful of services.
- One rule keeps the topology readable: services may depend on the shared stack freely; a service depending on another service needs justifying.
- The discipline erodes through parameter lists, not architecture decisions. Audit unused parameters in CI.
You start with one stack. It grows. Somewhere past forty or fifty resources you notice that a one-line change to a single function redeploys the custom domain, the authorizer, every log group and every other function you own — and then takes several minutes to tell you whether it worked.
So you split it. One stack per service: orders, billing, notifications, each with its own template and its own deploy.
The first thing that happens is that all three of them try to create the same DNS record.
That's the fork, and the problem it exposes is this: some things can only exist once. The second you have eight stacks, you have to decide who owns them. Get that decision wrong and you spend the next two years paying for it in ways that are hard to attribute back to the original mistake.
Two kinds of "only once"
Worth separating these, because they fail completely differently and the distinction shapes everything below.
Physically singular. A DNS record for api.example.com. An API Gateway custom domain. An S3 bucket name. A certificate for a given hostname. Two stacks declaring the same one and the second deploy fails — already exists, rollback, someone annoyed for twenty minutes.
That's the good outcome. It's loud, it's immediate, and nobody ships it.
Logically singular. A shared Lambda layer. An authorizer. A log destination. A WAF web ACL. A VPC endpoint.
Nothing stops you creating eight of these. There's no name collision, no quota breach, no error. Eight stacks deploy cleanly, eight authorizers verify tokens, and the system works. You find out eighteen months later, when the token format changes and you update seven of them.
This second category is where the money goes, and the word "shared" is what hides it — it describes an intention, and intentions aren't enforced by anything. The first category enforces itself. The second needs an owner.
The three ways teams get this wrong
Every stack owns its own copy. Each service defines its own authorizer, bundles its own shared code, claims its own subdomain. This is the logically-singular failure in its purest form: it works immediately and it's obviously wrong at any scale. When the token format changes you have eight implementations to update, and you'll discover the ninth in production.
One stack owns everything, including business logic. The "core" stack accumulates the shared resources and the functions that felt too important to put elsewhere. Now your highest-churn code sits in the stack with your highest-blast-radius resources, and every deploy of a shared utility risks your custom domain.
Nobody owns them and they were made by hand. The domain was created in the console in 2023 by someone who has left. It's not in any template. It's not in any repository. It's discovered during an incident.
The third one is the most common, and it's the one that turns a two-hour outage into a two-day one.
What I'd do instead
Have a root stack. Give it exactly one job: construct the things that must exist once, and hand them to everything else.
A word on that name before going further, because CloudFormation overloads it and the ambiguity will otherwise follow you through the rest of this article. In AWS's own vocabulary, a root stack is specifically the top of a nested-stack hierarchy — the one you sam deploy, with AWS::CloudFormation::Stack resources beneath it. I mean something looser: the stack that owns the estate's shared resources. It might be a nested root. It might be a completely independent stack that never nests anything and communicates only through exports or SSM. If "platform stack" or "shared stack" reads better to you, use that — the argument doesn't change.
Whatever you call it: nothing else goes in it. No business logic, no domain functions, no queues that belong to a service. It should be the least interesting file in the repository, and it should change roughly four times a year.
Concretely, it owns:
- The custom domain and its DNS record
- Shared Lambda layers
- A shared authorizer, if you have one
- Cross-cutting log and trace destinations
- VPC endpoints and shared security groups
- Anything with a name that must be globally unique
And in practice it's dull enough to read in one sitting:
# platform/template.yaml — resources and outputs, nothing else
Resources:
ApiDomain:
Type: AWS::ApiGateway::DomainName
Properties:
DomainName: !Sub "api-${Environment}.example.com"
RegionalCertificateArn: !Ref ApiCertificate
SharedLayer:
Type: AWS::Serverless::LayerVersion
Properties:
ContentUri: layers/shared/
CompatibleRuntimes: [nodejs22.x]
TokenAuthorizerFunction:
Type: AWS::Serverless::Function
Properties:
CodeUri: authorizer/
Handler: index.handlerThree resources, no handlers, no routes, no queues. If reviewing a change to this file makes anyone nervous, something is in it that shouldn't be.
This is the composition root idea from dependency injection, applied to infrastructure. In an application, the composition root is the single place that assembles the object graph — the one file allowed to call new, so that everything else receives its dependencies instead of constructing them. Here the root stack is the one template allowed to create a shared resource; every service template can only receive one.
The value isn't the wiring. It's the constraint. Once shared resources have exactly one owner, "who's allowed to create this?" has exactly one answer, and the logically-singular category stops being able to duplicate silently.
flowchart TD
P[Root stack<br/>domain · layer · authorizer]
P -.ARNs / parameter paths.-> O[Orders]
P -.-> B[Billing]
P -.-> N[Notifications]
B ==>|service-to-service:<br/>the edge that needs justifying| OThe property this buys you is worth naming, because it's the one you'll actually feel: you can understand any single service by reading its own template plus the block that feeds it. You never hold the whole estate in your head to reason about one part of it. That's the difference between an estate a new engineer can join and one where only two people know how anything connects.
Now the hard part: how do services get the shared values?
This is where the real decision is, and it's usually made by accident. There are three mechanisms, and they trade off along the same axis: how tightly do you want deployment coupled?
Nested stacks
The root declares each service as an AWS::Serverless::Application (or AWS::CloudFormation::Stack) and passes shared values down as parameters.
OrdersService:
Type: AWS::Serverless::Application
Properties:
Location: services/orders/template.yaml
Parameters:
AuthorizerArn: !GetAtt TokenAuthorizerFunction.Arn
SharedLayerArn: !Ref SharedLayerCloudFormation understands the whole graph and orders it correctly. One sam deploy, one changeset, one rollback — nobody has to know the deploy order, because nobody is choosing it.
That's genuinely valuable, and it takes away independence completely. A one-line change to one service redeploys the root and everything under it; blast radius is the whole estate. A failure in the ninth nested stack rolls back the other eight. And you'll eventually hit the 500-resource-per-stack ceiling, which arrives faster than you'd think when every function brings a role, a log group and a permission.
Use nested stacks when the estate is small enough that "deploy everything" is acceptable, and when you'd rather have coordination handled for you than have independence.
Exports and Fn::ImportValue
The root exports values; services import them by name.
# root stack
Outputs:
TokenAuthorizerArn:
Value: !GetAtt TokenAuthorizerFunction.Arn
Export:
Name: platform-TokenAuthorizerArn
# orders stack, deployed entirely separately
Parameters:
AuthorizerArn:
Default: !ImportValue platform-TokenAuthorizerArnIndependent stacks, but CloudFormation enforces the dependency: you cannot change or delete an exported value while anything imports it. That sounds like safety — and it is, right up until you need to replace the authorizer and discover you must remove every consumer first, in order, across twelve repositories.
Two constraints worth knowing before you commit to this. Export names must be unique per account per region, so the naming scheme is an estate-wide decision on day one. And exports don't cross a region or an account boundary at all — if a disaster-recovery region is anywhere in your future, exports are the mechanism that won't come with you.
Exports are excellent for values that will genuinely never change. They are a trap for anything you might want to swap.
Indirection through SSM Parameter Store
The root writes ARNs to well-known parameter paths. Services read them at deploy time.
# root stack
AuthorizerArnParam:
Type: AWS::SSM::Parameter
Properties:
Name: !Sub "/platform/${Environment}/authorizer-arn"
Type: String
Value: !GetAtt TokenAuthorizerFunction.Arn
# orders stack
AuthorizerArn: "{{resolve:ssm:/platform/prod/authorizer-arn}}"Loosest coupling, least protection. Nothing stops you deleting a parameter something depends on. But it's the only mechanism that lets you replace a shared resource without touching consumers: write the new ARN to the same parameter path, and services pick it up on their next deploy.
You can claw back some of the safety you gave up by pinning a version — {{resolve:ssm:/platform/prod/authorizer-arn:7}} resolves that specific parameter version rather than latest. That converts "picks up changes automatically" into "picks up changes when someone edits the number", which is occasionally exactly what you want for a high-blast-radius value. Same caveat as any pin: record why, next to the pin.
Choosing
| Mechanism | Deploy independence | What enforces correctness | Replacing the shared resource |
|---|---|---|---|
| Nested stacks | None — one deploy, one rollback | CloudFormation orders the graph | One changeset |
| Exports | Full | CloudFormation refuses to break an import | Unwind every consumer first, in order |
| SSM parameter | Full | Nothing | Write the new ARN to the same path |
For estates past a handful of services, the right-hand column is usually what decides it, and that's where I land on SSM. The coupling becomes a contract on a name rather than a hard link — with the honest cost that you now need discipline and monitoring where CloudFormation was previously giving you a guarantee.
One thing worth knowing regardless of which you pick: {{resolve:ssm:...}} resolves at deploy time, so the value is baked into the deployed resource. Reading the parameter at runtime instead gives you the ability to change a dependency without redeploying — and buys you a per-invocation API call, a new failure mode, and a caching problem. Deploy-time resolution is the right default. Runtime resolution is for values that genuinely change without a deployment.
The rule I'd write on the wall
Services may depend on the root freely. A service depending on another service requires a conversation.
That sounds bureaucratic. It isn't — it's the single rule that keeps the topology comprehensible, because service-to-service infrastructure dependencies are cheap to add and brutally expensive to unwind.
It takes one line to make the billing stack read a queue name from the orders stack. That one line converts a star into a graph. Now those two stacks have a deployment order, can't be torn down independently, and the next engineer has to know about the relationship to reason about either one. Add four more of those over a year and you've rebuilt the monolith you split up, except now it's distributed and the coupling isn't in a file anyone reads.
Sometimes it's genuinely necessary. When it is, make it explicit and visible — a declared dependency the tooling understands beats a hardcoded ARN or a name assembled from a convention. A convention-built name (${env}-orders-events) avoids the deployment dependency, which looks like a win, right until someone renames the queue and nothing tells them what broke. Pick the failure you'd rather have. I'll take the deployment constraint, because it's visible before the change ships.
Where this discipline actually erodes
Not in a big architectural decision. In parameter lists.
Adding a parameter to a service is a two-line change in two files. Removing one requires proving nothing uses it. So parameter lists ratchet — they only grow. Give it eighteen months and you'll have services receiving values they never reference, and each unused parameter is a small lie: it tells the next reader that this service participates in something it doesn't.
Two habits fix this cheaply:
Audit unused parameters. A script that parses each template, collects declared parameters, and greps for references catches every one of them in seconds. Run it in CI.
Assume the plumbing is wrong and check it. The classic failure is a service template growing a required parameter that its parent never passes. Nothing catches this — not your editor, not the build, not review, because nobody compares a parameter list against a template they don't have open. It fails during the deploy, which in CloudFormation means it fails during the rollback. It's also entirely statically detectable, and worth a couple of hundred lines of script.
The test
Open your root template. If you can't describe what it owns in one sentence, it owns too much.
The root stack is infrastructure for your infrastructure. It should be short, stable, and dull enough that nobody has strong feelings about it. If it's the file people are nervous to change, the problem isn't the file — it's that you've put things in it that belong somewhere they can fail on their own.