Skip to content

AWS Well-Architected · Multi-account

Well-Architected, in practice.

The six pillars, mapped to what actually runs in my AWS setup. And the decision that limits the damage when something goes wrong: 6 AWS accounts instead of one.

Well-Architected
Organizations
IAM
CloudTrail

The framework

Six pillars, and how my stack answers them

AWS frames each pillar as questions. These are my answers, from the setup behind my own products.

Operational excellence

Can you run it, change it safely, and learn when something goes wrong?

  • Every resource is Terraform orchestrated by Terragrunt, so changes are reviewed as a plan before they happen.
  • Errors in Lambda logs are forwarded through CloudWatch log subscriptions to an alerting function.
  • An architecture diagram lives next to the code and is updated whenever a unit changes.
  • CloudWatch
  • Lambda
  • Terraform

Security

Who can do what, and what happens if a credential or a component is compromised?

  • Separate AWS accounts for production, development, shared services, audit and log archive, under one Organization.
  • Infrastructure changes run through AWS SSO and assume a dedicated role in the target account. No credentials or account IDs live in the repo.
  • The database sits in private subnets, encrypted with a customer-managed KMS key that rotates automatically.
  • The jumpbox has no public IP and no SSH key: access goes through EC2 Instance Connect Endpoint.
  • API Gateway checks Cognito tokens before any Lambda runs; S3 buckets are private behind CloudFront.
  • Organizations
  • IAM
  • KMS
  • Cognito

Reliability

Does it keep working when parts of it fail, and can it recover?

  • Managed, serverless building blocks (Lambda, API Gateway, SQS) instead of servers to keep alive.
  • Slow or failing work goes through SQS with dead-letter queues, so a broken message is parked, not lost.
  • The VPC spans multiple Availability Zones, and RDS keeps automated backups.
  • Small Terraform state per unit: a failed change to one piece cannot corrupt the state of another.
  • SQS
  • Lambda
  • RDS
  • VPC

Performance efficiency

Are you using the right resources for the job, and do they scale with demand?

  • CloudFront serves static assets from S3 at the edge, compressed; only dynamic requests reach Lambda.
  • Lambda and API Gateway scale with traffic without capacity planning.
  • Scheduled and background work runs off the request path, so user-facing calls stay fast.
  • CloudFront
  • S3
  • API Gateway

Cost optimisation

Are you paying only for what delivers value?

  • Serverless by default: pay per request, nothing idle overnight.
  • A small fck-nat instance instead of a managed NAT Gateway, which has a fixed hourly charge per gateway.
  • Development only holds the resources it actually needs; it does not mirror production.
  • Consolidated billing in the management account, with each account’s spend visible on its own.
  • Lambda
  • EC2
  • Organizations

Sustainability

Are you minimising the resources your workload consumes?

  • No always-on servers for application code: compute runs only when there is work.
  • Edge caching means repeated requests do not repeat the work.
  • The jumpbox runs on arm64 (Graviton), and nothing is provisioned in development “just in case”.
  • CloudFront
  • Lambda
  • EC2
The six pillars at a glance
PillarThe questionOne way my stack answers it
Operational excellenceCan you run it, change it safely, and learn when something goes wrong?Every resource is Terraform orchestrated by Terragrunt, so changes are reviewed as a plan before they happen.
SecurityWho can do what, and what happens if a credential or a component is compromised?Separate AWS accounts for production, development, shared services, audit and log archive, under one Organization.
ReliabilityDoes it keep working when parts of it fail, and can it recover?Managed, serverless building blocks (Lambda, API Gateway, SQS) instead of servers to keep alive.
Performance efficiencyAre you using the right resources for the job, and do they scale with demand?CloudFront serves static assets from S3 at the edge, compressed; only dynamic requests reach Lambda.
Cost optimisationAre you paying only for what delivers value?Serverless by default: pay per request, nothing idle overnight.
SustainabilityAre you minimising the resources your workload consumes?No always-on servers for application code: compute runs only when there is work.

Blast radius

How much breaks when something goes wrong?

Every system eventually has a bad day: a leaked key, a wrong script, a runaway loop. The question is how far it spreads. An AWS account is the hardest boundary AWS gives you for permissions, quotas and billing, so I use 6 of them. Pick an incident and compare.

One AWS account

  • Billing
  • Org policies
  • DNS zones
  • Deploy roles
  • Security findings
  • Audit logs
  • Test auth
  • Experiments
  • Websites
  • APIs
  • Customer database
  • User sign-in
  • Email

13/13

13 of 13 resources inside the blast radius. One account means one boundary. A credential that can touch the test setup can usually reach the customer database, DNS and the logs too.

The Organization

6 accounts, each with one job

Grouped into organizational units so policies apply to a whole group at once.

  • Root

    Management

    Organization settings, guardrail policies and consolidated billing. No workloads.

  • Infrastructure OU

    Shared services

    DNS zones and the cross-account roles other accounts use.

  • Security OU

    Audit

    A separate place to review security findings across every account.

  • Security OU

    Log archive

    Audit logs stored away from the accounts that produce them.

  • Workloads OU

    Development

    Testing changes before they reach real users.

  • Workloads OU

    Production

    Live websites, APIs, authentication, database and email.

One AWS account compared with a multi-account Organization
When…One accountMulti-account
A leaked credentialCan reach everything in the accountLimited to the account it belongs to
Service quotasShared by test and productionSeparate per account
Audit logsDeletable by whoever compromises the accountStored in a separate log archive account
Mistakes in automationOne wrong filter can touch productionBounded by the account the credentials belong to
Cost visibilityDepends on tagging disciplineSpend separated by account by default
GuardrailsIAM policies onlyOrganization-level policies on top of IAM

Running 6 accounts by hand would be its own risk. It works because every account, role and resource is defined in Terraform and Terragrunt, where the folder a unit lives in decides which account it deploys to.

Questions

Well-Architected and multi-account FAQ

What are the six pillars of the AWS Well-Architected Framework?

Operational excellence, security, reliability, performance efficiency, cost optimisation and sustainability. Each pillar is a set of questions and best practices for judging whether a workload is built well, and the pillars often trade off against each other.

What is blast radius in AWS?

Blast radius is how much can be affected when something goes wrong: a leaked credential, a bad deployment, a runaway script. Reducing it means drawing boundaries so a failure in one place cannot spread to everything else. AWS accounts are the strongest of those boundaries.

Why use multiple AWS accounts instead of one?

An AWS account is a hard boundary for permissions, service quotas and billing. Putting production, development, shared services and logs in separate accounts means a credential, a script or a quota problem in one cannot reach the others, and audit logs can be kept out of reach of the accounts that produce them.

Does running multiple AWS accounts cost more?

The accounts themselves are free; you pay for the resources in them. With AWS Organizations, every account rolls up into one consolidated bill. The real cost is setup: roles, DNS and deployments have to work across accounts, which is why I manage it all with Terraform and Terragrunt.

How many AWS accounts does a small team need?

More than one, sooner than most teams expect. A sensible start is a management account with no workloads, separate production and development accounts, and somewhere separate for logs. Shared services and a dedicated audit account can follow as the setup grows.

Want a second pair of eyes on your AWS setup?

Whether it's one account doing everything or a multi-account setup that's grown messy, I'm happy to look at it with you.

This is my interpretation of the AWS Well-Architected Framework applied to my own setup, not an official AWS Well-Architected Review. AWS, the AWS service icons and AWS Well-Architected are trademarks of Amazon.com, Inc. or its affiliates. No endorsement is implied.