One Account, One Branch, Real Customers: Bringing a Hand-Built Production Stack into Terraform
One Account, One Branch, Real Customers: Bringing a Hand-Built Production Stack into Terraform
In my last post about console-created resources, the fix was unglamorous: delete the hand-made resources in pre-prod and let Terraform create them fresh. That works when the thing you're deleting doesn't matter yet.
This one mattered.
A client brought me on as a contractor with a clear brief: get everything managed with infrastructure as code. Then came the part that shaped every decision afterward. They had one environment. One AWS account, one domain, one git branch, and everything deployed from that branch into that account. It had all been built in the console, and it had customers on it.
So "dev" was production. There was nothing else.
What was in the account
A serverless backend that had grown one click at a time:
- About 50 Lambda functions, Python, most of them talking to MySQL
- 38 API Gateway REST APIs and a handful of Lambda Function URLs
- RDS MySQL
- A video delivery path: S3, CloudFront, a Lambda@Edge resolver, and a WAF
- SES identities, a nightly scheduled sync job, log groups, IAM roles, the VPC around all of it
The goal was three environments in parity, dev, stage, and prod, each in its own account, all deployed from CI. The order mattered. You can't build stage and prod from code you don't have yet, and the only accurate description of the system was the live account itself. So step one was to make Terraform describe what already existed, exactly, before anything else got built.
The rule: import, don't apply
If you point a fresh Terraform root at a populated account and run apply, you get one of two outcomes. Either it fails on name collisions, or it succeeds and replaces something customers are using. Neither is a good first day.
So the method was terraform import and terraform state mv only, run deliberately, environment by environment. The finish line wasn't "apply went green." It was a plan with zero imports, zero moves, and nothing in it I hadn't meant to put there.
The first full plan against the live account was nowhere near that:
Plan: 324 to import, 124 to add, 216 to change, 78 to destroy.
That plan isn't something you apply. It's a to-do list. The imports were the work. Every "change" was a place where my code didn't match reality yet, and the fix was the code, not the account. Every "destroy" was either a resource I had modeled wrong or a resource that existed in the account and not in my code. All of them had to be worked down before anything ran for real.
Keep the scaffolding deletable
Terraform's import {} blocks are great and also clutter. Once a resource is in state, the block does nothing, but it keeps sitting next to the real configuration and confusing the next person to read the module.
So everything that only existed to adopt resources went into its own files:
api.imports.tf
function_urls.imports.tf
lambda.imports.tf
log_groups.imports.tf
schedule.imports.tf
config/<env>.imports.tfvars
Each one started with the same banner:
# IMPORT SCAFFOLDING: delete this file as a set with the other *.imports.tf
# files, config/*.imports.tfvars, and moved.tf once every environment has
# adopted its resources. Nothing here describes infrastructure.
The import-only variables were loaded by CI and a local helper script only when the file existed, so deleting them later would need no other change.
One small thing paid for itself quickly. Thirty-eight APIs times their routes and methods is a lot of import addresses. Writing them out by hand made the tfvars file 386 lines long. Expressing it as one line per API in a map, and letting the .imports.tf file expand that back into the full import set, got it down to 48 lines that a human could actually review.
The gotcha: your first import plan may print your passwords
This is the part I'd tell anyone doing this first.
The Lambda module ignored changes to environment. That's a reasonable choice while you're adopting, because it means Terraform won't try to rewrite variables it doesn't own yet. But when you import a Lambda, Terraform reads the live function from AWS and renders it in the plan. All of it. Including environment variables. Including database passwords.
sensitive() can't help. Terraform only masks values it manages as sensitive, and these were values it was deliberately ignoring.
The plan log went to CI, where anyone with access to the pipeline could read it.
The first fix was a blunt one: pipe plan and apply output through a small script that replaces the value of any key matching *PASSWORD*, *SECRET*, *TOKEN*, *API_KEY*, or *PRIVATE_KEY*, and use set -o pipefail so Terraform's exit code still fails the job.
The better fix came a few days later. Import those functions through the version of the module that does manage environment, with the values marked sensitive. The same imports then print (sensitive value), and the redaction script got deleted.
If you take one thing from this post: before you share the link to your first import pipeline, read its log yourself.
The only changes to the live account were deliberate
Import-only doesn't mean nothing in the live account ever changed. It means nothing changed by accident. The changes that did land were the ones the project existed for:
- Tags. Every Terraform-managed resource got
ManagedByandTerraformStatetags through providerdefault_tags, so anyone in the console can see what owns a resource and which state file to look in. - Private databases. The databases were in a public subnet. They moved to a private one.
- A bastion for database access. With the databases private, people still needed a way in. An SSM Session Manager bastion turned access into an IAM permission instead of an IP allowlist entry: granting is a policy attachment, revoking is immediate, and every session shows up in CloudTrail. A small tunnel script opens the port-forward so a database client can connect to
localhost. - A stronger WAF in front of the video distribution.
Each one went through the same plan review as everything else: the only adds and changes in the plan were the ones I was there to make.
One bastion lesson: if your AMI comes from the "latest Amazon Linux" SSM parameter, every new Amazon release shows up in your plan as a bastion replacement. ignore_changes = [ami] keeps the instance on its image until you decide to replace it.
Then build the environments that didn't exist
With the live account described in code, it became the dev environment in the repo. Stage and prod were new AWS accounts and new git branches, built fresh from the same root.
Prod was deliberately not a copy of how the original account had grown up:
- Multi-AZ RDS, deletion protection, 14-day backups, a final snapshot on delete, private subnets
- Secrets Manager containers created by Terraform, with the values written separately, so no credential ever lands in the state bucket
- Before the first apply, a check that the state bucket and CI deployer role existed, no RDS instances were present, and none of the few Lambdas already in the account collided with the names Terraform was about to create
Cutting over: separate infrastructure from code
For the cutover, I split the prod pipeline in two. A push to the prod branch planned and applied Terraform automatically, but the Lambda code deploy waited for a manual run.
deploy:lambdas:prod:
stage: deploy
needs: [publish:lambdas]
when: manual
resource_group: prod-lambdas
That let prod take infrastructure changes on its own schedule while a customer-facing code change waited for sign-off. Infrastructure first, then code, never both by surprise.
Once prod was live, CI went back to its steady state: dev applies and deploys on push, the nightly job moves from the original account to prod and is switched off in dev, and prod deploys automatically like stage.
Cleanup: delete the scaffolding
The last commit of the adoption deleted every *.imports.tf file, the imports tfvars, moved.tf, and the CI plumbing that passed them in. Its message is the scoreboard:
- dev (the original account): 332 imports (38 REST APIs, 4 Function URLs, 8 log groups, the nightly schedule, 3 Lambdas), with no apply
- prod: 10 state moves and 3 log-group imports
- stage: nothing to adopt
All three environments now plan 0 to import, 0 moved. The scaffolding did its job and left.
What I'd tell the next person
- Find out which environment is really production before you touch anything. Here it was easy, because there was only one. Often it isn't.
- Import and move state; don't apply your way to parity. The finish line is a plan with zero imports, zero moves, and only the changes you intended.
- Put import scaffolding in files you can delete as a set, then delete them.
- Read your first import plan's log for secrets. Ignored attributes are still printed.
- Make the live changes deliberately, and only the ones that justify themselves: tags, private databases, a bastion, a better WAF.
- Build the new environments fresh from the same code, and split infrastructure apply from code deploy for the cutover.
Last time, Terraform met reality and reality lost. This time reality had customers on it, so Terraform had to learn to describe it first.