Introduction
Part 1 was the research and Part 2 was the deployment: fifteen runs of the lza1-RunEngine automation before one came back with exit code 0. Both ended up pointing at the same thing, so here it is.
This is the day-2 post. I wire lza-mcp-server into the container deployment. Then I do the only test that actually means anything: take a real failure from Part 2, point diagnoseDeploymentErrors at it, and see whether it finds in seconds what cost me an afternoon of describe-stack-events archaeology.
It doesn't find it. I still think the server is worth running.
Account IDs and bucket names are redacted below. Error text, tool names, and parameter names are verbatim.
Building the server
Nothing here is surprising, but two details cost me time, so the table first:
| Step | What I did |
|---|---|
| Checkout | git clone https://github.com/awslabs/lza-mcp-server.git, then git checkout v1.1.0 |
| Build directory | Not the repo root. It's src/lza-mcp-server/ (Makefile, Dockerfile, mcp-config-finch.json) |
| Build | make build LZA_MIN_VERSION=v1.16.0 |
| Container runtime | The Makefile auto-detects docker before finch. Docker 29.4.0 here, so no override |
| Image | lza-mcp-server:local, 282 MB |
| Config volume | ~/Developer/lza-config mounted at /app/lza-config |
| Client config | .mcp.json, server key lza |
The README covers the happy path and not much else. To find my way around the code, I leaned on the generated DeepWiki docs, which map out the tool modules and deployment-type detection better than the repo does. Treat them as what they are: AI-generated from the source rather than written by the maintainers, but as an orientation aid, they saved me a lot of file-opening.
The build clones LZA once per version
download-lza-versions.sh walks the GitHub releases API and does a git clone --depth 1 of the LZA repo for every release in the range, pulls source/packages/@aws-accelerator/config/lib/schemas out of it, resolves the $refs, and throws the clone away. The Makefile default is LZA_MIN_VERSION=v1.12.0 with no maximum, which means every release from v1.12.0 to the latest.
I built with LZA_MIN_VERSION=v1.16.0 because my landing zone runs v1.16.2 and the schema tools only need to answer questions about the version that's actually deployed.
The build is not reproducible, and that's the real bug
The image built cleanly and the container started. Then every tool module died on import:
ModuleNotFoundError: No module named 'mcp.server.fastmcp'. This is mcp 2.x, where
FastMCP was renamed to MCPServer (from mcp.server.mcpserver import MCPServer) and
other APIs changed; see the migration guide ... or pin 'mcp<2' to keep running v1 code.The interesting part is where the bad version comes from. pyproject.toml at v1.1.0 declares mcp = {extras = ["cli"], version = ">=1.23.0"}, which is unbounded. There's no committed poetry.lock because make release runs clean-lock and deletes it deliberately. So I fixed the constraint there, rebuilt, and nothing changed.
That file never governs runtime dependencies. The Dockerfile installs a hardcoded list with plain pip, and then installs the application itself with --no-deps:
pip install --no-cache-dir --target /build/deps \
loguru pydantic pyyaml mcp fastapi uvicorn psutil boto3
pip install --no-cache-dir --target /build/deps --no-deps -e .The project's own dependency metadata is bypassed entirely, and mcp resolves to whatever is newest on PyPI at build time. When the MCP Python SDK shipped 2.0 and renamed FastMCP to MCPServer, every fresh build of the v1.1.0 tag started producing a broken image. The code is fine. The build is not.
The fix is PR #2, which pins the Dockerfile line and keeps the cli extra while mirroring the floor from pyproject.toml:
loguru pydantic pyyaml "mcp[cli]>=1.23.0,<2.0.0" fastapi uvicorn psutil boto3Its comment names the scope precisely: roughly 33 modules import mcp.server.fastmcp, so the pin stays until that migration is done. That's what I applied locally.
The general lesson is worth more than the incident. When "build the container yourself" is the only distribution channel, unpinned build dependencies make every user's build different, and a release tag stops meaning anything.
Credentials, and the security flags worth keeping
scripts/extract-aws-credentials.sh is the entry point rather than docker itself. It runs aws configure export-credentials --profile "$AWS_PROFILE" --format env, which evaluates the result, and execs the container command with bare -e AWS_ACCESS_KEY_ID / -e AWS_SECRET_ACCESS_KEY / -e AWS_SESSION_TOKEN / -e AWS_REGION, so the values pass through from the wrapper's environment.
Nothing from ~/.aws is bind-mounted. Personally, I much prefer that to the usual -v ~/.aws:/root/.aws:ro pattern. The container never sees the profile file, the SSO cache, or any other account's credentials, only the resolved temporary keys for one profile.
As a result, the container's credentials are as short-lived as the SSO session. When the profile expires, the tools start failing with credential errors, and the fix is aws sso login plus refreshing the exported profile, then restarting the MCP server. I hit that repeatedly.
The documented docker run line is also more hardened than most MCP server docs bother with, and I kept all of it:
--security-opt=no-new-privileges:true
--cap-drop=ALL
--read-only
--tmpfs /tmp:rw,noexec,nosuid,size=200mRead-only root, a noexec,nosuid tmpfs for scratch, no capabilities, no privilege escalation, and the only writable path is the config bind mount. Good example to hold up when you write your own.
One thing I have not done: iam-policies/ ships three policies, including external-ecs-container-policy.json for exactly this deployment mode. I'm still running with AWSAdministratorAccess, which is precisely what the README's security note tells you not to rely on. Applying the least-privilege policy and reporting what breaks is a better post than showing a setup that only works because it's admin, so that's on my list.
Which account do you point it at?
This one has a clean, demonstrable answer, and it isn't the management account.
From the management account, AWS connectivity succeeds but LZA is invisible:
"deployment": {
"success": false,
"deployment_type": null,
"message": "No LZA deployment detected in region eu-central-1"
}From the LZA Deployment account, the same checkAwsConnectivity call auto-detects everything:
"deployment": {
"success": true,
"deployment_type": "external_pipeline",
"qualifiers": [{ "qualifier": "lza1", "runner": "ecs_container" }],
"message": "External pipeline LZA deployment detected with 1 qualifier(s)"
}Detection is driven by SSM parameters under /accelerator/, which is why the management account comes up empty: the installer's parameters live where the installer stack was deployed. The tool's own troubleshooting text says "for external pipeline deployments: check the pipeline account", which is correct advice you only get to read after the call has already failed.
Two vocabulary notes for anyone searching the docs. Container deployment is classified as external_pipeline with runner: ecs_container, so "external pipeline" is the umbrella here and the ECS container mode is a runner underneath it. Go looking for an ecs_container deployment type and you won't find one.
And prefix and qualifier are two different things, which is the same seam that produced the AWSAccelerator-Suspended-Guardrails incident in Part 2, now showing up as a client-configuration problem instead of a permissions one:
| Value here | Where it comes from | |
|---|---|---|
LZA_PREFIX | AWSAccelerator | Env var in .mcp.json. Drives {prefix}-Pipeline, /accelerator/{prefix}-InstallerStack/version, {prefix}-* stacks |
| Qualifier | lza1 | Not configured at all. Discovered by checkAwsConnectivity, then passed per tool call |
There's no LZA_QUALIFIER env var to set. checkAwsConnectivity returns qualifier_names: ["lza1"] with instructions to ask which qualifier you want and then "pass it to subsequent tools that require it". So the discovery call is the first step of every session, and the two values stay independent inputs throughout.
The day-2 test: diagnoseDeploymentErrors against run 8
Part 2 gave me fifteen runs to choose from. Run 8 is the one that cost the most time: the SSM automation role got an AccessDenied on ecs:DescribeTasks from a service control policy, mid-deployment. If the MCP server can find that, it saves me an afternoon.
I passed run 8's execution ID. It didn't find the failure, and it missed it in four distinct ways.
1. execution_id is ignored for ECS deployments
The schema accepts execution_id. The response described task d59ac7a2bd9747199bb4d22297d2a596, which is run 15, the successful one. The tool description is honest about the behavior ("finds the most recent deployment log stream"). Still, the parameter is accepted without complaint rather than rejected as unsupported, so nothing tells you the answer is about a different run than the one you asked about.
Run 8's log stream still exists. ecs/lza-deployment-container/407bb2f14674432b88e26831665bdfd8 is sitting right there in /ecs/lza1-lza-deployment. The data was available, but the tool has no way to point to it.
Passing run 8's ECS task ID instead of its automation UUID is what gives the game away. The container answer is identical, but the response grows a block it didn't have before:
"pipeline_diagnosis": {
"success": false,
"error": "Invalid execution ID: Execution ID must be a valid UUID format (e.g., 12345678-1234-1234-1234-123456789012)",
"message": "Execution ID does not meet AWS format requirements"
}sources_checked goes from ["ECS container"] to ["CodePipeline", "ECS container"]. So execution_id is consumed only by the CodePipeline branch, where it's validated as a CodePipeline execution UUID. The ECS branch never receives it and picks its log stream by maximum LastEventTime, unconditionally. That explains: a UUID passes the pipeline validator and vanishes, while a task ID, the only identifier that could target an ECS log stream, gets rejected by a validator belonging to the wrong subsystem.
2. It labels a successful run as a failed deployment
The single entry in failed_deployments carries deployment_status: "COMPLETE". Those two things cannot both be true.
3. The error it extracts is a warning
2026-09-13 13:51:23.762 | warn | index | No template found for account <ACCOUNT_ID>
in region eu-central-1. Returning empty template. Error: ValidationError:
Stack with id AWSAccelerator-NetworkVpcStack-<ACCOUNT_ID>-eu-central-1 does not existThat is a normal warn-level line from a healthy first-time deploy. Of course the stack doesn't exist yet, it's about to be created. The extraction looks like it's keying on the substring Error: rather than on log level or outcome.
4. Run 8's failure was never in the container log
This is the fundamental one. The container ran fine. What failed was the SSM Automation step WaitForTaskCompletion, when the automation role hit the SCP deny on ecs:DescribeTasks. diagnoseDeploymentErrors only reads the ECS container log stream, so this entire class of failure - anything in the automation wrapper rather than inside the container - is invisible to it by construction. The response contains no mention of AccessDenied or service control policy anywhere.
So the tool I most wanted after Part 2 could not have helped with the failure that cost the most time, and would have actively misled me. It would have shown me a green run's warning line and called it the error.
To be fair to it: run 8 was arguably outside its remit by design, since the failure happened outside the container. The fair rematch is run 11, which failed inside the container, and I haven't run that one yet.
What upstream looks like
I checked before writing any of this up, because "the tool has a bug" is only worth saying if nobody has said it already.
- There is no
mainbranch. The remote has exactly two,release/v1.0.0andrelease/v1.1.0, with the latter as default and identical to thev1.1.0tag. The public code is exactly what I tested. - The whole issue tracker is two items. #1 is a May 2026 request for CodeConnection support. #2 is the Dockerfile pin.
- PR #2 is still open, filed 2026-08-24 by an outside contributor, two thumbs-up, zero comments, roughly three weeks with no maintainer response. The fix I applied locally exists in no published artifact, so anyone building the documented way today gets a broken image.
- Nothing anywhere mentions
diagnoseDeploymentErrors,execution_idor ECS diagnosis.
One structural detail probably explains the shape of that. The repo carries .gitlab/ and a .gitlab-ci.yml, and its merge commits read "Merge branch '…' into 'integ'". Development happens on an internal GitLab, and GitHub is a mostly read-only mirror. That changes what "open source" means here in practice, because a GitHub PR has no obvious path into a release, which is a plausible reason a one-line build fix has sat for three weeks.
The other day-2 question: what does one change cost?
Reading through startDeployment made me ask something the MCP server can't answer, and it matters more for daily work than any tool in the list. On my CodePipeline landing zone, the smallest config change costs a full 45- to 75-minute cycle. Is the container the same?
Yes, and with less nuance, because there is no change detection anywhere in the container path. container/scripts/run-pipeline.sh hardcodes the list:
AllStacks=( 'key' 'logging' 'organizations' 'security-audit' 'network-prep' 'security' \
'operations' 'network-vpc' 'security-resources' 'identity-center' \
'network-associations' 'customizations' 'finalize' )A plain for loop walks that array after a for Item1 in prepare accounts loop and the bootstrap block. Every stage gets the same three commands: yarn run lza --stage X for the module runner, then cdk.ts synth --stage X, then cdk.ts deploy --stage X. Nothing consults a diff, the config ETag, or CloudFormation drift. The only thing that makes an unchanged stage cheap is CloudFormation's own no-op on an unchanged stack.
Here's run 15, stage by stage, from CloudWatch:
| Stage | Start (UTC) | Elapsed to next |
|---|---|---|
| prepare | 13:48:37 | 1m32s |
| accounts | 13:50:09 | 0m49s |
| bootstrap | 13:50:59 | 0m16s |
| network-vpc (module only) | 13:51:16 | 0m07s |
| key | 13:51:23 | 0m17s |
| logging | 13:51:41 | 1m02s |
| organizations | 13:52:44 | 17m43s |
| security-audit | 14:10:27 | 1m14s |
| network-prep | 14:11:42 | 6m18s |
| security | 14:18:00 | 2m51s |
| operations | 14:20:52 | 39m20s |
| security-resources | 15:00:12 | 5m26s |
| identity-center | 15:05:38 | 1m33s |
| network-associations | 15:07:12 | 3m45s |
| customizations | 15:10:57 | 0m39s |
| finalize | 15:11:36 | 0m19s |
83 minutes, and operations plus organizations are 57 of them.
And every retry starts over at stage 1. Run 8 reached accounts, run 12 reached logging, run 14 reached organizations, run 15 went all the way. The script exits on a stage failure, and the next task re-enters at prepare. There is no resume-from-failed-stage, which is a good part of why Part 2 took an afternoon.
The topology is the real difference
CodePipeline is not linear, and that's what my comparison table in Part 1 didn't capture. Its Deploy stage packs nine actions into one stage separated by runOrder, and Bootstrap fans out fifteen parallel actions:
flowchart TD
SRC[Source] --> PREP[Prepare] --> ACC[Accounts]
ACC --> BOOT["Bootstrap stage<br/>15 parallel actions, runOrder 1"]
BOOT --> REV{{"Review stage<br/>Pre-Approval diff + manual Approve"}}
REV --> KEY["Key<br/>runOrder 1"]
KEY --> LOGG["Logging<br/>runOrder 2"]
LOGG --> ORG[Organization] --> SA[SecurityAudit]
SA --> NP["Network_Prepare<br/>runOrder 1"] & SEC["Security<br/>runOrder 1"] & OPS["Operations<br/>runOrder 1"]
NP & SEC & OPS --> W2{{"all of runOrder 1 complete"}}
W2 --> NV["Network_VPCs<br/>runOrder 2"] & SR["Security_Resources<br/>runOrder 2"] & IC["Identity_Center<br/>runOrder 2"]
NV & SR & IC --> NA["Network_Associations<br/>runOrder 3"]
NA --> CU["Customizations<br/>runOrder 4"] --> FIN["Finalize<br/>runOrder 5"]Everything from Network_Prepare down to Finalize is one CodePipeline stage called Deploy. The runOrder numbers are what serialize it internally, and actions sharing a number run at the same time.
The container collapses all of that into one sequential process:
flowchart TD
SSM["SSM Automation - lza1-RunEngine"] --> RT[RunTask]
RT --> TASK
subgraph TASK["One Fargate task · run-lza.sh → run-pipeline.sh"]
direction TB
Z["download + unzip config zip from S3"] --> BM["bootstrap-management.sh"]
BM --> VC["yarn validate-config"]
VC --> P1
subgraph P1["Phase 1 - for prepare, accounts"]
direction TB
M1["yarn run lza --stage X<br/>(module runner)"] --> S1["cdk.ts synth --stage X"] --> D1["cdk.ts deploy --stage X"]
D1 -.->|next| M1
end
P1 --> BS["Phase 2 - bootstrap<br/>module → synth → cdk bootstrap"]
BS --> P3
subgraph P3["Phase 3 - for AllStacks (13 stages, hardcoded order)"]
direction TB
M2["yarn run lza --stage X"] --> S2["cdk.ts synth --stage X"] --> D2["cdk.ts deploy --stage X"]
D2 -.->|next| M2
end
P3 --> UP["aws s3 sync cdk.out → S3"]
end
TASK --> WAIT["WaitForTaskCompletion (polls up to 12h)"]
WAIT --> EXIT["CheckTaskExitCode"]Same full traversal, but the container gives up the parallelism. CodePipeline runs Network_Prepare, Security, and Operations concurrently. The container runs them back to back. That's where 45 to 75 minutes becomes 88.
Worth stating, because it isn't just a missing UI: there is no Review stage in the container path at all. run-pipeline.sh has no approval gate and no diff step. That's the mechanical reason behind the "no diff" line in the container FAQ I quoted in Part 1.
So can you deploy just one stage?
First a vocabulary correction, because I got this wrong in my own notes: the granularity everywhere is a stage like operations or network-vpc, not a stack. One stage deploys many stacks across accounts and regions.
The script does support partial runs. Line 85 builds SomeStacks=( $stack1 $stack2 … $stack14 ) and line 223 branches on whether it's empty. Set those environment variables, and it deploys only the stages you name. There's a synthOnly=true path too, which synthesizes everything, syncs cdk.out to S3, and exits without deploying.
The problem is that three layers above the script each fail to pass them through:
- The task definition sets 25 environment variables. None of them is
stack1…stack14orsynthOnly. - The SSM document
{qualifier}-RunEnginedeclares six parameters:AutomationAssumeRole,TaskDefinition,Cluster,SubnetId,SecurityGroupId,PythonRuntime- nothing for environment or overrides. - Its RunTask step calls
client.run_task(taskDefinition=…, cluster=…, launchType='FARGATE', networkConfiguration={…})with nooverridesargument. Even a document parameter would have nowhere to go.
Through the supported path, it's all sixteen stages, fixed order, sequential, every time.
Which brings me back to the AWS luminarlz CLI, where I first got excited about running stages in isolation. Its lza stage deploy runs this:
yarn run ts-node --transpile-only cdk.ts deploy --require-approval never --stage <stage> --config-dir <dir> --partition awsThat is the identical command run-pipeline.sh runs, executed on your machine in a temporary LZA checkout at your deployed version, with your own credentials. It never starts a pipeline, never starts a task, never touches SSM. That means it doesn't matter for the deployment mode and would work against a container deployment exactly as it does against my CodePipeline one.
One gap to know about before you rely on it: luminarlz doesn't run the module runner. Its stage deploy is synth plus deploy only, with no yarn run lza --stage X anywhere in the source. The container runs the module runner before every stage, and that's the non-CloudFormation half of the work: creating OUs, inviting and moving accounts, the Control Tower landing zone, registering OUs, root user management, stack policies. Probably irrelevant for operations or network-vpc. Clearly not for prepare or accounts.
So the honest framing isn't "luminarlz is the only way to run one stage". It's that running one stage means running the CDK app yourself, and luminarlz is the ergonomic wrapper for doing that. The container deployment has no in-band answer at all.
Conclusion
My rule of thumb after two days with it: use lza-mcp-server for the surfaces where it's good, and don't retire your CloudWatch tabs yet.
The good surfaces are real. checkAwsConnectivity auto-detecting deployment type, runner, and qualifier in one call is the kind of thing I'd otherwise piece together from three console pages. The schema tools answer config questions against the version you actually run. The container is properly hardened, and the credential handling is better than most.
The diagnosis surface is where I'd set expectations low. On a container deployment, it reads one log stream, the most recent one, and it cannot be pointed anywhere else. It called a COMPLETE run a failed deployment and handed me a warning line as the root cause. If your failure happened in the SSM Automation wrapper rather than inside the container, which, in this deployment mode, is a large and interesting class of failure, it will not see it.
None of that is reported upstream, so I'll file it. Given that PR #2 has sat open for three weeks against a GitHub mirror of an internal GitLab, an issue is probably the more realistic route than a patch.
And if you take one thing from the second half of this post: on container deployment, every change is a full 80-plus minute traversal of all sixteen stages, with no resume and no diff. Budget accordingly, and keep a local checkout handy for times when you only want one stage.
Happy deploying!
Update: Twelve days after this, I decommissioned the whole thing. The order you tear it down in is forced, and getting it wrong is genuinely unrecoverable - I wrote up the deadlock and the full sweep in Decommissioning the Landing Zone Accelerator and Control Tower, Twelve Days Later.
This post was written with AI assistance and verified plus enhanced by a human.



