fix: retry ECR push on IAM propagation access-denied errors - #169
Merged
Merged
Conversation
During an initial infra launch, the image_pusher IAM identity is created seconds before the push runs. AWS IAM is eventually consistent, so authenticating with ECR (or the docker push itself) can fail with access-denied until the identity's permissions propagate. Add aws/iampropagation, which retries an operation with exponential backoff for up to 3 minutes when it fails with an access-denied error, logging each attempt to stderr so the user can see we're waiting on IAM propagation and why. Wrap ECR Push and ListArtifactVersions with it, replacing the narrower AccessDeniedException retryer on the ECR client (5 attempts / ~1m window) that the observed failure outlasted.
ssickles
approved these changes
Sep 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
During an initial infra launch/push/deploy, an app can fail to push because the
image_pusherIAM identity is created seconds before the push runs. AWS IAM is eventually consistent, so ECR rejects the freshly-created identity even though its permissions are correctly defined:The existing ECR client retryer (5 attempts, ~1 minute window) retried this exact error but gave up before IAM finished propagating.
Changes
aws/iampropagationpackage:IsPropagationErrorclassifies AWS access-denied errors (smithyAccessDenied/AccessDeniedException/InvalidClientTokenId/UnrecognizedClientException, plus the plain-textdenied: ... not authorized to performerrors surfaced by docker push).Retryretries the whole operation with exponential backoff (2s → 30s) for up to 3 minutes, failing immediately on any other error. Each attempt logs to the user's stderr so they can see we're waiting on IAM propagation and why:Pushretags once, then runs auth + docker push inside the retry loop (covers both aGetAuthorizationTokendenial and a mid-push layer-upload denial).ListArtifactVersionsgets the same wrapper since the launch path calls it in the same race window.AccessDeniedExceptionretryer from the ECR client so retry loops don't nest.Trade-off
A genuinely misconfigured pusher identity now fails after ~3 minutes instead of ~1, but with explicit logging of what it was waiting for, and the final error explains both possibilities.
Follow-up
Port the same wrapper to GAR/ACR/S3/GCS/Blob pushers (same class of propagation race on other clouds).