Preserve provider instances and expose initialization failure policy - #116
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When one configured provider mounts partially and then fails, its default slot can overwrite a healthy account even though initialization continues. Restore the provider mount table and ordering before handling that failure, include the failed instance ID in the observability event, and enqueue readiness only after remapping succeeds.
Add an optional documented
provider.load_failurecapability for host policy before lifecycle routing. It receives a deep copy of the exact failed provider entry and the original exception, may record or mount an unavailable-instance marker, and may deliberately abort. Without it, the existing warn-and-continue policy remains. Core does not choose replacement accounts, change priorities, prune configuration, or add automatic fallback.Validation: 31 focused Python initialization/multi-instance/observability/readiness tests passed using the modified Python helper over released Core 2.0.1. New rollback and failure-policy cases exercise both MockCoordinator and the real native coordinator, including overwritten/removed mounts, preserved order, import/mount/remap failure, policy abort, and continued healthy loading. New tests pass Ruff; the helper retains six existing lint findings and adds none. Native code and provider protocols are unchanged; no native rebuild, real-provider calls, or Unified deployment was performed.
This is one prerequisite for microsoft/amplifier-unified#182; it does not alone fix the full preparation/routing/host behavior. Rollback covers the mount table, not arbitrary external side effects in third-party mount functions. Configuration and exception data remain private to the callback; hosts must allowlist safe diagnostics.