Skip to content

fix: stale-URL teardown cycle, PORT pin, DB_* env-sync guard - #7

Open
RayST3 wants to merge 1 commit into
mainfrom
fix/render-deploy-layer
Open

RayST3 wants to merge 1 commit into
mainfrom
fix/render-deploy-layer

Conversation

@RayST3

@RayST3 RayST3 commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Three gaps found comparing the Render deploy layer against its agentos-railway reference. The portable core is byte-identical to upstream@04744ac; all of this is in the deploy layer.

  1. Teardown poisoned the next deploy. The port dropped upstream's down.sh env-file cleanup, so a torn-down deployment left a dead AGENTOS_URL in .env.production. Render mints a new onrender.com URL on relaunch, and up.sh only pinned when the value was empty, so the new service got no AGENTOS_URL at all: scheduled jobs silently never fire and MCP OAuth advertises a dead origin. env-sync.sh did not rescue it either, since it auto-pins only when the file carries no AGENTOS_URL, so it pushed the corpse.

    Fixed from both ends. down.sh ports comment_out_env_block from upstream and comments out the dead AGENTOS_URL and the JWT_VERIFICATION_KEY (PEM continuation lines included) once both resources are confirmed gone. up.sh compares against the live URL instead of testing for emptiness, normalizing trailing slashes and scoping the staleness test to onrender.com hosts so a deliberate custom domain or tunnel survives.

  2. PORT was left to detection. The Dockerfile CMD hardcodes 8000 with no EXPOSE; Render defaults PORT to 10000 and is only "usually" able to detect a server bound elsewhere, and a failed detection fails the deploy. railway.json had no such exposure, carrying an explicit startCommand instead. render.yaml now declares PORT=8000.

  3. env-sync.sh could repoint production at the wrong database. It skipped only RENDER_, so a DB_ line in an env file (example.env ships them commented) would override the fromDatabase wiring in render.yaml with local/compose values. DB_* is now skipped too, DB_DRIVER included: db/url.py defaults it to postgresql+psycopg, which is what Railway sets explicitly.

README and AGENTS.md updated to match the behavior.

Verified: all four scripts pass bash -n; comment_out_env_block tested against a multi-line PEM fixture and is idempotent; up.sh's env parser reads the cleaned file as unset for both keys; the staleness condition checked across unset/stale/current/trailing-slash/custom-domain/tunnel; render.yaml parses and PORT matches the Dockerfile CMD.

Three gaps found comparing the Render deploy layer against its
agentos-railway reference. The portable core is byte-identical to
upstream@04744ac; all of this is in the deploy layer.

1. Teardown poisoned the next deploy. The port dropped upstream's
   down.sh env-file cleanup, so a torn-down deployment left a dead
   AGENTOS_URL in .env.production. Render mints a new onrender.com URL
   on relaunch, and up.sh only pinned when the value was empty, so the
   new service got no AGENTOS_URL at all: scheduled jobs silently never
   fire and MCP OAuth advertises a dead origin. env-sync.sh did not
   rescue it either, since it auto-pins only when the file carries no
   AGENTOS_URL, so it pushed the corpse.

   Fixed from both ends. down.sh ports comment_out_env_block from
   upstream and comments out the dead AGENTOS_URL and the
   JWT_VERIFICATION_KEY (PEM continuation lines included) once both
   resources are confirmed gone. up.sh compares against the live URL
   instead of testing for emptiness, normalizing trailing slashes and
   scoping the staleness test to onrender.com hosts so a deliberate
   custom domain or tunnel survives.

2. PORT was left to detection. The Dockerfile CMD hardcodes 8000 with
   no EXPOSE; Render defaults PORT to 10000 and is only "usually" able
   to detect a server bound elsewhere, and a failed detection fails the
   deploy. railway.json had no such exposure, carrying an explicit
   startCommand instead. render.yaml now declares PORT=8000.

3. env-sync.sh could repoint production at the wrong database. It
   skipped only RENDER_*, so a DB_* line in an env file (example.env
   ships them commented) would override the fromDatabase wiring in
   render.yaml with local/compose values. DB_* is now skipped too,
   DB_DRIVER included: db/url.py defaults it to postgresql+psycopg,
   which is what Railway sets explicitly.

README and AGENTS.md updated to match the behavior.

Verified: all four scripts pass bash -n; comment_out_env_block tested
against a multi-line PEM fixture and is idempotent; up.sh's env parser
reads the cleaned file as unset for both keys; the staleness condition
checked across unset/stale/current/trailing-slash/custom-domain/tunnel;
render.yaml parses and PORT matches the Dockerfile CMD.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@RayST3
RayST3 requested review from kausmeows and ysolanky and a lite review from Copilot August 28, 2026 05:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR hardens the Render deploy layer to avoid stale environment state across teardown/redeploy cycles, ensuring scheduled jobs and MCP OAuth continue working correctly after a Blueprint relaunch.

Changes:

  • Re-pin AGENTOS_URL in up.sh when the env-file value is stale vs the live Render service URL (with trailing-slash normalization and onrender-scoped staleness logic).
  • Prevent env-sync.sh from pushing local DB_* values (and DB_DRIVER) that would override Render’s fromDatabase wiring.
  • Make down.sh comment out now-dead AGENTOS_URL and multi-line JWT_VERIFICATION_KEY blocks after confirming deletion; pin PORT=8000 in render.yaml; update docs accordingly.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
scripts/render/up.sh Re-pin AGENTOS_URL when stale vs the live service URL to prevent broken scheduler/OAuth after relaunch.
scripts/render/env-sync.sh Skip DB_* keys during env sync to avoid silently pointing production at the wrong database.
scripts/render/down.sh Comment out dead AGENTOS_URL and full multi-line JWT_VERIFICATION_KEY blocks after teardown to prevent poisoning future deploys.
render.yaml Pin PORT=8000 to match the Dockerfile’s hard-coded uvicorn port.
README.md Document DB_* skipping and teardown’s env-file cleanup behavior.
AGENTS.md Update Render blueprint notes to reflect PORT pin and deploy-script behavior.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/render/up.sh
Comment on lines +251 to +252
if [[ -z "$AGENTOS_URL" || ( "${AGENTOS_URL%/}" != "${APP_URL%/}" && "$AGENTOS_URL" == *.onrender.com* ) ]]; then
[[ -n "$AGENTOS_URL" ]] && echo -e "${DIM}AGENTOS_URL pointed at ${AGENTOS_URL} (stale) — re-pinning${NC}"
Comment on lines +124 to +129
DB_HOST | DB_PORT | DB_USER | DB_PASS | DB_DATABASE | DB_DRIVER)
# render.yaml wires these from the managed Postgres (fromDatabase).
# A DB_* line in an env file is the local/compose value; pushing it
# would silently repoint production at a database that isn't there.
echo -e "${DIM} Skipping ${current_key} (wired by render.yaml from agentos-db)${NC}"
;;
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants