Skip to content

feat: REST API for external job agents to pull and execute jobs #1153

Description

@jsbroks

Summary

Provide a REST API that lets an external provider act as a job agent — pulling the jobs assigned to it, executing them, and reporting status back — without ctrlplane needing to push/dispatch into the provider's environment.

Background / problem

Today's job agents (e.g. GitHub Actions, Argo Workflows) are push / dispatch-style: ctrlplane initiates the job inside the agent's system. That doesn't fit an external provider that:

  • can't (or shouldn't) be reached inbound by ctrlplane, and
  • wants to integrate generically over HTTP rather than via a bespoke, per-system integration.

There is currently no generic way for such a provider to pull the jobs assigned to its job agent and run them.

Proposed API (sketch)

A job agent authenticates as itself and interacts over REST. Illustrative shape:

  • Acquire work (long-poll): GET /v1/job-agents/{id}/jobs/next
    • Holds the request open until a job is available for this job agent or a timeout elapses (returns empty on timeout).
    • Returns a single job (id, deployment / version / release-target references, config / inputs) and atomically marks it claimed by this agent.
  • Report status / progress: POST /v1/jobs/{jobId}/status
    • { status: running | successful | failed | ..., message?, externalRef?, outputs? }
    • Idempotent on (jobId, status) so retries don't double-apply.
  • Renew claim / heartbeat (if lease-based): POST /v1/jobs/{jobId}/heartbeat
  • Optional logs: POST /v1/jobs/{jobId}/logs

(Paths/shapes are illustrative — the point is the pull + status-report contract.)

Concurrency

Assume a job agent is represented by a single external agent. Even so, two failure modes must be handled:

  1. Double-pickup — a job must be handed out at most once, even under overlapping long-poll requests, client retries, or at-least-once delivery. "Acquire next job" therefore needs to be an atomic, idempotent claim, not a plain read.
  2. Crash mid-job — if the agent dies while holding a claimed job, the job must become reclaimable rather than stuck indefinitely (e.g. a lease/heartbeat or visibility timeout that requeues the job when the claim expires).

The exact claim / reclaim mechanism (lease + heartbeat vs. visibility timeout vs. atomic-claim-only) is left for design discussion.

Open questions

  • Auth / identity — per-job-agent API key or token?
  • Claim mechanism — lease + heartbeat, visibility timeout, or atomic-claim-only?
  • Job payload — how are inputs, outputs, and logs modeled over the API?
  • Backpressure — limits / fairness on long-poll connections.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions