Skip to content

Latest commit

 

History

History
243 lines (174 loc) · 15.3 KB

File metadata and controls

243 lines (174 loc) · 15.3 KB
title Errors
description HTTP status codes returned by ScrapeUnblocker and what they mean.

ScrapeUnblocker uses standard HTTP status codes. 2xx means the page was delivered, 4xx points at the request or at the target site, and 5xx means a problem on our side or a request that ran out of time. Only a few responses are billed - they are marked in the table; everything else is free.

Status codes

Code Meaning Billed Where to look
200 The page was delivered. Also used for a dead origin's own 5xx error page - then the X-Origin-Status header carries the target's status. See Dead URLs Yes Response body; X-Origin-Status if present
400 Bad request: the x-scrapeunblocker-key header is missing, the URL is malformed or not http/https, an unsupported proxy_country, or Hostname cannot be resolved (the domain does not exist) Only Hostname cannot be resolved Check the URL and the key header
401 Authentication problem - the key is not recognised, or the account behind it has no valid subscription No See 401 unauthorized below and Authentication
402 Billing problem - quota exceeded, credit limit exceeded, or repeated payment failures No See 402 payment required below
403 Blocked by the target site's bot protection on every available bypass path No Try a different proxy_country, or see handling failures
404 With X-Origin-Status: 404: the target site itself says the page does not exist; the body is the target's page. See Dead URLs. Without the header: no image element found (only on /getImage), a plugin found no such item, or the API path is wrong Only with X-Origin-Status X-Origin-Status header, then the body
408 Browser run timed out (only on /getImage) No Retry
410 The target site itself says the page is gone (X-Origin-Status: 410); the body is the target's page. See Dead URLs Yes Drop the URL
415 The URL is not an HTML document (an image, PDF or file download) No Use /getImage for images
422 Validation error - missing required field or wrong type. With steps, also step_failed when an action could not run. With parsed_data=true, also no_data_extracted: the page rendered but no structured data could be extracted from it (no HTML in the response) Only step_failed (the page was rendered); no_data_extracted is free The response body pinpoints the problem field or step. For no_data_extracted, call without parsed_data to get the HTML
429 Concurrent requests limit exceeded (body Concurrency limit exceeded, header X-SU-Deny-Reason: concurrency) No Wait for a running request to finish, then retry. See concurrency limits
500 Internal error on our side No Retry with backoff
502 Our load balancer lost the connection to an API node (HTML body). Dedicated plugin endpoints also return 502 with a plain-text message when the target returned no usable results No Retry with backoff
503 No API capacity right now (No server is available to handle this request, or Render unavailable.) - a problem on our side, not the target's No Retry with backoff
504 The request reached its time cap: Request timed out. (60 s on /getPageSource, 90 s with steps) or a SERP / plugin timeout. Usually a slow target No Retry. If persistent for one domain, contact support
523 Origin unreachable - the target answered nobody over any of our exits (down, or refusing connections in bursts). Not a bot block No Retry after a pause

Dead URLs: X-Origin-Status

When the target site itself answers that the page does not exist (404 or 410), that answer is the result for your URL - we reached the site, the URL is just dead. ScrapeUnblocker returns the target's own status (404 or 410) with the page exactly as the target served it (the body can be empty; with get_cookies=true it is the usual JSON, and with parsed_data=true the body is {"data": {"page_type": "not_found", "data": {}}} - there is nothing to extract from a not-found page) and adds:

Header Value
X-Origin-Status The target's own status, e.g. 404, 410, 502
X-ScrapeUnblocker-Note A short human-readable explanation

These calls are billed like any delivered page, and retrying will not change the answer. X-Origin-Status is what tells the target's 404 apart from a 404 of our own (a wrong API path, or /getImage finding no image) - branch on it:

r = requests.post(
    "https://api.scrapeunblocker.com/getPageSource",
    params={"url": url},
    headers={"x-scrapeunblocker-key": "YOUR_API_KEY"},
    timeout=125,
)
if r.status_code in (404, 410) and r.headers.get("X-Origin-Status"):
    mark_as_dead(url)  # the page no longer exists on the target site

When the target's origin server returns its own 5xx error page (for example nginx 502 Bad Gateway), you get a 200 with that page, X-Origin-Status carrying the origin's status and X-Upstream-Unavailable: true, also billed. It stays a 200 on purpose: most HTTP clients retry 5xx automatically, and every retry would be billed.

Until 2026-09-30 a target's `404` / `410` also came back as a `200` with `X-Origin-Status`. Code that checks `r.status_code == 200` first should now handle `404` / `410` with the header as a dead URL, not as an error to retry.

A URL whose domain does not exist gets 400 Hostname cannot be resolved instead, also billed.

Is it the target, or is ScrapeUnblocker down?

If you run failover to another provider, only these mean a problem on our side: connection errors, read timeouts, 500, 502, 503, and 504 when it hits many different domains at the same time. All other codes above describe the target site or the request.

  • A call never hangs indefinitely: if our load balancer is unreachable you get a connection error, and if no API node is up you get an immediate 503. A healthy call takes at most its time cap plus up to 60 s of queueing when the whole fleet is busy. Use a connect timeout of about 10 s and a read timeout of about 125 s.
  • Decide on a rate, not on a single response - one failed call is often one difficult domain. For example, switch over when at least half of the calls in the last minute (at least 20 calls, across at least 3 domains) end with one of those signals, and switch back once probe calls succeed.
  • Responses our load balancer generates itself (502, 503, 504) are never billed.

401 unauthorized

A 401 means the request was rejected at our edge, before it ever reached the scraping engine. Because nothing was scraped, a 401 does not count against your quota and is never billed.

Unlike validation errors, 401 responses have a plain-text body (Content-Type: text/plain), not JSON. There are two distinct messages, and they mean different things:

Response body Meaning
Unauthorized The key you sent is not recognised
No valid subscription The key is recognised, but the account behind it has no active subscription

Unauthorized

HTTP/1.1 401 Unauthorized
Content-Type: text/plain

Unauthorized

The value of your x-scrapeunblocker-key header does not match any known key. Common causes:

  • A typo, or a truncated copy-paste. Keys are long; make sure the whole value was copied.
  • Trailing whitespace or a newline in the header value - especially when the key is read from a file rather than an environment variable.
  • An empty header value. Sending x-scrapeunblocker-key with nothing after it counts as an unknown key, not as a missing header.
  • A rotated or revoked key. After you generate a new key in the dashboard, the old one stops working once its short grace period ends.
  • Environment mismatch. Production keys only work against api.scrapeunblocker.com. A key issued for one environment sent to another is an unknown key there.

No valid subscription

HTTP/1.1 401 Unauthorized
Content-Type: text/plain

No valid subscription

The key itself is valid, but the account it belongs to currently has no subscription period covering today - for example the free trial has ended and no plan was chosen, or a plan lapsed and was not renewed. Pick a plan in the dashboard and access resumes within about a minute; no key change is needed.

A billing problem on an **active** subscription returns `402`, not `401` - quota exceeded, credit limit exceeded, or a card that failed repeatedly. `401` is strictly about who you are, `402` about what you owe.

Missing header returns 400, not 401

If you omit the x-scrapeunblocker-key header entirely, the response is 400 Bad Request with the body Missing x-scrapeunblocker-key. This is deliberate: it separates "you forgot to authenticate" from "you authenticated, and it was rejected", so client code can tell a wiring bug from a credential problem.

HTTP/1.1 400 Bad Request
Content-Type: text/plain

Missing x-scrapeunblocker-key

402 payment required

A 402 means your key and account are recognised and in good standing as credentials - the request was stopped for a billing reason. Like 401, it is refused at our edge before anything is scraped, so a 402 consumes no quota and is never billed.

The body is plain text (Content-Type: text/plain), not JSON. There are three messages:

Response body Meaning Fix
Quota exceeded You have used every request your plan allows this billing period Upgrade your plan, or wait for the period to reset
Credit limit exceeded Your unpaid balance has grown past your account's credit limit Pay the outstanding invoice
Payment failed - update payment method A card payment for an open invoice has failed three times in a row Update your card, then pay the invoice

If more than one applies, the most serious wins: payment failure outranks credit limit, which outranks quota.

All three clear themselves. Our load balancer refreshes key statuses about once a minute, so once you upgrade or the invoice is paid, access comes back within roughly a minute - no key change, no support ticket, no redeploy.

Quota exceeded

HTTP/1.1 402 Payment Required
Content-Type: text/plain

Quota exceeded

Your usage for the current billing period has passed your plan's quota. If your plan allows overages, this only fires once you are past quota plus the overage allowance - inside that band requests still succeed and the extra usage is invoiced. Any active coupon credit is spent before plan quota, so a key with remaining credit is never quota-blocked.

The counter resets at the start of your next billing period, which starts on your subscription's anniversary day, not on the first of the month. To get moving sooner, upgrade in the dashboard - the new quota applies on the next status refresh. Current usage against quota is visible in the dashboard, so this is the one 402 you can see coming.

Credit limit exceeded

HTTP/1.1 402 Payment Required
Content-Type: text/plain

Credit limit exceeded

This applies to accounts that accrue usage-based charges. We add up what you currently owe - the amount remaining on your open invoices, plus metered usage already consumed on active subscriptions but not yet invoiced - and compare it against your account's credit limit. Past the limit, the key is paused.

When this triggers we also finalise and attempt payment on the outstanding invoices automatically, so in the common case where your card is good it settles itself and access returns within about a minute. If payment does not go through, pay the invoice from the dashboard. A higher credit limit can be arranged through support.

Payment failed - update payment method

HTTP/1.1 402 Payment Required
Content-Type: text/plain

Payment failed - update payment method

An invoice on your account is open and its payment has been attempted and declined three times. Those attempts are our payment provider's automatic retries spread over several days, so reaching this state means a card has been failing for a while - typically expired, cancelled, or short of funds.

Update your payment method in the dashboard and settle the open invoice. As soon as the invoice is paid the block lifts on the next status refresh, within about a minute.

Subscribing to a new plan does **not** clear this on its own. The old unpaid invoice stays open, so the block stays in place until that specific invoice is paid, even if the new subscription is active and paid for.

402 versus its neighbours

Code It means
401 We do not accept your credentials - unknown key, or no subscription at all. See 401 unauthorized
402 We accept your credentials; the account owes money or is out of quota
429 Nothing is wrong with the account - you have more requests in flight at once than your plan's concurrent requests limit. See concurrency limits

422 validation error shape

When you send an invalid request body, /getPageSource, /serpApi, and /getImage all return a structured validation error:

{
  "detail": [
    {
      "loc": ["query", "url"],
      "msg": "field required",
      "type": "value_error.missing"
    }
  ]
}

loc is the path to the problem field. msg is human-readable. type is a stable machine-readable identifier.

403 - blocked vs. invalid key

A 403 from ScrapeUnblocker never means your API key is wrong. Invalid keys return 401. A 403 always means: the target site blocked us on every bypass route we tried.

When you see 403:

  1. Try a different proxy_country. Some sites geo-fence or geo-rotate their bot protection. A US site may be unreachable from EU IPs and vice versa.
  2. Wait and retry. Rate-based blocks expire after a few minutes.
  3. Contact support if the same URL repeatedly fails - we may need to add a custom plugin for that domain.

More detail in the handling failures guide.

Retries and idempotency

All three endpoints are safe to retry. Requests are idempotent in the sense that retrying with the same parameters does not double-charge or create duplicate state on your account. We recommend exponential backoff for transient 5xx errors and 523 (failed calls are free, so a retry costs nothing extra):

import time
import requests

def fetch_with_retry(url, max_attempts=3):
    for attempt in range(max_attempts):
        r = requests.post(
            "https://api.scrapeunblocker.com/getPageSource",
            params={"url": url},
            headers={"x-scrapeunblocker-key": "YOUR_API_KEY"},
            timeout=125,
        )
        if r.status_code == 200:
            return r
        if r.status_code in (500, 502, 503, 504, 523) and attempt < max_attempts - 1:
            time.sleep(2 ** attempt)
            continue
        r.raise_for_status()
    return r