From b21cbf4a797dd546c19575161616e5e02977db39 Mon Sep 17 00:00:00 2001 From: Nikita Date: Wed, 2 Sep 2026 18:26:09 +0200 Subject: [PATCH 1/2] PRDCT-676: give the Generic Extractor tutorial a reading order, and classify all 39 pages MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The tutorial is a real seven-part progression built on one API, but every part ended mid-thought: nothing told the reader where to go next, and only the index carried a list. Each part now ends with a Next pointer following the nav order — REST → JSON → basic configuration → pagination → jobs → mapping — and the last one hands off to the configuration reference rather than dead-ending. The index keeps its existing Next Steps list. Every one of the 39 pages also gets a type marker, the convention already used in src/content/docs/cli/. The tree had none. Types: tutorial for the seven tutorial pages, explanation for the hub and the parameter map, how-to for running, publishing, incremental and the SSH proxy, reference for the rest. The markers deliberately do NOT claim a verification date. Nothing here was fact-checked against keboola/generic-extractor — this is a form pass — so each marker says so and points at PRDCT-676. A later accuracy pass can grep for them and replace the note with a real source and date. Form only: no page created, split, moved or deleted, no nav change, no prose rewritten. Co-Authored-By: Claude Fable 5 --- .../configuration/api/authentication/api_key/index.md | 2 ++ .../configuration/api/authentication/basic/index.md | 2 ++ .../configuration/api/authentication/bearer_token/index.md | 2 ++ .../configuration/api/authentication/index.md | 2 ++ .../configuration/api/authentication/login/index.md | 2 ++ .../configuration/api/authentication/oauth10/index.md | 2 ++ .../configuration/api/authentication/oauth20-login/index.md | 2 ++ .../configuration/api/authentication/oauth20/index.md | 2 ++ .../configuration/api/authentication/oauth_cc/index.md | 2 ++ .../configuration/api/authentication/query/index.md | 2 ++ .../extractors/generic-extractor/configuration/api/index.md | 2 ++ .../configuration/api/pagination/cursor/index.md | 2 ++ .../generic-extractor/configuration/api/pagination/index.md | 2 ++ .../configuration/api/pagination/multiple/index.md | 2 ++ .../configuration/api/pagination/offset/index.md | 2 ++ .../configuration/api/pagination/pagenum/index.md | 2 ++ .../configuration/api/pagination/response-param/index.md | 2 ++ .../configuration/api/pagination/response-url/index.md | 2 ++ .../generic-extractor/configuration/aws-signature/index.md | 2 ++ .../generic-extractor/configuration/config/index.md | 2 ++ .../configuration/config/jobs/children/index.md | 2 ++ .../generic-extractor/configuration/config/jobs/index.md | 2 ++ .../configuration/config/mappings/index.md | 2 ++ .../extractors/generic-extractor/configuration/index.md | 2 ++ .../generic-extractor/configuration/iterations/index.md | 2 ++ .../generic-extractor/configuration/ssh-proxy/index.md | 2 ++ .../extractors/generic-extractor/functions/index.md | 2 ++ .../extractors/generic-extractor/incremental/index.md | 2 ++ .../docs/components/extractors/generic-extractor/index.md | 2 ++ .../components/extractors/generic-extractor/map/index.md | 2 ++ .../extractors/generic-extractor/publish/index.md | 2 ++ .../extractors/generic-extractor/running/index.md | 2 ++ .../extractors/generic-extractor/tutorial/basic/index.md | 4 ++++ .../extractors/generic-extractor/tutorial/index.md | 6 +++++- .../extractors/generic-extractor/tutorial/jobs/index.md | 4 ++++ .../extractors/generic-extractor/tutorial/json/index.md | 4 ++++ .../extractors/generic-extractor/tutorial/mapping/index.md | 4 ++++ .../generic-extractor/tutorial/pagination/index.md | 4 ++++ .../extractors/generic-extractor/tutorial/rest/index.md | 3 +++ 39 files changed, 92 insertions(+), 1 deletion(-) diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md index f5fb13933..c444c5ca7 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/api_key/ --- + + API Key token authentication sends a token in either a header or query parameter of each API request. E.g., Headers: `X-StorageApi-Token:your_token` diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md index 2c7b22fe1..39500843e 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/basic/ --- + + Basic Authentication provides the [HTTP Basic Authentication](https://en.wikipedia.org/wiki/Basic_access_authentication) method. It requires entering a username and password in the configuration and sends the encoded values in the `Authorization` header. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md index 1acedad01..11f249d25 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/bearer_token/ --- + + Bearer token authentication sends a token in the `Authorization` header of each API request. This method is available through UI. You can select the `Bearer Token` method and fill in the token. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md index 279485536..c28d29241 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various configuration options.* diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md index 449e9b2ac..e583a75f2 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/login/ --- + + Use the Login authentication to send a one-time **login request** to obtain temporary credentials for authentication of all the other API requests. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md index 0e57c4dd6..140d10b7b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth10/ --- + + **Note** that this configuration option is not yet supported and the test endpoint button will not work. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md index adde56120..52f354468 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth20-login/ --- + + **Note** that this configuration option is not yet supported and the test endpoint button will not work. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md index 9eac3d715..731ad2ea5 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth20/ --- + + OAuth 2.0 Authentication is one of [two OAuth methods](/components/extractors/generic-extractor/configuration/api/authentication/#oauth) and is supported only for [components registered in the developer portal](/components/extractors/generic-extractor/publish/). diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md index d1b35ce6b..ca073322a 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth_cc/ --- + + oAuth 2.0 Client Credentials authentication performs the [oAuth 2.0 client_credentials flow](https://auth0.com/docs/get-started/authentication-and-authorization-flow/client-credentials-flow). This method is available through the UI and is implemented via the [Login](/components/extractors/generic-extractor/configuration/api/authentication/login/) method. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md index 561d1706d..63bb699b0 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/query/ --- + + Query Authentication provides the simplest authentication method, in which the credentials are sent in the [request URL](/components/extractors/generic-extractor/tutorial/rest#url). This method is most often used with APIs that authenticate using API tokens and diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md index 5d79e3ea1..49c5d935b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/basic/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md index 2f4a753c2..c85a9e08d 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/cursor/ --- + + The Cursor Scroller can be used with an API which expects the client to maintain a cursor (pointer) to the last obtained item. For example, on the first request, it returns items with ID 1-100; for the second diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md index 150d6d576..44fdefe68 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/ --- + + *If new to Generic Extractor, learn about [pagination in our tutorial](/components/extractors/generic-extractor/tutorial/pagination/) first.* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various configuration options.* diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md index 25e3b5065..317035077 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/multiple/ --- + + Setting the pagination method to `multiple` allows you to use **multiple scrollers on a single API**. This type of pagination contains the definition of all scrollers used in the entire configuration. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md index 40d10725f..42675a7fa 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/offset/ --- + + The Offset scroller handles a pagination strategy in which the API splits the results into pages of the same size (limit parameter) and navigates through them using the **item offset** parameter. This diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md index b9730b3f1..48f60cb66 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/pagenum/ --- + + The Page Number Scroller handles a pagination strategy in which the API splits the results into pages of the same size (limit parameter) and navigates through them using the **page offset** parameter. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md index 4cda14f9c..7b306dfba 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/response-param/ --- + + The Response Parameter Scroller can be used with APIs that provide a certain kind of value in the response which must be used in the next request. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md index 61d01e6fd..8d70127ad 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/response-url/ --- + + The Response URL Scroller can be used with APIs that provide the URL of the next page in the response. This scroller is suitable for APIs supporting the diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md index ca5d7f1b3..442b8ee07 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/aws-signature/ --- + + Generic Extractor allows signing requests with [**AWS Signature Version 4**](https://docs.aws.amazon.com/general/latest/gr/signature-version-4.html). Signing is the process of adding authentication information to your requests. When you use AWS tools, the extractor signs your API request. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md index 97e8a7c7e..2206c380b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/config/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md index ba498fb5a..2ebe8311f 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/config/jobs/children/ --- + + *If new to Generic Extractor, learn about [jobs in our tutorial](/components/extractors/generic-extractor/tutorial/jobs/) first.* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md index 758b88b4d..0dd71fbf3 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/config/jobs/ --- + + *If new to Generic Extractor, learn about [jobs in our tutorial](/components/extractors/generic-extractor/tutorial/jobs/) first.* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/mappings/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/mappings/index.md index 6d78deefa..da015e8df 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/mappings/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/mappings/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/config/mappings/ --- + + *If you are new to Generic Extractor, learn about [mapping in our tutorial](/components/extractors/generic-extractor/tutorial/mapping/) first.* *Use the [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various configuration options.* diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/index.md index 795f146b5..d3f2e859b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md index dac9a6a1d..5bc7baef0 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md @@ -6,6 +6,8 @@ redirect_from: - /extend/generic-extractor/iterations/ --- + + The `iterations` section allows you to **execute a configuration multiple times, each time with different values**. The most typical use for `iterations` is extraction of the same data from multiple accounts. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md index 6be7eabff..99a3ecb0e 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/ssh-proxy/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various diff --git a/src/content/docs/components/extractors/generic-extractor/functions/index.md b/src/content/docs/components/extractors/generic-extractor/functions/index.md index f427526ec..bc1d98b2a 100644 --- a/src/content/docs/components/extractors/generic-extractor/functions/index.md +++ b/src/content/docs/components/extractors/generic-extractor/functions/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/functions/ --- + + Functions are simple pre-defined functions that diff --git a/src/content/docs/components/extractors/generic-extractor/incremental/index.md b/src/content/docs/components/extractors/generic-extractor/incremental/index.md index dd391bafc..b358125d1 100644 --- a/src/content/docs/components/extractors/generic-extractor/incremental/index.md +++ b/src/content/docs/components/extractors/generic-extractor/incremental/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/incremental/ --- + + Extracting data incrementally is universally beneficial — it **speeds up the extraction** and **lowers the load** on both the API and [Keboola Storage](/storage/) (thus saving diff --git a/src/content/docs/components/extractors/generic-extractor/index.md b/src/content/docs/components/extractors/generic-extractor/index.md index 2a7a3650a..1707536a3 100644 --- a/src/content/docs/components/extractors/generic-extractor/index.md +++ b/src/content/docs/components/extractors/generic-extractor/index.md @@ -6,6 +6,8 @@ redirect_from: - /components/extractors/other/generic/ --- + + Generic Extractor is a [Keboola component](/overview/) that acts like a customizable [HTTP REST](/components/extractors/generic-extractor/tutorial/rest/) client. It can be configured to extract data diff --git a/src/content/docs/components/extractors/generic-extractor/map/index.md b/src/content/docs/components/extractors/generic-extractor/map/index.md index fb6995bb3..8ecc86d7d 100644 --- a/src/content/docs/components/extractors/generic-extractor/map/index.md +++ b/src/content/docs/components/extractors/generic-extractor/map/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/map/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* Use the following sample configuration to navigate among various **configuration options**: diff --git a/src/content/docs/components/extractors/generic-extractor/publish/index.md b/src/content/docs/components/extractors/generic-extractor/publish/index.md index c5e80f6b1..d019e3593 100644 --- a/src/content/docs/components/extractors/generic-extractor/publish/index.md +++ b/src/content/docs/components/extractors/generic-extractor/publish/index.md @@ -6,6 +6,8 @@ redirect_from: - /extend/generic-extractor/registration/ --- + + It is possible to publish a Generic Extractor configuration as a completely separate component. This enables sharing the API extractor between various projects and simplifies its further configuration. diff --git a/src/content/docs/components/extractors/generic-extractor/running/index.md b/src/content/docs/components/extractors/generic-extractor/running/index.md index 7127e7300..e0b4a2ce2 100644 --- a/src/content/docs/components/extractors/generic-extractor/running/index.md +++ b/src/content/docs/components/extractors/generic-extractor/running/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/running/ --- + + Generic Extractor is normally run from within the Keboola user interface. It can be found in the **Extractors** section and all you need to do is provide its configuration JSON. No other settings are necessary. diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md index 5bb512ea7..5714ae4bb 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/basic/ --- + + Before configuring Generic Extractor, you should have a basic understanding of [REST API](/components/extractors/generic-extractor/tutorial/rest/) and @@ -216,3 +218,5 @@ of doing much more; see other parts of this tutorial for an explanation of pagin (resources) to be extracted. - [Mapping](/components/extractors/generic-extractor/tutorial/mapping/) — describes how the JSON response is converted into CSV files that will be imported into Storage. + +**Next:** [Pagination →](/components/extractors/generic-extractor/tutorial/pagination/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/index.md index 945d4e0e2..f517f5c39 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/ --- + + In this tutorial, we will guide you through configuring Generic Extractor for a new API. In our case, MailChimp — an email marketing service. @@ -79,4 +81,6 @@ configuration here: - [Jobs](/components/extractors/generic-extractor/tutorial/jobs/) — describes the API endpoints (resources) to be extracted. - [Mapping](/components/extractors/generic-extractor/tutorial/mapping/) — describes how the JSON - response is converted into CSV files that will be imported into Storage. \ No newline at end of file + response is converted into CSV files that will be imported into Storage. + +**Next:** [REST HTTP API introduction →](/components/extractors/generic-extractor/tutorial/rest/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md index 2aae6e239..76571cfdf 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/jobs/ --- + + On your way through the Generic Extractor tutorial, you have learned about @@ -238,3 +240,5 @@ MailChimp API and giving us a lot of trouble, is best to be ignored. The answer You might also have noticed some duplicate records in the table `in.c-ge-tutorial.campaigns__campaign_id__content` along the way. You'll look into this as well. + +**Next:** [Mapping →](/components/extractors/generic-extractor/tutorial/mapping/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md index 13f3e7dd1..4c0143205 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/json/ --- + + [JSON (JavaScript Object Notation)](http://www.json.org/) is an easy-to-work-with format for describing structured data. Before you start working with JSON, familiarize yourself with basic programming jargon. It is also recommended @@ -145,3 +147,5 @@ The order of items in an object is not important. It is also worth noting that ` ## Summary This page contains a little introduction to JSON documents. We intentionally avoided many details, but you should now understand what JSON is, and how to write some stuff in it. + +**Next:** [Basic configuration →](/components/extractors/generic-extractor/tutorial/basic/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md index d7df96615..a7ba1aaf1 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/mapping/ --- + + In the previous part of the tutorial, you [extracted the content of a MailChimp campaign](/components/extractors/generic-extractor/tutorial/jobs/). Now, it's time to clean up the response. @@ -350,3 +352,5 @@ The key of the mapping supports dot notation to traverse into children. So, if t ``` As you changed the delimiter from the default `.` to `/`, it's no longer parsed as two separate keys `created` and `date`, but rather just a single key `created.date`. + +**Next:** [Configuration reference →](/components/extractors/generic-extractor/configuration/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md index 34084748e..a5f47a450 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/pagination/ --- + + Pagination breaks a result with a large number of items into separate pages and is used very commonly in many API calls. @@ -165,3 +167,5 @@ getting incomplete data. The next two parts of our tutorial deal with setting up (resources) to be extracted. - [Mapping](/components/extractors/generic-extractor/tutorial/mapping/) — describes how the JSON response is converted into CSV files that will be imported into Storage. + +**Next:** [Jobs →](/components/extractors/generic-extractor/tutorial/jobs/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md index 1225930dc..26fd5cb0b 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/rest/ --- + + An [API (Application Programming Interface)](https://en.wikipedia.org/wiki/Application_programming_interface) is an [interface](https://en.wikipedia.org/wiki/Interface_(computing)) to an application, or a **service** @@ -134,3 +136,4 @@ to get responses from virtually any HTTP REST API. Since the REST rules are not is not possible to ensure that Generic Extractor will be capable of reading 100% of APIs, even when declared as RESTful by someone. +**Next:** [JSON introduction →](/components/extractors/generic-extractor/tutorial/json/) From 1896028cee81f11b50c52a5d120fd4ff82152ab7 Mon Sep 17 00:00:00 2001 From: Nikita Date: Thu, 3 Sep 2026 01:34:35 +0200 Subject: [PATCH 2/2] PRDCT-676: sentence-case the Generic Extractor headings MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit House style is sentence case; this tree was Title Case throughout, inherited from the dev-docs original. 176 headings across 33 pages. Case-only, and verified to be so: the built anchor ids are byte-identical before and after — 476 ids, none added, none removed — because Starlight slugifies through github-slugger, which lowercases anyway. No link, in this repo or in the keboola org, can break on this. Left capitalised: product names (Generic Extractor, Keboola), acronyms (API, URL, JSON, OAuth, HTTP, SSH, AWS, CSV), scroller and function names that are identifiers rather than prose (Has-More, StrToTime), and the first word after a step number. Co-Authored-By: Claude Fable 5 --- .../api/authentication/basic/index.md | 6 +-- .../api/authentication/login/index.md | 20 ++++----- .../api/authentication/oauth20-login/index.md | 6 +-- .../api/authentication/oauth20/index.md | 6 +-- .../api/authentication/oauth_cc/index.md | 2 +- .../api/authentication/query/index.md | 8 ++-- .../configuration/api/index.md | 26 ++++++------ .../api/pagination/cursor/index.md | 8 ++-- .../configuration/api/pagination/index.md | 22 +++++----- .../api/pagination/multiple/index.md | 2 +- .../api/pagination/offset/index.md | 10 ++--- .../api/pagination/pagenum/index.md | 10 ++--- .../api/pagination/response-param/index.md | 10 ++--- .../api/pagination/response-url/index.md | 10 ++--- .../configuration/aws-signature/index.md | 2 +- .../configuration/config/index.md | 8 ++-- .../config/jobs/children/index.md | 30 ++++++------- .../configuration/config/jobs/index.md | 28 ++++++------- .../configuration/config/mappings/index.md | 20 ++++----- .../generic-extractor/configuration/index.md | 6 +-- .../configuration/iterations/index.md | 4 +- .../configuration/ssh-proxy/index.md | 6 +-- .../generic-extractor/functions/index.md | 42 +++++++++---------- .../generic-extractor/incremental/index.md | 8 ++-- .../extractors/generic-extractor/index.md | 10 ++--- .../generic-extractor/publish/index.md | 2 +- .../generic-extractor/running/index.md | 10 ++--- .../generic-extractor/tutorial/basic/index.md | 6 +-- .../generic-extractor/tutorial/index.md | 4 +- .../generic-extractor/tutorial/jobs/index.md | 4 +- .../generic-extractor/tutorial/json/index.md | 4 +- .../tutorial/mapping/index.md | 6 +-- .../generic-extractor/tutorial/rest/index.md | 6 +-- 33 files changed, 176 insertions(+), 176 deletions(-) diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md index 39500843e..0765dc228 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md @@ -11,7 +11,7 @@ Basic Authentication provides the [HTTP Basic Authentication](https://en.wikiped method. It requires entering a username and password in the configuration and sends the encoded values in the `Authorization` header. -### User Interface +### User interface In the user interface, you simply select the `Basic Authorization` method and enter the username and password. @@ -41,11 +41,11 @@ They are also prefixed by the hash `#` character, which means they are stored [e If the API expects something else than a username and password in the `Authorization` header, or if it requires a custom authorization header, use the [Default Headers option](/components/extractors/generic-extractor/configuration/api/#headers). -## Configuration Parameters +## Configuration parameters This `basic` type of authentication has no configuration parameters. The login and password must be provided in the [`config` section](/components/extractors/generic-extractor/configuration/config/) of the Generic Extractor configuration. -## Basic Configuration Example +## Basic configuration example Assume you have an API which requires you to use the HTTP Basic authentication to send the login and password in the `Authorization` header. Assume that your login is `JohnDo` and password is `secret`. The following configuration solves the situation: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md index e583a75f2..f449e3d1b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md @@ -11,7 +11,7 @@ redirect_from: Use the Login authentication to send a one-time **login request** to obtain temporary credentials for authentication of all the other API requests. -## User Interface +## User interface Note that this configuration option is not yet covered. You can add the JSON configuration using the `Custom` auth method. @@ -53,7 +53,7 @@ A sample Login authentication looks like this: } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `login` type of authentication: - `loginRequest` (required, object) — a [job-like](/components/extractors/generic-extractor/configuration/config/jobs/) object describing the login request; it has the following properties: @@ -79,7 +79,7 @@ is called only once before all other requests. To call the login request before ## Examples Below are several examples showing you how to use various login authentication related features in Generic Extractor. -### Configuration with Headers +### Configuration with headers Let's say you have an API which requires every API call to be authorized with the `X-ApiToken` header. The value of that header (an API token) is obtained by calling the `/login` endpoint with the headers `X-Login` and `X-Password`. The `/login` endpoint response looks like this: @@ -130,7 +130,7 @@ will contain the header: See [example [EX079]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/079-login-auth-headers). -### Configuration with Headers and Text Response +### Configuration with headers and text response Let's say you have an API like the above, but it returns the login response as a plain text: a1b2c3d435f6 @@ -183,7 +183,7 @@ will contain the header: See [example [EX128]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/128-login-auth-text). -### Configuration with Query Parameters +### Configuration with query parameters Let's say you have an API which requires an [HTTP POST](https://en.wikipedia.org/wiki/POST_(HTTP)) request with `username` and `password` to the endpoint `/login/form`. On a successful login, it returns the following response: @@ -246,7 +246,7 @@ so the second API call will be sent as: See [example [EX080]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/080-login-auth-query). Notice that the example uses completely different URL for the login request. -### Parameter Overriding +### Parameter overriding The above examples show how to use query parameters and headers separately. However, they can be mixed freely; they can also be mixed with parameters and headers entered elsewhere in the configuration. The following example shows how parameters from different places are merged together: @@ -407,7 +407,7 @@ This causes Generic Extractor to call the **login request** every hour. See [example [EX082]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/082-login-auth-expires). -### Expiration from Response +### Expiration from response In case the credentials provided by the **login request** have a time-limited validity, use the `expires` option. If the validity of the credentials is returned in the response, modify the [first example](#configuration-with-headers) to this: @@ -455,7 +455,7 @@ This assumes that the response of the **login request** looks like this: See [example [EX083]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/083-login-auth-expires-date). -### Relative Expiration from Response +### Relative expiration from response In case the API returns credentials validity in the **login request** and that validity is expressed in seconds, use the `expires` option together with setting `relative` to `true`. The result is the behavior of the [first example](#expiration-basic) but the value is taken @@ -506,7 +506,7 @@ This assumes that the response of the **login request** looks like this: See [example [EX084]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/084-login-auth-expires-seconds). -### Login Authentication with Functions +### Login authentication with functions Suppose you have an API which requires you to send a username and password separated by a colon and base64 encoded — for example, `JohnDoe:TopSecret` (base64 encoded to `Sm9obkRvZTpUb3BTZWNyZXQ=`) in the `X-Authorization` header to an `/auth` endpoint. The login endpoint then returns a token @@ -570,7 +570,7 @@ uses the `login` authorization method to send them to the special `/auth` endpoi See [example [EX100]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/100-function-login-headers). -### Login Authentication with Login and API Request +### Login authentication with login and API request Suppose you have an API similar to the one in the [previous example](#login-authentication-with-functions). It requires you to send a username and password separated by a colon and base64 encoded — for example, `JohnDoe:TopSecret` (base64 encoded to `Sm9obkRvZTpUb3BTZWNyZXQ=`) in the diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md index 52f354468..2f30c38e0 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md @@ -42,7 +42,7 @@ for authentication of all the other API requests. A sample OAuth Login authentic } ``` -## Configuration Parameters +## Configuration parameters The configuration parameters are identical to the [Login](/components/extractors/generic-extractor/configuration/api/authentication/login/) method. The difference, however, is in the [function context](/components/extractors/generic-extractor/functions/#oauth-20-login-authentication-context). The **login request** is assumed to require the OAuth2 authorization and its response must be in JSON format (plaintext is not supported). @@ -50,7 +50,7 @@ The **login request** is assumed to require the OAuth2 authorization and its res ## Examples The following examples demonstrate how to use OAuth with a basic login request and Google API in Generic Extractor. -### Basic Configuration +### Basic configuration The following configuration shows how to set up an OAuth **login request**: ```json @@ -133,7 +133,7 @@ and sent to other API requests (`/users`). See [example [EX105]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/105-oauth2-login). -### Google API Configuration +### Google API configuration The following example shows how to set up the OAuth authentication for Google APIs. The access token is refreshed with each API call. #### Generate access tokens diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md index 731ad2ea5..da346eedd 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md @@ -78,7 +78,7 @@ Note that the properties `appKey` and `#appSecret` must exist even if not used b to empty strings. For more information about OAuth 2, see the [official documentation](https://oauth.net/2/) or learn [more about Keboola-OAuth integration](/extend/common-interface/oauth). -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `oauth20` authentication type: - `format` (optional, string) — If the OAuth service provider response is JSON, use the only possible @@ -94,7 +94,7 @@ are available in the [OAuth function context](/components/extractors/generic-ext ## Examples The following two examples demonstrate the support for OAuth 2 in Generic Extractor. -### Bearer Authentication +### Bearer authentication The most basic OAuth authentication method is with "Bearer Token". If you have an API which supports this authentication method, the following configuration can be used: @@ -146,7 +146,7 @@ the header `Authorization: Bearer SomeToken1234abcd567ef` using the See [example [EX103]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/103-oauth2-bearer). -### HMAC Authentication +### HMAC authentication If you have an API which requires an [HMAC](https://en.wikipedia.org/wiki/Hash-based_message_authentication_code) signed token, generate the correct signature using [functions](/components/extractors/generic-extractor/functions). The following example assumes you obtain the following response from the API upon authentication: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md index ca073322a..3f3f2d79b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md @@ -13,7 +13,7 @@ This method is available through the UI and is implemented via the [Login](/comp ![img.png](/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc.png) -### Configuration Parameters +### Configuration parameters - `Login Request type` - `Basic Auth`: The client_id and client_secret are sent in the Authorization header as a Basic authorization, e.g. `Authorization: Basic base64(client_id:client_secret)`. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md index 63bb699b0..7efef1c17 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md @@ -13,7 +13,7 @@ This method is most often used with APIs that authenticate using API tokens and signatures. Dynamic values of query parameters can be generated using [user functions](/components/extractors/generic-extractor/functions/). -## User Interface +## User interface In the user interface, you simply select the `Query` method and enter the key-value pairs of the query parameters. ![img.png](/components/extractors/generic-extractor/configuration/api/authentication/query.png) @@ -38,12 +38,12 @@ A sample Query authentication configuration looks like this: } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `query` type of authentication: - `query` (required, object): An object whose properties represent key-value pairs of the URL query. -## Basic Configuration Example +## Basic configuration example Let's say you have an API that requires an `api-token` parameter (with value 2267709) to be sent with each request. The following authentication configuration does exactly that: @@ -62,7 +62,7 @@ configuration remains organized. See [example [EX077]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/077-query-auth). -## Configuration With Encrypted Token Example +## Configuration with encrypted token example Usually, you want the value used for authentication to be encrypted (the `api-token` parameter with the value 2267709 in our example), so you do not expose it to other users or store it in the configuration versions history. The following authentication configuration, combined with the parameter defined in the [`config`](/components/extractors/generic-extractor/configuration/config/) section, does that (the value with the prefix `#` is encrypted upon saving the configuration): diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md index 49c5d935b..ea1caa408 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md @@ -94,7 +94,7 @@ Authentication (authorization) needs to be configured for any API which is not p Because there are many authorization methods used by different APIs, there are also many [configuration options](/components/extractors/generic-extractor/configuration/api/authentication/). -## Retry Configuration +## Retry configuration By default, Generic Extractor **automatically retries failed HTTP requests** — repeatedly, and on most errors. This is one of the big advantages over writing your own extractor from scratch. Tweak the retry setting to optimize the speed of an extraction or to avoid unwanted flooding of the API. @@ -126,7 +126,7 @@ There are two retry strategies: - Either the API sends a `Retry-After` header (or its equivalent), or - Generic Extractor uses an [exponential backoff algorithm](https://en.wikipedia.org/wiki/Exponential_backoff). -### API Retry Strategy +### API retry strategy Per the HTTP specification, the API may send the [`Retry-After`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Retry-After) header which should contain number of seconds to pause/sleep before the next request. Generic Extractor supports some extensions to this. First, the *Retry Header* name may be customized. Second, the header @@ -140,7 +140,7 @@ time of the next request The second and third options are often called **Rate Limit Reset** as they describe when the next successful request can be made (i.e., the limit is reset). -### Backoff Strategy +### Backoff strategy The exponential backoff in Generic Extractor is defined as `truncate(2^(retry\_number - 1)) * 1000` seconds. This means that the first retry (zero-based index) will be after 0 seconds (`(2^(0-1)) = 0.5`, truncated to 0). The retry delays are the following: @@ -181,7 +181,7 @@ in the [debug](/components/extractors/generic-extractor/running/#debug-mode) mes If the exponential backoff is used, you will see its sequence of times. See an [example](/components/extractors/generic-extractor/configuration/api/#retry-configuration). -## Default HTTP Options +## Default HTTP options The `http` configuration option allows you to set the timeouts, default headers and parameters sent with each API call (defined later in the [`jobs` section](/components/extractors/generic-extractor/configuration/config/jobs/#request-parameters)). @@ -201,7 +201,7 @@ the headers and values are their values — for instance: See the full [example](/components/extractors/generic-extractor/configuration/api/#default-headers). -### Request Parameters +### Request parameters The `http.defaultOptions.params` configuration allows you to **set the [request parameters](/components/extractors/generic-extractor/tutorial/rest/#url) to be sent with each API request**. The same rules apply as to the @@ -209,7 +209,7 @@ sent with each API request**. The same rules apply as to the See an [example](/components/extractors/generic-extractor/configuration/api/#default-headers). -### Required Headers +### Required headers Similar to the `http.headers` option, the `http.requiredHeaders` option allows you to **set the HTTP header for every API request**. The difference is that the `requiredHeaders` configuration specifies **only the header names**. The actual values must be provided in the [`config`](/components/extractors/generic-extractor/configuration/config/) @@ -240,7 +240,7 @@ Failing to provide the header values in the `config` section will cause an error See the full [example](/components/extractors/generic-extractor/configuration/api/#required-headers). -### Ignore Errors +### Ignore errors The `ignoreErrors` option allows you to force Generic Extractor to ignore certain extraction errors. The option lists HTTP codes for which any errors occurring during downloading and JSON parsing the response will be ignored. The `ignoreErrors` option error is an array of HTTP @@ -274,7 +274,7 @@ API implementations and should not be used blindly if other solutions may be app [`responseFilter`](/components/extractors/generic-extractor/configuration/config/jobs/#response-filter). When ignoring errors, **you might miss even those errors that require your attention.** -### Connect Timeout +### Connect timeout The `connectTimeout` option is a float describing the number of seconds to wait while trying to connect to a server. Default value is `30` seconds. Use `0` to wait indefinitely, we do not recommend it. @@ -285,7 +285,7 @@ Default value is `30` seconds. Use `0` to wait indefinitely, we do not recommend } ``` -### Request Timeout +### Request timeout The `requestTimeout` option is a float describing the total timeout of the request in seconds. Default value is `300` seconds. Use `0` to wait indefinitely, we do not recommend it. @@ -298,7 +298,7 @@ Default value is `300` seconds. Use `0` to wait indefinitely, we do not recommen ## Examples -### Retry Configuration +### Retry configuration Assume that you have an API which implements throttling in the following way: when the number of requests is exceeded, it returns an empty response with the status code `202` and a timestamp when a new requests can be made in the `X-RetryAfter` HTTP header. @@ -323,7 +323,7 @@ Notice that it is necessary to add the response code `202` to the existing defau See [example [EX037]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/037-retry-header). -### Default Headers +### Default headers Assume that you have an API which returns a JSON response only if the client sends an `Accept: application/json` header. Additionally, if the client sends an `Accept-Encoding: gzip` header, the HTTP transmission will be compressed (and thus faster). @@ -343,7 +343,7 @@ The following configuration sends both headers with every API request: See [example [EX038]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/038-default-headers). -### Default Parameters +### Default parameters Assume that you have an API requiring all requests to contain a filter for the account to which they belong. This is done by passing the `account=XXX` parameter. The following configuration sends the parameter with every API request: @@ -366,7 +366,7 @@ may also be used. See [example [EX039]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/039-default-parameters). -### Required Headers +### Required headers Assume that an API requires the header `X-AppKey` to be sent with each API request. The following API configuration can be used: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md index c85a9e08d..446779c91 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md @@ -26,7 +26,7 @@ request, you must tell the API to start with ID 101. } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `cursor` method of pagination: - `idKey` (required, string) — path to the key which contains the value of the cursor; the path is entered relative to the exported items. @@ -43,7 +43,7 @@ The request parameter specified in the `param` configuration overwrites the para [job parameters](/components/extractors/generic-extractor/configuration/config/jobs/#request-parameters). Other job parameters are carried over without modification (see an [example](#reverse-configuration)). -### Stopping Condition +### Stopping condition The pagination ends **when the `dataField` of the response contains no items**. Because of this, each run with the `cursor` scroller produces a similar warning: @@ -54,7 +54,7 @@ This is expected behavior. [Common stopping conditions](/components/extractors/g ## Examples This section contains two API pagination examples where the Cursor Scroller is used. -### Basic Configuration +### Basic configuration Let's say you have an API which has an endpoint `/users` returning the following response: ```json @@ -91,7 +91,7 @@ Notice that the `idKey` parameter is relative to the extracted array of items (` See [example [EX060]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/060-pagination-cursor-basic). -### Reverse Configuration +### Reverse configuration Some APIs return items starting with the newest item and therefore need to be queried for offset in reverse order. Let's say a request to `/users?startWith=last` will produce: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md index 44fdefe68..e08b43ec4 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md @@ -43,7 +43,7 @@ An example pagination configuration looks like this: } ``` -## Paging Strategy +## Paging strategy Generic Extractor supports the following paging strategies (scrollers); they are configured using the `method` option: @@ -54,7 +54,7 @@ using the `method` option: - [`cursor`](/components/extractors/generic-extractor/configuration/api/pagination/cursor/) — uses the identifier of the item in response to maintain a scrolling cursor. - [`multiple`](/components/extractors/generic-extractor/configuration/api/pagination/multiple/) — allows to set different scrollers for different API endpoints. -### Choosing Paging Strategy +### Choosing paging strategy If the API responses contain direct links to the next set of results, use the [`response.url` method](/components/extractors/generic-extractor/configuration/api/pagination/response-url/). This applies to the APIs following the [JSON API specification](https://jsonapi.org/). The response usually @@ -102,7 +102,7 @@ If the API uses different paging methods for different endpoints, use the [`multiple` method](/components/extractors/generic-extractor/configuration/api/pagination/multiple/) together with any of the above methods. -## Stopping Strategy +## Stopping strategy Generic Extractor stops scrolling - based on the `nextPageFlag` condition configuration. @@ -137,7 +137,7 @@ the first page, it is not same as the previous page and therefore another reques is the same as the previous page, the same check kicks in and the extraction is stopped too. However, the results from the first page will be duplicated. -### Next Page Flag +### Next page flag The above describes automatic behavior of Generic Extractor regarding scrolling stopping. Using **Next Page Flag** allows you to do a **manual setup of the stopping strategy**: Generic Extractor analyzes the response, looks for a particular field (the flag) and decides whether to continue scrolling based on the value or presence of that flag. @@ -169,7 +169,7 @@ Example `nextPageFlag` setting: See our [Next Page Flag Examples](#next-page-flag-examples). -### Force Stop +### Force stop Force stop configuration allows you to stop scrolling when some extraction limits are hit. The supported options are: @@ -234,7 +234,7 @@ and [example [EX116]](https://github.com/keboola/generic-extractor/tree/master/d and [example [EX140]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/140-pagination-forcestop-child-filter) (combining with child jobs). -### Limit Stop +### Limit stop Limit stop configuration allows you to stop scrolling when a specified number of items is extracted. The supported options are: @@ -283,7 +283,7 @@ See [example [EX126]](https://github.com/keboola/generic-extractor/tree/master/d For `count` configuration, see [example [EX127]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/127-pagination-stop-field) (a modified version of [EX051](https://github.com/keboola/generic-extractor/tree/master/doc/examples/051-pagination-pagenum-basic). -### Combining Multiple Stopping Strategies +### Combining multiple stopping strategies All stopping strategies are evaluated simultaneously and for the scrolling to continue, none of the stopping conditions must be met. In other words, the scrolling continues until any of the stopping conditions is true. To this you need to account specific stopping strategies for @@ -312,14 +312,14 @@ will stop if **any** of the following is true: - The `isLast` field is present in the response and is true (`nextPageFlag`). - The `isLast` field is not present in the response. -## Next Page Flag Examples +## Next page flag examples In this section, we want to show you the following examples of the Next Page Flag stopping strategy: - Has-More Scrolling - Non-Boolean Has-More Scrolling - Is-Last Scrolling -### Has-More Scrolling +### Has-More scrolling Assume that the API returns a response which contains a `hasMore` field. The field is present in every response and has always the value `true` except for the last response where it is `false`. The following pagination configuration can be used to configure the stopping strategy: @@ -341,7 +341,7 @@ In this case, setting `ifNotSet` is not necessary. See [example [EX045]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/045-next-page-flag-has-more) and [example [EX139]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/139-pagination-hasmore-child-filter) (combining with child jobs). -### Non-Boolean Has-More Scrolling +### Non-Boolean Has-More scrolling Assume that the API returns a response which contains a `hasMore` field. The field is present only in the last response and has the value `"no"` there. The following pagination configuration can be used to configure the stopping strategy: @@ -365,7 +365,7 @@ to false. In this case setting `ifNotSet` is mandatory. See [example [EX046]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/046-next-page-flag-has-more-2). -### Is-Last Scrolling +### Is-Last scrolling Assume that the API returns a response which contains an `isLast` field. The field is present only in the last response and has the value `true` there. The following pagination configuration can be used to configure the stopping strategy: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md index 317035077..9b222657a 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md @@ -54,7 +54,7 @@ The name of the scroller must be used in a specific [job `scroller` parameter](/ A `default` scroller can be set (must be one of the names defined in `scrollers`). In that case, all jobs without an assigned scroller will use the default one. -### Stopping Condition +### Stopping condition There are no specific stopping conditions for the `multiple` pagination. Each scroller acts upon its normal stopping conditions. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md index 42675a7fa..775b84189 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md @@ -29,7 +29,7 @@ An example configuration: } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `offset` method of pagination: - `limit` (required, integer) — page size @@ -45,7 +45,7 @@ the [job parameters](/components/extractors/generic-extractor/configuration/conf 100 items at most and you set the limit 1000, it would cause the extraction to stop after the first page. This is because the [underflow condition](/components/extractors/generic-extractor/configuration/api/pagination/#stopping-strategy) would be triggered. -### Stopping Condition +### Stopping condition Scrolling is stopped **when the result contains less items than requested** — specified in the `limit` configuration (underflow). This also includes an instance when no items are returned, or the response is empty. @@ -88,7 +88,7 @@ All [common stopping conditions](/components/extractors/generic-extractor/config ## Examples This section contains three examples of API pagination using the Offset Scroller. -### Basic Scrolling +### Basic scrolling This is the simplest scrolling setup: ```json @@ -103,7 +103,7 @@ The next request has `limit=20` and `offset=20`, for example, `/users?limit=20&o See [example [EX043]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/043-paging-stop-underflow) and [example [EX044] with a more structured response](https://github.com/keboola/generic-extractor/tree/master/doc/examples/044-paging-stop-underflow-struct). -### Renaming Parameters +### Renaming parameters The `limitParam` and `offsetParam` configuration options allow you to rename the limit and offset for the needs of a specific API: @@ -121,7 +121,7 @@ and `skip=0`, for example, `/users?count=2&skip=0`. See [example [EX049]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/049-pagination-offset-rename). -### Overriding Limit and Offset +### Overriding limit and offset It is possible to override both the limit and offset parameters of a specific API job. This is useful in case you want to use different limits for different API endpoints. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md index 48f60cb66..2bd1c3089 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md @@ -26,7 +26,7 @@ If you need to use the item offset, use the [Offset Scroller](/components/extrac } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `pagenum` method of pagination: - `limit` (optional, integer) — page size @@ -36,7 +36,7 @@ The following configuration parameters are supported for the `pagenum` method of value is `true`. - `firstPage` (optional, integer) — index of the first page; the default value is `1`. -### Stopping Condition +### Stopping condition The `pagenum` scroller uses similar stopping condition as the [`offset` scroller](/components/extractors/generic-extractor/configuration/api/pagination/offset/#stopping-condition). Scrolling is stopped in case of an underflow — when the result contains **less items than requested** (including zero). However, in the `pagenum` scroller, the **`limit` parameter is not required** and has **no default value**. This means that if you omit it, @@ -45,7 +45,7 @@ the scrolling will stop only if an empty page is encountered. ## Examples This section contains three API pagination examples where the Page Number Scroller is used. -### Basic Scrolling +### Basic scrolling The most simple scrolling setup is the following: ```json @@ -59,7 +59,7 @@ The next request will have `page=2`, for example `/users?page=2`. See [example [EX051]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/051-pagination-pagenum-basic). -### Renaming Parameters +### Renaming parameters The `limitParam` and `pageParam` configuration options allow you to rename the limit and offset for the needs of a specific API: @@ -80,7 +80,7 @@ and `set=1`; for example, `/users?set=1&count=20`. See [example [EX052]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/052-pagination-pagenum-rename). -### Overriding Parameters +### Overriding parameters It is possible to override the limit parameter of a specific API job. This is useful when you want to use different limits for different API endpoints. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md index 7b306dfba..60b5e359e 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md @@ -25,7 +25,7 @@ of value in the response which must be used in the next request. } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `response.param` method of pagination: - `responseParam` (required, string) — path to the key which contains the value used for scrolling @@ -38,7 +38,7 @@ parameters](/components/extractors/generic-extractor/configuration/config/jobs/# - `scrollRequest` (optional, object) — [job-like](/components/extractors/generic-extractor/configuration/config/jobs/) object (supported fields are `endpoint`, `method` and `params`) which allows to sent an initial scrolling request (see an [example](#using-scroll-request)). -### Stopping Condition +### Stopping condition The pagination ends **when the value of `responseParam` parameters is empty** — the key is not present at all, is null, is an empty string, or is `false`. Take care when configuring the `responseParam` parameter. If you, for example, misspell the name of the key, the extraction will not go beyond the first page. @@ -47,7 +47,7 @@ the key, the extraction will not go beyond the first page. ## Examples The following API pagination examples demonstrate the use of the Response Parameter Scroller. -### Basic Configuration +### Basic configuration Assume you have an API which returns, for instance, the next page number inside the response: ```json @@ -100,7 +100,7 @@ is sent to `/users?page=2`. See [example [EX057]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/057-pagination-response-param-basic). -### Overriding Parameters +### Overriding parameters The following configuration passes the parameter `orderBy` to every request: ```json @@ -143,7 +143,7 @@ and the second request to `/users?page=2&orderBy=id`. See [example [EX058]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/058-pagination-response-param-override). -### Using Scroll Request +### Using scroll request The response param scroller supports sending of an initial scrolling request. This can be used in situations where the API requires special initialization of a scrolling endpoint; for instance, the [Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/5.2/search-request-scroll.html). diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md index 8d70127ad..068e932f2 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md @@ -24,7 +24,7 @@ next page in the response. This scroller is suitable for APIs supporting the } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `response.url` pagination method: - `urlKey` (optional, string) — path in the response to the field which contains the URL of the next request; @@ -41,7 +41,7 @@ the default value is `.`. See the [examples below](#examples). -### Stopping Condition +### Stopping condition The pagination ends **when the value of the `urlKey` parameter is empty** — the key is not present at all, is null, is an empty string or is `false`. Be careful when configuring the `urlKey` parameter. If you, for example, misspell the key name, the extraction will not go beyond the first page. @@ -50,7 +50,7 @@ key name, the extraction will not go beyond the first page. ## Examples This section provides three API pagination examples where the Response URL Scroller is used. -### Basic Configuration +### Basic configuration To configure pagination for an API that supports the [JSON API specification](https://jsonapi.org/format/#fetching-pagination), use the configuration below: @@ -86,7 +86,7 @@ If the URL is *relative* (`users?page=2`), it is appended to the endpoint URL. See [example [EX054]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/054-pagination-response-url-basic). -### Merging Parameters +### Merging parameters To pass additional parameters to each of the page URLs, use the `includeParams` parameter: ```json @@ -144,7 +144,7 @@ would probably break the paging. See [example [EX055]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/055-pagination-response-url-params). -### Overriding Parameters +### Overriding parameters Sometimes the API does not pass the entire URL, but only the [query string](/components/extractors/generic-extractor/tutorial/rest/#url) parameters which should be used for querying the next page. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md index 442b8ee07..77add3355 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md @@ -31,7 +31,7 @@ A sample AWS signature configuration looks like this: See [example [EX143]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/143-aws-signature-request). -## AWS Signature Credentials +## AWS signature credentials - **accessKeyId** — AWS access key ID - **#secretKey** — AWS secret access key - **serviceName** — Signing to a particular service name diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md index 2206c380b..9d9f3910a 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md @@ -56,7 +56,7 @@ The Jobs configuration describes the API endpoints (resources) which will be ext includes configuring the HTTP method and parameters. The `jobs` configuration is **required** and is described in a [separate article](/components/extractors/generic-extractor/configuration/config/jobs/). -## Output Bucket +## Output bucket The `outputBucket` option defines the name of the [Storage Bucket](/storage/buckets/) in which the extracted tables will be stored. The configuration is **required** unless the extractor is [published](/components/extractors/generic-extractor/publish/) as a standalone component with the @@ -110,7 +110,7 @@ which case it is essentially equal to [`api.http.headers`](/components/extractor See [example [EX074]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/074-http-headers). -## Incremental Output +## Incremental output The `incrementalOutput` boolean option allows you to load the extracted data into [Storage](/storage/) incrementally. This flag in no way affects the data extraction. When `incrementalOutput` is set to `true`, the contents of the target table in Storage will not be cleared. @@ -121,7 +121,7 @@ is described in a [dedicated article](/components/extractors/generic-extractor/i See [example [EX075]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/075-incremental-output). -## User Data +## User data The `userData` option allows you to add arbitrary data to extracted records. It is an object with arbitrary property names which are added as columns to all records extracted from parent jobs. The property values are the columns values. It is also possible to use @@ -167,7 +167,7 @@ contains a column with the same name as a `userData` property, the `userData` co See [example [EX076]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/076-user-data). -## Compatibility Level +## Compatibility level As we develop the Generic Extractor, some of the new features might lead to minor differences in extraction results. When such a situation arises, a new *compatibility level* is introduced. The `compatLevel` setting allows you to force the old compatibility level and **temporarily** maintain the old behavior. The current diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md index 2ebe8311f..918554eaa 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md @@ -119,7 +119,7 @@ This is useful when using [User Defined functions](/components/extractors/generi or without having a placeholder in the `endpoint`. But then all the child requests would be the same and that is usually not what you intend to do. -### Placeholder Level +### Placeholder level Optionally, the placeholder name may be prefixed by a nesting **level**. Nesting allows you to refer to properties in other objects than the direct parent. The level is written as the placeholder name prefix, delimited by a colon `:`. For example, `2:user-id`. @@ -154,7 +154,7 @@ to not contain the value `" employee"` (which is probably not what you intended ## Examples This section contains a number of examples using child jobs. -### Basic Example +### Basic example Let's say that you have an API with two endpoints: - `/users/` — Returns a list of users. @@ -257,7 +257,7 @@ property (see the next example). The auto-generated name is rather ugly. See [example [EX021]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/021-basic-child-job). -### Basic Job With Data Type +### Basic job with data type To avoid automatic table names, it is advisable to always use the `dataType` property for child jobs: @@ -297,7 +297,7 @@ user-detail: See [example [EX022]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/022-basic-child-job-datatype). -### Basic Job With Array Values +### Basic job with array values It is also possible that the main job returns objects which contain direct references to the children: @@ -356,7 +356,7 @@ user-child: See [example [EX135]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/135-basic-child-job-array). -### Accessing Nested ID +### Accessing nested ID If the placeholder value is nested within the response object, you can use dot notation to access child properties of the response object. For instance, if the parent response with a list of users returns a response similar to this: @@ -427,7 +427,7 @@ Notice that the parent reference column name is the concatenation of the `parent See [example [EX023]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/023-child-job-nested-id). -### Accessing Deeply Nested Id +### Accessing deeply nested id The placeholder path is configured **relative to** the extracted object. Assume that the parent endpoint returns a complicated response like this: @@ -495,7 +495,7 @@ may be confusing because the endpoint property in that child job is set relative See [example [EX024]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/024-child-job-deeply-nested-id). -### Naming Conflict +### Naming conflict Because a new column is added to the table representing child properties, it is possible that you run into a naming conflict. That is, if the child response with user details looks like this: @@ -536,7 +536,7 @@ to create the column `parent_id` with the placeholder value, overwriting the ori See [example [EX025]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/025-naming-conflict). -### Nesting Level +### Nesting level By default, the placeholder value is taken from the object retrieved in the parent job. As long as the child jobs are nested only one level deep, there is no other option anyway. Let's see what happens with a deeper nesting. @@ -679,7 +679,7 @@ Notice that each table contains additional columns with the placeholder property See [example [EX026]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/026-basic-deeper-nesting). -### Nesting Level Alternative +### Nesting level alternative Because the required user and order IDs are present in multiple requests (in the list and in the detail), there are multiple ways how the jobs may be configured. For example, the following configuration produces the exact same result as the above configuration: @@ -768,7 +768,7 @@ the deepest child will really contain the `orderId` value. See [example [EX027]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/027-basic-deeper-nesting-alternative). -### Deep Job Nesting +### Deep job nesting Let's look at how to retrieve more nested API resources: ```json @@ -876,7 +876,7 @@ where the `parent_id` column refers the `5:user-id` placeholder. See [example [EX028]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/028-advanced-deep-nesting). -### Nested Array +### Nested array Suppose now that the endpoint `/users` returns a more complicated response: @@ -998,7 +998,7 @@ The `users-2\_members\_items` contains the same results as the `users` table, bu This makes the response in the `users` table quite useless, but the job is still required to generate the child jobs to obtain the `user-detail` table. See [example [EX106]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/106-child-jobs-array). -### Simple Filter +### Simple filter Let's assume that you have an API which has two endpoints: - `users` — Returns a list of users. @@ -1093,7 +1093,7 @@ the details are retrieved only for the desired users. See [example [EX029]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/029-simple-filter). -### Not Like Filter +### Not like filter Apart from the standard comparison operators, the recursive filter allows to use a **like** comparison operator `~`. It expects that the value contains a placeholder `%`, which matches any number of characters. The following configuration: @@ -1129,7 +1129,7 @@ following `user-detail` table will be extracted: See [example [EX030]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/030-not-like-filter). -### Combining Filters +### Combining filters Multiple filters can be combined using the [logical](https://en.wikipedia.org/wiki/Boolean_algebra#Basic_operations) `&` (and) and `|` (or) operators. For example, the following configuration retrieves details for users who have @@ -1162,7 +1162,7 @@ The following `user-detail` will be produced: See [example [EX031]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/031-combined-filter). -### Multiple Filter Combinations +### Multiple filter combinations Although you can join a multiple filter expression with logical operators as in the above example, there is no support for parentheses. The following configuration combines multiple filters: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md index 0dd71fbf3..e4d3aef48 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md @@ -59,7 +59,7 @@ way. Each response is processed in the following steps: 3. Flatten the object structure into one or more tables. 4. Create the required tables in Storage and load data into them. -## Merging Responses +## Merging responses The first two steps are the responsibility of [Jobs](/components/extractors/generic-extractor/configuration/config/jobs/) resulting in an array of objects. Generic Extractor then tries to find a common super-set of properties of all objects, for example, with the following response: @@ -127,23 +127,23 @@ Assume the following [API definition](/components/extractors/generic-extractor/c } ``` -### Relative URL Fragment +### Relative URL fragment The relative endpoint **must not start** with a slash; so, with `endpoint` set to `campaign`, the final resource URL would be `https://example.com/3.0/campaign`. -### Absolute Domain URL +### Absolute domain URL The absolute endpoint **must start** with a slash. So, with `/endpoint` set to `campaign`, the final resource URL would be `https://example.com/campaign`. This means that the path part specified in the `baseURL` is ignored and fully replaced by the value specified in `endpoint`. -### Absolute Full URL +### Absolute full URL The full absolute URL must start with a protocol. So, with the endpoint set to `https://eu.example.com/campaign`, this would be the final resource URL and the path specified in the `baseURL` is completely ignored. -### Specifying Endpoint +### Specifying endpoint The following table summarizes possible outcomes: |`baseURL`|`endpoint`|actual URL| @@ -167,7 +167,7 @@ Also, closely follow the target API specification regarding trailing slashes. Fo both `https://example.com/3.0/campaign` and `https://example.com/3.0/campaign/` URLs may be accepted and valid. For other APIs, however, only one version may be supported. -## Request Parameters +## Request parameters The `params` section defines [request parameters](/components/extractors/generic-extractor/tutorial/rest). They may be optional or required, depending on the target API specification. The `params` section is an object with arbitrary properties (or, more precisely, parameters understood by the target @@ -257,7 +257,7 @@ or, in a more readable [URLDecoded](https://urldecode.org/) form: Also, the `Content-Type: application/x-www-form-urlencoded` HTTP header will be added to the request. -## Data Type +## Data type The `dataType` parameter assigns a name to the object(s) obtained from the endpoint. Setting it is optional. If not set, a name will be generated automatically from the `endpoint` value and parent jobs. @@ -285,7 +285,7 @@ for example, in a situation where two API endpoints return the same resource: In the above case, only a single `tickets` table will be produced in the output bucket. It will contain records from both API endpoints. -## Data Field +## Data field The `dataField` parameter is used to determine what part of the API **response** will be extracted. The following rules apply by default: @@ -320,7 +320,7 @@ as an object with the `path` property. For instance, these two configurations ar ] ``` -### Data Field Delimiter +### Data field delimiter The path to the response property is by default expected to be dot separated. That is — a path `members.active` refers to the property `active` nested inside the property `members`. If you need to refer to a property containing a dot, you have to change the data field path delimiter to some other character. This can be @@ -356,7 +356,7 @@ inside the property `members.active` you have to use: The `delimiter` character is completely arbitrary but must be something that is not used in the property names in the response. See [example [EX120]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/120-datafield-separator). -## Response Filter +## Response filter The `responseFilter` option allows you to skip parts of the API response from processing. This can be useful in these cases: @@ -687,7 +687,7 @@ The following table will be extracted: See [example [EX009]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/009-nested-array). -## Examples with Complicated Objects +## Examples with complicated objects The above examples show how simple objects are extracted from different objects. Generic Extractor can also extract objects with non-scalar properties. The default [JSON to CSV mapping](/components/extractors/generic-extractor/configuration/config/mappings/) flattens nested objects and produces secondary tables from nested arrays. @@ -910,7 +910,7 @@ auto-generated key to the parent *Users* table. Also notice that the See [example [EX012]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/012-deeply-nested-object). -## Response Filter Examples +## Response filter examples ### Skip flattening If you have an API response like this: @@ -1040,7 +1040,7 @@ See [example [EX014]](https://github.com/keboola/generic-extractor/tree/master/d