diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md index f5fb13933..c444c5ca7 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/api_key/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/api_key/ --- + + API Key token authentication sends a token in either a header or query parameter of each API request. E.g., Headers: `X-StorageApi-Token:your_token` diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md index 5a6ea5314..913d1b8ad 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/basic/index.md @@ -5,11 +5,13 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/basic/ --- + + Basic Authentication provides the [HTTP Basic Authentication](https://en.wikipedia.org/wiki/Basic_access_authentication) method. It requires entering a username and password in the configuration and sends the encoded values in the `Authorization` header. -### User Interface +### User interface In the user interface, you simply select the `Basic Authorization` method and enter the username and password. @@ -39,11 +41,11 @@ They are also prefixed by the hash `#` character, which means they are stored [e If the API expects something else than a username and password in the `Authorization` header, or if it requires a custom authorization header, use the [Default Headers option](/components/extractors/generic-extractor/configuration/api/#headers). -## Configuration Parameters +## Configuration parameters This `basic` type of authentication has no configuration parameters. The login and password must be provided in the [`config` section](/components/extractors/generic-extractor/configuration/config/) of the Generic Extractor configuration. -## Basic Configuration Example +## Basic configuration example Assume you have an API which requires you to use the HTTP Basic authentication to send the login and password in the `Authorization` header. Assume that your login is `JohnDo` and password is `secret`. The following configuration solves the situation: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md index 1acedad01..11f249d25 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/bearer_token/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/bearer_token/ --- + + Bearer token authentication sends a token in the `Authorization` header of each API request. This method is available through UI. You can select the `Bearer Token` method and fill in the token. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md index 8c863ed97..65cc7d04c 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various configuration options.* diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md index 449e9b2ac..f449e3d1b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/login/index.md @@ -5,11 +5,13 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/login/ --- + + Use the Login authentication to send a one-time **login request** to obtain temporary credentials for authentication of all the other API requests. -## User Interface +## User interface Note that this configuration option is not yet covered. You can add the JSON configuration using the `Custom` auth method. @@ -51,7 +53,7 @@ A sample Login authentication looks like this: } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `login` type of authentication: - `loginRequest` (required, object) — a [job-like](/components/extractors/generic-extractor/configuration/config/jobs/) object describing the login request; it has the following properties: @@ -77,7 +79,7 @@ is called only once before all other requests. To call the login request before ## Examples Below are several examples showing you how to use various login authentication related features in Generic Extractor. -### Configuration with Headers +### Configuration with headers Let's say you have an API which requires every API call to be authorized with the `X-ApiToken` header. The value of that header (an API token) is obtained by calling the `/login` endpoint with the headers `X-Login` and `X-Password`. The `/login` endpoint response looks like this: @@ -128,7 +130,7 @@ will contain the header: See [example [EX079]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/079-login-auth-headers). -### Configuration with Headers and Text Response +### Configuration with headers and text response Let's say you have an API like the above, but it returns the login response as a plain text: a1b2c3d435f6 @@ -181,7 +183,7 @@ will contain the header: See [example [EX128]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/128-login-auth-text). -### Configuration with Query Parameters +### Configuration with query parameters Let's say you have an API which requires an [HTTP POST](https://en.wikipedia.org/wiki/POST_(HTTP)) request with `username` and `password` to the endpoint `/login/form`. On a successful login, it returns the following response: @@ -244,7 +246,7 @@ so the second API call will be sent as: See [example [EX080]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/080-login-auth-query). Notice that the example uses completely different URL for the login request. -### Parameter Overriding +### Parameter overriding The above examples show how to use query parameters and headers separately. However, they can be mixed freely; they can also be mixed with parameters and headers entered elsewhere in the configuration. The following example shows how parameters from different places are merged together: @@ -405,7 +407,7 @@ This causes Generic Extractor to call the **login request** every hour. See [example [EX082]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/082-login-auth-expires). -### Expiration from Response +### Expiration from response In case the credentials provided by the **login request** have a time-limited validity, use the `expires` option. If the validity of the credentials is returned in the response, modify the [first example](#configuration-with-headers) to this: @@ -453,7 +455,7 @@ This assumes that the response of the **login request** looks like this: See [example [EX083]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/083-login-auth-expires-date). -### Relative Expiration from Response +### Relative expiration from response In case the API returns credentials validity in the **login request** and that validity is expressed in seconds, use the `expires` option together with setting `relative` to `true`. The result is the behavior of the [first example](#expiration-basic) but the value is taken @@ -504,7 +506,7 @@ This assumes that the response of the **login request** looks like this: See [example [EX084]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/084-login-auth-expires-seconds). -### Login Authentication with Functions +### Login authentication with functions Suppose you have an API which requires you to send a username and password separated by a colon and base64 encoded — for example, `JohnDoe:TopSecret` (base64 encoded to `Sm9obkRvZTpUb3BTZWNyZXQ=`) in the `X-Authorization` header to an `/auth` endpoint. The login endpoint then returns a token @@ -568,7 +570,7 @@ uses the `login` authorization method to send them to the special `/auth` endpoi See [example [EX100]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/100-function-login-headers). -### Login Authentication with Login and API Request +### Login authentication with login and API request Suppose you have an API similar to the one in the [previous example](#login-authentication-with-functions). It requires you to send a username and password separated by a colon and base64 encoded — for example, `JohnDoe:TopSecret` (base64 encoded to `Sm9obkRvZTpUb3BTZWNyZXQ=`) in the diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md index 0e57c4dd6..140d10b7b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth10/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth10/ --- + + **Note** that this configuration option is not yet supported and the test endpoint button will not work. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md index adde56120..2f30c38e0 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth20-login/ --- + + **Note** that this configuration option is not yet supported and the test endpoint button will not work. @@ -40,7 +42,7 @@ for authentication of all the other API requests. A sample OAuth Login authentic } ``` -## Configuration Parameters +## Configuration parameters The configuration parameters are identical to the [Login](/components/extractors/generic-extractor/configuration/api/authentication/login/) method. The difference, however, is in the [function context](/components/extractors/generic-extractor/functions/#oauth-20-login-authentication-context). The **login request** is assumed to require the OAuth2 authorization and its response must be in JSON format (plaintext is not supported). @@ -48,7 +50,7 @@ The **login request** is assumed to require the OAuth2 authorization and its res ## Examples The following examples demonstrate how to use OAuth with a basic login request and Google API in Generic Extractor. -### Basic Configuration +### Basic configuration The following configuration shows how to set up an OAuth **login request**: ```json @@ -131,7 +133,7 @@ and sent to other API requests (`/users`). See [example [EX105]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/105-oauth2-login). -### Google API Configuration +### Google API configuration The following example shows how to set up the OAuth authentication for Google APIs. The access token is refreshed with each API call. #### Generate access tokens diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md index 9eac3d715..da346eedd 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth20/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth20/ --- + + OAuth 2.0 Authentication is one of [two OAuth methods](/components/extractors/generic-extractor/configuration/api/authentication/#oauth) and is supported only for [components registered in the developer portal](/components/extractors/generic-extractor/publish/). @@ -76,7 +78,7 @@ Note that the properties `appKey` and `#appSecret` must exist even if not used b to empty strings. For more information about OAuth 2, see the [official documentation](https://oauth.net/2/) or learn [more about Keboola-OAuth integration](/extend/common-interface/oauth). -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `oauth20` authentication type: - `format` (optional, string) — If the OAuth service provider response is JSON, use the only possible @@ -92,7 +94,7 @@ are available in the [OAuth function context](/components/extractors/generic-ext ## Examples The following two examples demonstrate the support for OAuth 2 in Generic Extractor. -### Bearer Authentication +### Bearer authentication The most basic OAuth authentication method is with "Bearer Token". If you have an API which supports this authentication method, the following configuration can be used: @@ -144,7 +146,7 @@ the header `Authorization: Bearer SomeToken1234abcd567ef` using the See [example [EX103]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/103-oauth2-bearer). -### HMAC Authentication +### HMAC authentication If you have an API which requires an [HMAC](https://en.wikipedia.org/wiki/Hash-based_message_authentication_code) signed token, generate the correct signature using [functions](/components/extractors/generic-extractor/functions). The following example assumes you obtain the following response from the API upon authentication: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md index d1b35ce6b..3f3f2d79b 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc/index.md @@ -5,13 +5,15 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/oauth_cc/ --- + + oAuth 2.0 Client Credentials authentication performs the [oAuth 2.0 client_credentials flow](https://auth0.com/docs/get-started/authentication-and-authorization-flow/client-credentials-flow). This method is available through the UI and is implemented via the [Login](/components/extractors/generic-extractor/configuration/api/authentication/login/) method. ![img.png](/components/extractors/generic-extractor/configuration/api/authentication/oauth_cc.png) -### Configuration Parameters +### Configuration parameters - `Login Request type` - `Basic Auth`: The client_id and client_secret are sent in the Authorization header as a Basic authorization, e.g. `Authorization: Basic base64(client_id:client_secret)`. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md index 561d1706d..7efef1c17 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/authentication/query/index.md @@ -5,13 +5,15 @@ redirect_from: - /extend/generic-extractor/configuration/api/authentication/query/ --- + + Query Authentication provides the simplest authentication method, in which the credentials are sent in the [request URL](/components/extractors/generic-extractor/tutorial/rest#url). This method is most often used with APIs that authenticate using API tokens and signatures. Dynamic values of query parameters can be generated using [user functions](/components/extractors/generic-extractor/functions/). -## User Interface +## User interface In the user interface, you simply select the `Query` method and enter the key-value pairs of the query parameters. ![img.png](/components/extractors/generic-extractor/configuration/api/authentication/query.png) @@ -36,12 +38,12 @@ A sample Query authentication configuration looks like this: } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `query` type of authentication: - `query` (required, object): An object whose properties represent key-value pairs of the URL query. -## Basic Configuration Example +## Basic configuration example Let's say you have an API that requires an `api-token` parameter (with value 2267709) to be sent with each request. The following authentication configuration does exactly that: @@ -60,7 +62,7 @@ configuration remains organized. See [example [EX077]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/077-query-auth). -## Configuration With Encrypted Token Example +## Configuration with encrypted token example Usually, you want the value used for authentication to be encrypted (the `api-token` parameter with the value 2267709 in our example), so you do not expose it to other users or store it in the configuration versions history. The following authentication configuration, combined with the parameter defined in the [`config`](/components/extractors/generic-extractor/configuration/config/) section, does that (the value with the prefix `#` is encrypted upon saving the configuration): diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md index 5d79e3ea1..ea1caa408 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/basic/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various @@ -92,7 +94,7 @@ Authentication (authorization) needs to be configured for any API which is not p Because there are many authorization methods used by different APIs, there are also many [configuration options](/components/extractors/generic-extractor/configuration/api/authentication/). -## Retry Configuration +## Retry configuration By default, Generic Extractor **automatically retries failed HTTP requests** — repeatedly, and on most errors. This is one of the big advantages over writing your own extractor from scratch. Tweak the retry setting to optimize the speed of an extraction or to avoid unwanted flooding of the API. @@ -124,7 +126,7 @@ There are two retry strategies: - Either the API sends a `Retry-After` header (or its equivalent), or - Generic Extractor uses an [exponential backoff algorithm](https://en.wikipedia.org/wiki/Exponential_backoff). -### API Retry Strategy +### API retry strategy Per the HTTP specification, the API may send the [`Retry-After`](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Retry-After) header which should contain number of seconds to pause/sleep before the next request. Generic Extractor supports some extensions to this. First, the *Retry Header* name may be customized. Second, the header @@ -138,7 +140,7 @@ time of the next request The second and third options are often called **Rate Limit Reset** as they describe when the next successful request can be made (i.e., the limit is reset). -### Backoff Strategy +### Backoff strategy The exponential backoff in Generic Extractor is defined as `truncate(2^(retry\_number - 1)) * 1000` seconds. This means that the first retry (zero-based index) will be after 0 seconds (`(2^(0-1)) = 0.5`, truncated to 0). The retry delays are the following: @@ -179,7 +181,7 @@ in the [debug](/components/extractors/generic-extractor/running/#debug-mode) mes If the exponential backoff is used, you will see its sequence of times. See an [example](/components/extractors/generic-extractor/configuration/api/#retry-configuration). -## Default HTTP Options +## Default HTTP options The `http` configuration option allows you to set the timeouts, default headers and parameters sent with each API call (defined later in the [`jobs` section](/components/extractors/generic-extractor/configuration/config/jobs/#request-parameters)). @@ -199,7 +201,7 @@ the headers and values are their values — for instance: See the full [example](/components/extractors/generic-extractor/configuration/api/#default-headers). -### Request Parameters +### Request parameters The `http.defaultOptions.params` configuration allows you to **set the [request parameters](/components/extractors/generic-extractor/tutorial/rest/#url) to be sent with each API request**. The same rules apply as to the @@ -207,7 +209,7 @@ sent with each API request**. The same rules apply as to the See an [example](/components/extractors/generic-extractor/configuration/api/#default-headers). -### Required Headers +### Required headers Similar to the `http.headers` option, the `http.requiredHeaders` option allows you to **set the HTTP header for every API request**. The difference is that the `requiredHeaders` configuration specifies **only the header names**. The actual values must be provided in the [`config`](/components/extractors/generic-extractor/configuration/config/) @@ -238,7 +240,7 @@ Failing to provide the header values in the `config` section will cause an error See the full [example](/components/extractors/generic-extractor/configuration/api/#required-headers). -### Ignore Errors +### Ignore errors The `ignoreErrors` option allows you to force Generic Extractor to ignore certain extraction errors. The option lists HTTP codes for which any errors occurring during downloading and JSON parsing the response will be ignored. The `ignoreErrors` option error is an array of HTTP @@ -272,7 +274,7 @@ API implementations and should not be used blindly if other solutions may be app [`responseFilter`](/components/extractors/generic-extractor/configuration/config/jobs/#response-filter). When ignoring errors, **you might miss even those errors that require your attention.** -### Connect Timeout +### Connect timeout The `connectTimeout` option is a float describing the number of seconds to wait while trying to connect to a server. Default value is `30` seconds. Use `0` to wait indefinitely, we do not recommend it. @@ -283,7 +285,7 @@ Default value is `30` seconds. Use `0` to wait indefinitely, we do not recommend } ``` -### Request Timeout +### Request timeout The `requestTimeout` option is a float describing the total timeout of the request in seconds. Default value is `300` seconds. Use `0` to wait indefinitely, we do not recommend it. @@ -296,7 +298,7 @@ Default value is `300` seconds. Use `0` to wait indefinitely, we do not recommen ## Examples -### Retry Configuration +### Retry configuration Assume that you have an API which implements throttling in the following way: when the number of requests is exceeded, it returns an empty response with the status code `202` and a timestamp when a new requests can be made in the `X-RetryAfter` HTTP header. @@ -321,7 +323,7 @@ Notice that it is necessary to add the response code `202` to the existing defau See [example [EX037]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/037-retry-header). -### Default Headers +### Default headers Assume that you have an API which returns a JSON response only if the client sends an `Accept: application/json` header. Additionally, if the client sends an `Accept-Encoding: gzip` header, the HTTP transmission will be compressed (and thus faster). @@ -341,7 +343,7 @@ The following configuration sends both headers with every API request: See [example [EX038]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/038-default-headers). -### Default Parameters +### Default parameters Assume that you have an API requiring all requests to contain a filter for the account to which they belong. This is done by passing the `account=XXX` parameter. The following configuration sends the parameter with every API request: @@ -364,7 +366,7 @@ may also be used. See [example [EX039]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/039-default-parameters). -### Required Headers +### Required headers Assume that an API requires the header `X-AppKey` to be sent with each API request. The following API configuration can be used: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md index 2f4a753c2..446779c91 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/cursor/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/cursor/ --- + + The Cursor Scroller can be used with an API which expects the client to maintain a cursor (pointer) to the last obtained item. For example, on the first request, it returns items with ID 1-100; for the second @@ -24,7 +26,7 @@ request, you must tell the API to start with ID 101. } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `cursor` method of pagination: - `idKey` (required, string) — path to the key which contains the value of the cursor; the path is entered relative to the exported items. @@ -41,7 +43,7 @@ The request parameter specified in the `param` configuration overwrites the para [job parameters](/components/extractors/generic-extractor/configuration/config/jobs/#request-parameters). Other job parameters are carried over without modification (see an [example](#reverse-configuration)). -### Stopping Condition +### Stopping condition The pagination ends **when the `dataField` of the response contains no items**. Because of this, each run with the `cursor` scroller produces a similar warning: @@ -52,7 +54,7 @@ This is expected behavior. [Common stopping conditions](/components/extractors/g ## Examples This section contains two API pagination examples where the Cursor Scroller is used. -### Basic Configuration +### Basic configuration Let's say you have an API which has an endpoint `/users` returning the following response: ```json @@ -89,7 +91,7 @@ Notice that the `idKey` parameter is relative to the extracted array of items (` See [example [EX060]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/060-pagination-cursor-basic). -### Reverse Configuration +### Reverse configuration Some APIs return items starting with the newest item and therefore need to be queried for offset in reverse order. Let's say a request to `/users?startWith=last` will produce: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md index 150d6d576..e08b43ec4 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/ --- + + *If new to Generic Extractor, learn about [pagination in our tutorial](/components/extractors/generic-extractor/tutorial/pagination/) first.* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various configuration options.* @@ -41,7 +43,7 @@ An example pagination configuration looks like this: } ``` -## Paging Strategy +## Paging strategy Generic Extractor supports the following paging strategies (scrollers); they are configured using the `method` option: @@ -52,7 +54,7 @@ using the `method` option: - [`cursor`](/components/extractors/generic-extractor/configuration/api/pagination/cursor/) — uses the identifier of the item in response to maintain a scrolling cursor. - [`multiple`](/components/extractors/generic-extractor/configuration/api/pagination/multiple/) — allows to set different scrollers for different API endpoints. -### Choosing Paging Strategy +### Choosing paging strategy If the API responses contain direct links to the next set of results, use the [`response.url` method](/components/extractors/generic-extractor/configuration/api/pagination/response-url/). This applies to the APIs following the [JSON API specification](https://jsonapi.org/). The response usually @@ -100,7 +102,7 @@ If the API uses different paging methods for different endpoints, use the [`multiple` method](/components/extractors/generic-extractor/configuration/api/pagination/multiple/) together with any of the above methods. -## Stopping Strategy +## Stopping strategy Generic Extractor stops scrolling - based on the `nextPageFlag` condition configuration. @@ -135,7 +137,7 @@ the first page, it is not same as the previous page and therefore another reques is the same as the previous page, the same check kicks in and the extraction is stopped too. However, the results from the first page will be duplicated. -### Next Page Flag +### Next page flag The above describes automatic behavior of Generic Extractor regarding scrolling stopping. Using **Next Page Flag** allows you to do a **manual setup of the stopping strategy**: Generic Extractor analyzes the response, looks for a particular field (the flag) and decides whether to continue scrolling based on the value or presence of that flag. @@ -167,7 +169,7 @@ Example `nextPageFlag` setting: See our [Next Page Flag Examples](#next-page-flag-examples). -### Force Stop +### Force stop Force stop configuration allows you to stop scrolling when some extraction limits are hit. The supported options are: @@ -232,7 +234,7 @@ and [example [EX116]](https://github.com/keboola/generic-extractor/tree/master/d and [example [EX140]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/140-pagination-forcestop-child-filter) (combining with child jobs). -### Limit Stop +### Limit stop Limit stop configuration allows you to stop scrolling when a specified number of items is extracted. The supported options are: @@ -281,7 +283,7 @@ See [example [EX126]](https://github.com/keboola/generic-extractor/tree/master/d For `count` configuration, see [example [EX127]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/127-pagination-stop-field) (a modified version of [EX051](https://github.com/keboola/generic-extractor/tree/master/doc/examples/051-pagination-pagenum-basic). -### Combining Multiple Stopping Strategies +### Combining multiple stopping strategies All stopping strategies are evaluated simultaneously and for the scrolling to continue, none of the stopping conditions must be met. In other words, the scrolling continues until any of the stopping conditions is true. To this you need to account specific stopping strategies for @@ -310,14 +312,14 @@ will stop if **any** of the following is true: - The `isLast` field is present in the response and is true (`nextPageFlag`). - The `isLast` field is not present in the response. -## Next Page Flag Examples +## Next page flag examples In this section, we want to show you the following examples of the Next Page Flag stopping strategy: - Has-More Scrolling - Non-Boolean Has-More Scrolling - Is-Last Scrolling -### Has-More Scrolling +### Has-More scrolling Assume that the API returns a response which contains a `hasMore` field. The field is present in every response and has always the value `true` except for the last response where it is `false`. The following pagination configuration can be used to configure the stopping strategy: @@ -339,7 +341,7 @@ In this case, setting `ifNotSet` is not necessary. See [example [EX045]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/045-next-page-flag-has-more) and [example [EX139]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/139-pagination-hasmore-child-filter) (combining with child jobs). -### Non-Boolean Has-More Scrolling +### Non-Boolean Has-More scrolling Assume that the API returns a response which contains a `hasMore` field. The field is present only in the last response and has the value `"no"` there. The following pagination configuration can be used to configure the stopping strategy: @@ -363,7 +365,7 @@ to false. In this case setting `ifNotSet` is mandatory. See [example [EX046]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/046-next-page-flag-has-more-2). -### Is-Last Scrolling +### Is-Last scrolling Assume that the API returns a response which contains an `isLast` field. The field is present only in the last response and has the value `true` there. The following pagination configuration can be used to configure the stopping strategy: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md index 25e3b5065..9b222657a 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/multiple/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/multiple/ --- + + Setting the pagination method to `multiple` allows you to use **multiple scrollers on a single API**. This type of pagination contains the definition of all scrollers used in the entire configuration. @@ -52,7 +54,7 @@ The name of the scroller must be used in a specific [job `scroller` parameter](/ A `default` scroller can be set (must be one of the names defined in `scrollers`). In that case, all jobs without an assigned scroller will use the default one. -### Stopping Condition +### Stopping condition There are no specific stopping conditions for the `multiple` pagination. Each scroller acts upon its normal stopping conditions. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md index 40d10725f..775b84189 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/offset/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/offset/ --- + + The Offset scroller handles a pagination strategy in which the API splits the results into pages of the same size (limit parameter) and navigates through them using the **item offset** parameter. This @@ -27,7 +29,7 @@ An example configuration: } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `offset` method of pagination: - `limit` (required, integer) — page size @@ -43,7 +45,7 @@ the [job parameters](/components/extractors/generic-extractor/configuration/conf 100 items at most and you set the limit 1000, it would cause the extraction to stop after the first page. This is because the [underflow condition](/components/extractors/generic-extractor/configuration/api/pagination/#stopping-strategy) would be triggered. -### Stopping Condition +### Stopping condition Scrolling is stopped **when the result contains less items than requested** — specified in the `limit` configuration (underflow). This also includes an instance when no items are returned, or the response is empty. @@ -86,7 +88,7 @@ All [common stopping conditions](/components/extractors/generic-extractor/config ## Examples This section contains three examples of API pagination using the Offset Scroller. -### Basic Scrolling +### Basic scrolling This is the simplest scrolling setup: ```json @@ -101,7 +103,7 @@ The next request has `limit=20` and `offset=20`, for example, `/users?limit=20&o See [example [EX043]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/043-paging-stop-underflow) and [example [EX044] with a more structured response](https://github.com/keboola/generic-extractor/tree/master/doc/examples/044-paging-stop-underflow-struct). -### Renaming Parameters +### Renaming parameters The `limitParam` and `offsetParam` configuration options allow you to rename the limit and offset for the needs of a specific API: @@ -119,7 +121,7 @@ and `skip=0`, for example, `/users?count=2&skip=0`. See [example [EX049]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/049-pagination-offset-rename). -### Overriding Limit and Offset +### Overriding limit and offset It is possible to override both the limit and offset parameters of a specific API job. This is useful in case you want to use different limits for different API endpoints. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md index b9730b3f1..2bd1c3089 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/pagenum/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/pagenum/ --- + + The Page Number Scroller handles a pagination strategy in which the API splits the results into pages of the same size (limit parameter) and navigates through them using the **page offset** parameter. @@ -24,7 +26,7 @@ If you need to use the item offset, use the [Offset Scroller](/components/extrac } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `pagenum` method of pagination: - `limit` (optional, integer) — page size @@ -34,7 +36,7 @@ The following configuration parameters are supported for the `pagenum` method of value is `true`. - `firstPage` (optional, integer) — index of the first page; the default value is `1`. -### Stopping Condition +### Stopping condition The `pagenum` scroller uses similar stopping condition as the [`offset` scroller](/components/extractors/generic-extractor/configuration/api/pagination/offset/#stopping-condition). Scrolling is stopped in case of an underflow — when the result contains **less items than requested** (including zero). However, in the `pagenum` scroller, the **`limit` parameter is not required** and has **no default value**. This means that if you omit it, @@ -43,7 +45,7 @@ the scrolling will stop only if an empty page is encountered. ## Examples This section contains three API pagination examples where the Page Number Scroller is used. -### Basic Scrolling +### Basic scrolling The most simple scrolling setup is the following: ```json @@ -57,7 +59,7 @@ The next request will have `page=2`, for example `/users?page=2`. See [example [EX051]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/051-pagination-pagenum-basic). -### Renaming Parameters +### Renaming parameters The `limitParam` and `pageParam` configuration options allow you to rename the limit and offset for the needs of a specific API: @@ -78,7 +80,7 @@ and `set=1`; for example, `/users?set=1&count=20`. See [example [EX052]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/052-pagination-pagenum-rename). -### Overriding Parameters +### Overriding parameters It is possible to override the limit parameter of a specific API job. This is useful when you want to use different limits for different API endpoints. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md index 4cda14f9c..60b5e359e 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-param/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/response-param/ --- + + The Response Parameter Scroller can be used with APIs that provide a certain kind of value in the response which must be used in the next request. @@ -23,7 +25,7 @@ of value in the response which must be used in the next request. } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `response.param` method of pagination: - `responseParam` (required, string) — path to the key which contains the value used for scrolling @@ -36,7 +38,7 @@ parameters](/components/extractors/generic-extractor/configuration/config/jobs/# - `scrollRequest` (optional, object) — [job-like](/components/extractors/generic-extractor/configuration/config/jobs/) object (supported fields are `endpoint`, `method` and `params`) which allows to sent an initial scrolling request (see an [example](#using-scroll-request)). -### Stopping Condition +### Stopping condition The pagination ends **when the value of `responseParam` parameters is empty** — the key is not present at all, is null, is an empty string, or is `false`. Take care when configuring the `responseParam` parameter. If you, for example, misspell the name of the key, the extraction will not go beyond the first page. @@ -45,7 +47,7 @@ the key, the extraction will not go beyond the first page. ## Examples The following API pagination examples demonstrate the use of the Response Parameter Scroller. -### Basic Configuration +### Basic configuration Assume you have an API which returns, for instance, the next page number inside the response: ```json @@ -98,7 +100,7 @@ is sent to `/users?page=2`. See [example [EX057]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/057-pagination-response-param-basic). -### Overriding Parameters +### Overriding parameters The following configuration passes the parameter `orderBy` to every request: ```json @@ -141,7 +143,7 @@ and the second request to `/users?page=2&orderBy=id`. See [example [EX058]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/058-pagination-response-param-override). -### Using Scroll Request +### Using scroll request The response param scroller supports sending of an initial scrolling request. This can be used in situations where the API requires special initialization of a scrolling endpoint; for instance, the [Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/5.2/search-request-scroll.html). diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md index 61d01e6fd..068e932f2 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/api/pagination/response-url/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/api/pagination/response-url/ --- + + The Response URL Scroller can be used with APIs that provide the URL of the next page in the response. This scroller is suitable for APIs supporting the @@ -22,7 +24,7 @@ next page in the response. This scroller is suitable for APIs supporting the } ``` -## Configuration Parameters +## Configuration parameters The following configuration parameters are supported for the `response.url` pagination method: - `urlKey` (optional, string) — path in the response to the field which contains the URL of the next request; @@ -39,7 +41,7 @@ the default value is `.`. See the [examples below](#examples). -### Stopping Condition +### Stopping condition The pagination ends **when the value of the `urlKey` parameter is empty** — the key is not present at all, is null, is an empty string or is `false`. Be careful when configuring the `urlKey` parameter. If you, for example, misspell the key name, the extraction will not go beyond the first page. @@ -48,7 +50,7 @@ key name, the extraction will not go beyond the first page. ## Examples This section provides three API pagination examples where the Response URL Scroller is used. -### Basic Configuration +### Basic configuration To configure pagination for an API that supports the [JSON API specification](https://jsonapi.org/format/#fetching-pagination), use the configuration below: @@ -84,7 +86,7 @@ If the URL is *relative* (`users?page=2`), it is appended to the endpoint URL. See [example [EX054]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/054-pagination-response-url-basic). -### Merging Parameters +### Merging parameters To pass additional parameters to each of the page URLs, use the `includeParams` parameter: ```json @@ -142,7 +144,7 @@ would probably break the paging. See [example [EX055]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/055-pagination-response-url-params). -### Overriding Parameters +### Overriding parameters Sometimes the API does not pass the entire URL, but only the [query string](/components/extractors/generic-extractor/tutorial/rest/#url) parameters which should be used for querying the next page. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md index ca5d7f1b3..77add3355 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/aws-signature/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/aws-signature/ --- + + Generic Extractor allows signing requests with [**AWS Signature Version 4**](https://docs.aws.amazon.com/general/latest/gr/signature-version-4.html). Signing is the process of adding authentication information to your requests. When you use AWS tools, the extractor signs your API request. @@ -29,7 +31,7 @@ A sample AWS signature configuration looks like this: See [example [EX143]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/143-aws-signature-request). -## AWS Signature Credentials +## AWS signature credentials - **accessKeyId** — AWS access key ID - **#secretKey** — AWS secret access key - **serviceName** — Signing to a particular service name diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md index 9b3e13d7f..be1191d58 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/config/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various @@ -54,7 +56,7 @@ The Jobs configuration describes the API endpoints (resources) which will be ext includes configuring the HTTP method and parameters. The `jobs` configuration is **required** and is described in a [separate article](/components/extractors/generic-extractor/configuration/config/jobs/). -## Output Bucket +## Output bucket The `outputBucket` option defines the name of the [Storage Bucket](/storage/buckets/) in which the extracted tables will be stored. The configuration is **required** unless the extractor is [published](/components/extractors/generic-extractor/publish/) as a standalone component with the @@ -108,7 +110,7 @@ which case it is essentially equal to [`api.http.headers`](/components/extractor See [example [EX074]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/074-http-headers). -## Incremental Output +## Incremental output The `incrementalOutput` boolean option allows you to load the extracted data into [Storage](/storage/) incrementally. This flag in no way affects the data extraction. When `incrementalOutput` is set to `true`, the contents of the target table in Storage will not be cleared. @@ -119,7 +121,7 @@ is described in a [dedicated article](/components/extractors/generic-extractor/i See [example [EX075]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/075-incremental-output). -## User Data +## User data The `userData` option allows you to add arbitrary data to extracted records. It is an object with arbitrary property names which are added as columns to all records extracted from parent jobs. The property values are the columns values. It is also possible to use @@ -165,7 +167,7 @@ contains a column with the same name as a `userData` property, the `userData` co See [example [EX076]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/076-user-data). -## Compatibility Level +## Compatibility level As we develop the Generic Extractor, some of the new features might lead to minor differences in extraction results. When such a situation arises, a new *compatibility level* is introduced. The `compatLevel` setting allows you to force the old compatibility level and **temporarily** maintain the old behavior. The current diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md index 4217dee6b..2fd91343e 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/children/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/config/jobs/children/ --- + + *If new to Generic Extractor, learn about [jobs in our tutorial](/components/extractors/generic-extractor/tutorial/jobs/) first.* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various @@ -117,7 +119,7 @@ This is useful when using [User Defined functions](/components/extractors/generi or without having a placeholder in the `endpoint`. But then all the child requests would be the same and that is usually not what you intend to do. -### Placeholder Level +### Placeholder level Optionally, the placeholder name may be prefixed by a nesting **level**. Nesting allows you to refer to properties in other objects than the direct parent. The level is written as the placeholder name prefix, delimited by a colon `:`. For example, `2:user-id`. @@ -152,7 +154,7 @@ to not contain the value `" employee"` (which is probably not what you intended ## Examples This section contains a number of examples using child jobs. -### Basic Example +### Basic example Let's say that you have an API with two endpoints: - `/users/` — Returns a list of users. @@ -255,7 +257,7 @@ property (see the next example). The auto-generated name is rather ugly. See [example [EX021]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/021-basic-child-job). -### Basic Job With Data Type +### Basic job with data type To avoid automatic table names, it is advisable to always use the `dataType` property for child jobs: @@ -295,7 +297,7 @@ user-detail: See [example [EX022]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/022-basic-child-job-datatype). -### Basic Job With Array Values +### Basic job with array values It is also possible that the main job returns objects which contain direct references to the children: @@ -354,7 +356,7 @@ user-child: See [example [EX135]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/135-basic-child-job-array). -### Accessing Nested ID +### Accessing nested ID If the placeholder value is nested within the response object, you can use dot notation to access child properties of the response object. For instance, if the parent response with a list of users returns a response similar to this: @@ -425,7 +427,7 @@ Notice that the parent reference column name is the concatenation of the `parent See [example [EX023]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/023-child-job-nested-id). -### Accessing Deeply Nested Id +### Accessing deeply nested id The placeholder path is configured **relative to** the extracted object. Assume that the parent endpoint returns a complicated response like this: @@ -493,7 +495,7 @@ may be confusing because the endpoint property in that child job is set relative See [example [EX024]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/024-child-job-deeply-nested-id). -### Naming Conflict +### Naming conflict Because a new column is added to the table representing child properties, it is possible that you run into a naming conflict. That is, if the child response with user details looks like this: @@ -534,7 +536,7 @@ to create the column `parent_id` with the placeholder value, overwriting the ori See [example [EX025]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/025-naming-conflict). -### Nesting Level +### Nesting level By default, the placeholder value is taken from the object retrieved in the parent job. As long as the child jobs are nested only one level deep, there is no other option anyway. Let's see what happens with a deeper nesting. @@ -677,7 +679,7 @@ Notice that each table contains additional columns with the placeholder property See [example [EX026]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/026-basic-deeper-nesting). -### Nesting Level Alternative +### Nesting level alternative Because the required user and order IDs are present in multiple requests (in the list and in the detail), there are multiple ways how the jobs may be configured. For example, the following configuration produces the exact same result as the above configuration: @@ -766,7 +768,7 @@ the deepest child will really contain the `orderId` value. See [example [EX027]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/027-basic-deeper-nesting-alternative). -### Deep Job Nesting +### Deep job nesting Let's look at how to retrieve more nested API resources: ```json @@ -874,7 +876,7 @@ where the `parent_id` column refers the `5:user-id` placeholder. See [example [EX028]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/028-advanced-deep-nesting). -### Nested Array +### Nested array Suppose now that the endpoint `/users` returns a more complicated response: @@ -996,7 +998,7 @@ The `users-2\_members\_items` contains the same results as the `users` table, bu This makes the response in the `users` table quite useless, but the job is still required to generate the child jobs to obtain the `user-detail` table. See [example [EX106]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/106-child-jobs-array). -### Simple Filter +### Simple filter Let's assume that you have an API which has two endpoints: - `users` — Returns a list of users. @@ -1091,7 +1093,7 @@ the details are retrieved only for the desired users. See [example [EX029]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/029-simple-filter). -### Not Like Filter +### Not like filter Apart from the standard comparison operators, the recursive filter allows to use a **like** comparison operator `~`. It expects that the value contains a placeholder `%`, which matches any number of characters. The following configuration: @@ -1127,7 +1129,7 @@ following `user-detail` table will be extracted: See [example [EX030]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/030-not-like-filter). -### Combining Filters +### Combining filters Multiple filters can be combined using the [logical](https://en.wikipedia.org/wiki/Boolean_algebra#Basic_operations) `&` (and) and `|` (or) operators. For example, the following configuration retrieves details for users who have @@ -1160,7 +1162,7 @@ The following `user-detail` will be produced: See [example [EX031]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/031-combined-filter). -### Multiple Filter Combinations +### Multiple filter combinations Although you can join a multiple filter expression with logical operators as in the above example, there is no support for parentheses. The following configuration combines multiple filters: diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md index 5d427b066..72bd8a321 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/config/jobs/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/config/jobs/ --- + + *If new to Generic Extractor, learn about [jobs in our tutorial](/components/extractors/generic-extractor/tutorial/jobs/) first.* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various @@ -57,7 +59,7 @@ way. Each response is processed in the following steps: 3. Flatten the object structure into one or more tables. 4. Create the required tables in Storage and load data into them. -## Merging Responses +## Merging responses The first two steps are the responsibility of [Jobs](/components/extractors/generic-extractor/configuration/config/jobs/) resulting in an array of objects. Generic Extractor then tries to find a common super-set of properties of all objects, for example, with the following response: @@ -125,23 +127,23 @@ Assume the following [API definition](/components/extractors/generic-extractor/c } ``` -### Relative URL Fragment +### Relative URL fragment The relative endpoint **must not start** with a slash; so, with `endpoint` set to `campaign`, the final resource URL would be `https://example.com/3.0/campaign`. -### Absolute Domain URL +### Absolute domain URL The absolute endpoint **must start** with a slash. So, with `/endpoint` set to `campaign`, the final resource URL would be `https://example.com/campaign`. This means that the path part specified in the `baseURL` is ignored and fully replaced by the value specified in `endpoint`. -### Absolute Full URL +### Absolute full URL The full absolute URL must start with a protocol. So, with the endpoint set to `https://eu.example.com/campaign`, this would be the final resource URL and the path specified in the `baseURL` is completely ignored. -### Specifying Endpoint +### Specifying endpoint The following table summarizes possible outcomes: |`baseURL`|`endpoint`|actual URL| @@ -165,7 +167,7 @@ Also, closely follow the target API specification regarding trailing slashes. Fo both `https://example.com/3.0/campaign` and `https://example.com/3.0/campaign/` URLs may be accepted and valid. For other APIs, however, only one version may be supported. -## Request Parameters +## Request parameters The `params` section defines [request parameters](/components/extractors/generic-extractor/tutorial/rest). They may be optional or required, depending on the target API specification. The `params` section is an object with arbitrary properties (or, more precisely, parameters understood by the target @@ -255,7 +257,7 @@ or, in a more readable [URLDecoded](https://urldecode.org/) form: Also, the `Content-Type: application/x-www-form-urlencoded` HTTP header will be added to the request. -## Data Type +## Data type The `dataType` parameter assigns a name to the object(s) obtained from the endpoint. Setting it is optional. If not set, a name will be generated automatically from the `endpoint` value and parent jobs. @@ -283,7 +285,7 @@ for example, in a situation where two API endpoints return the same resource: In the above case, only a single `tickets` table will be produced in the output bucket. It will contain records from both API endpoints. -## Data Field +## Data field The `dataField` parameter is used to determine what part of the API **response** will be extracted. The following rules apply by default: @@ -318,7 +320,7 @@ as an object with the `path` property. For instance, these two configurations ar ] ``` -### Data Field Delimiter +### Data field delimiter The path to the response property is by default expected to be dot separated. That is — a path `members.active` refers to the property `active` nested inside the property `members`. If you need to refer to a property containing a dot, you have to change the data field path delimiter to some other character. This can be @@ -354,7 +356,7 @@ inside the property `members.active` you have to use: The `delimiter` character is completely arbitrary but must be something that is not used in the property names in the response. See [example [EX120]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/120-datafield-separator). -## Response Filter +## Response filter The `responseFilter` option allows you to skip parts of the API response from processing. This can be useful in these cases: @@ -685,7 +687,7 @@ The following table will be extracted: See [example [EX009]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/009-nested-array). -## Examples with Complicated Objects +## Examples with complicated objects The above examples show how simple objects are extracted from different objects. Generic Extractor can also extract objects with non-scalar properties. The default [JSON to CSV mapping](/components/extractors/generic-extractor/configuration/config/mappings/) flattens nested objects and produces secondary tables from nested arrays. @@ -908,7 +910,7 @@ auto-generated key to the parent *Users* table. Also notice that the See [example [EX012]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/012-deeply-nested-object). -## Response Filter Examples +## Response filter examples ### Skip flattening If you have an API response like this: @@ -1038,7 +1040,7 @@ See [example [EX014]](https://github.com/keboola/generic-extractor/tree/master/d + *If you are new to Generic Extractor, learn about [mapping in our tutorial](/components/extractors/generic-extractor/tutorial/mapping/) first.* *Use the [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various configuration options.* @@ -70,12 +72,12 @@ The following configuration shows a sample mapping configuration for dataType `u } ``` -### User Interface +### User interface In the UI, the mapping can be created for each endpoint in the `Endpoints`.`Mapping section` by clicking `Create Mapping` toggle. ![Create mapping](/components/extractors/generic-extractor/tutorial/create_mapping_toggle.png) -#### Mapping Detection +#### Mapping detection You may opt to generate the mapping automatically by clicking the `Infer Mapping` button in the top right corner. @@ -94,7 +96,7 @@ Currently, the automatic detection outputs only single table mapping. You can co the `Nesting Level` property. For example, a depth of 1 transforms `{"address": {"street": "Main", "details": {"postcode": "170 00"}}}` into two columns: `address_street` and `address_details`. All elements that have ambiguous types or are beyond the specified depth are stored in a single column as JSON, e.g., with the [`force_type`](/components/extractors/generic-extractor/configuration/config/mappings/#mapping-without-processing) option. -### Column Mapping +### Column mapping Column mapping represents a basic mapping type that allows you to select extracted columns, rename them, and optionally set a primary key on them. The mapping configuration requires: @@ -106,12 +108,12 @@ configuration requires: - `forceType` (optional, boolean) — If set to `true`, the property will not be processed and will be stored as an encoded JSON (see an [example](#mapping-without-processing)). -### User Mapping +### User mapping User mapping has the same configuration as the [column mapping](#column-mapping). The only difference is that it applies to *virtual properties*. This is useful mainly for working with auto-generated properties/columns in child jobs (see an [example](#mapping-child-jobs)). -### Table Mapping +### Table mapping Table mapping allows you to create a new table from a particular property of the response object. Table mapping is, by default, used for arrays. The mapping configuration requires: @@ -155,7 +157,7 @@ The following configuration takes the `contacts` property from the response and ## Examples The following examples demonstrate how to map JSON responses to CSV files. -### Automatic Mapping +### Automatic mapping Without any configuration, the following JSON response: ```json @@ -218,7 +220,7 @@ when the API returns a completely empty response in which case no tables are cre When Manual mapping is used, the generated table structure always honors the mapping setting. See [example [EX137]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/137-mapping-tables-nested-empty). -### Basic Manual Mapping +### Basic manual mapping Maybe you are not interested in the user `interests` and want to simplify the user table to three columns: `country`, `name` and `id`. The following mapping configuration does the trick: @@ -286,7 +288,7 @@ correct settings, the following table will be produced: See [example [EX064]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/064-mapping-basic). -### Mapping Child Jobs +### Mapping child jobs Let's say that you have an API endpoint `/users` which returns a response similar to: ```json @@ -385,7 +387,7 @@ column `parent_id` does not really exist in the response as it is generated dyna See [example [EX065]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/065-mapping-child-jobs). -### Mapping without Processing +### Mapping without processing The `forceType` configuration property allows you to skip a part of the API response from processing. With the following API response: @@ -475,7 +477,7 @@ The same result can be achieved by using the [`responseFilter` job property](/co See [example [EX073]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/073-mapping-forceType). -### Table Mapping Examples +### Table mapping examples #### Basic table mapping Because all output columns must be listed in a mapping, using only column mapping settings skips diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/index.md index 795f146b5..0c139d55f 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* @@ -13,7 +15,7 @@ To get an overall idea of what to expect when configuring Generic Extractor, loo Then review a [sample configuration](#configuration-map) featuring all configuration options and their nesting. The **configuration map** is also available as a [separate article](/components/extractors/generic-extractor/map/). -### User Interface +### User interface Recently, we created a convenient user interface that allows you to build a configuration for the Generic Extractor without writing JSON code. You can set up and test the connection in a few clicks, just like you are used to in some other popular API development tools. @@ -32,7 +34,7 @@ In such cases, you will be notified in the UI what sections are not supported. ***NOTE:** The new UI does not affect the functionality of old configurations. All configurations will continue to work. However, in some cases, you might need to perform some manual adjustments in order to make the UI compatible.* -### JSON Configuration Sections +### JSON configuration sections *Click on the section names if you want to learn more.* - **parameters** @@ -77,7 +79,7 @@ Generic Extractor can be run from within the [**Keboola user interface**](/compo configuration [JSON](/components/extractors/generic-extractor/tutorial/json/) is needed), or [**locally**](/components/extractors/generic-extractor/running/#running-locally) (Docker is needed). -### Configuration Map +### Configuration map The following sample configuration shows various configuration options and their nesting. You can use the map to navigate between them. The parameter map is also available [separately](/components/extractors/generic-extractor/map/), and we recommend pinning it to your toolbar for quick reference. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md index dac9a6a1d..5a74d7f0d 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/iterations/index.md @@ -6,6 +6,8 @@ redirect_from: - /extend/generic-extractor/iterations/ --- + + The `iterations` section allows you to **execute a configuration multiple times, each time with different values**. The most typical use for `iterations` is extraction of the same data from multiple accounts. @@ -70,7 +72,7 @@ is honoured. ## Examples -### Iterating Parameters +### Iterating parameters Suppose, you have an API which takes a URL parameter `account_id`, which restricts the returned data to a certain account. The following configuration executes the entire configuration for two accounts — `345` and `456`: @@ -156,7 +158,7 @@ It looks as if the first execution is with `account_id=123`, but it is not the c will be executed only twice: the first time with `account_id=345` and the second time with `account_id=456`. See [example [EX112]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/112-iterations-params). -### Iterating Headers +### Iterating headers Suppose you have an API from which you want to extract data from two accounts (`JohnDoe` and `DoeJohn`). The API uses the [HTTP Basic Authentication](/components/extractors/generic-extractor/configuration/api/authentication/basic/) method, and in addition, each user has their own API token, which must be provided in the `X-Api-Token` header. diff --git a/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md b/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md index 6be7eabff..644fb8b55 100644 --- a/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md +++ b/src/content/docs/components/extractors/generic-extractor/configuration/ssh-proxy/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/configuration/ssh-proxy/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* *Use [Parameter Map](/components/extractors/generic-extractor/map/) to help you navigate among various @@ -33,7 +35,7 @@ to act as a gateway to your private network where your destination server reside Complete the following steps to set up an SSH proxy for Generic Extractor: -### 1. Set Up SSH Proxy Server +### 1. Set up SSH proxy server Here is a very basic [Dockerfile](https://docs.docker.com/engine/reference/builder/) example. All it does is run an sshd daemon and expose port 22. You can, of course, set this up in your system in a similar way without using Docker. @@ -66,7 +68,7 @@ See the following pages for more information about setting up SSH on your server - [OpenSSH configuration](https://help.ubuntu.com/community/SSH/OpenSSH/Configuring) - [Dockerized SSH service](https://docs.docker.com/engine/examples/running_ssh_service/) -### 2. Generate SSH Key Pair +### 2. Generate SSH key pair Generate an SSH key pair and copy the public key to your **SSH proxy server**. Paste it to the **public.key** file, and then append it to the authorized_keys file. @@ -75,7 +77,7 @@ mkdir ~/.ssh cat public.key >> ~/.ssh/authorized_keys ``` -### 3. Configure Generic Extractor SSH Proxy +### 3. Configure Generic Extractor SSH proxy ```json { diff --git a/src/content/docs/components/extractors/generic-extractor/functions/index.md b/src/content/docs/components/extractors/generic-extractor/functions/index.md index 61a0e0c45..e4da51b21 100644 --- a/src/content/docs/components/extractors/generic-extractor/functions/index.md +++ b/src/content/docs/components/extractors/generic-extractor/functions/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/functions/ --- + + Functions are simple pre-defined functions that @@ -89,7 +91,7 @@ These forms can be combined freely. They can be also nested in a virtually unlim } ``` -### User Interface +### User interface You may create functions in the user interface's `User Parameters` or `User Data` sections. You can also create the functions directly from other configuration contexts, e.g., when defining the query parameters on the endpoint. @@ -101,7 +103,7 @@ The UI also offers a convenient way to evaluate the function and see the results ![img.png](/components/extractors/generic-extractor/function_eval.gif) -## Supported Functions +## Supported functions ### md5 The [`md5` function](https://www.php.net/manual/en/function.md5.php) calculates the [MD5 hash](https://en.wikipedia.org/wiki/MD5) of a @@ -383,11 +385,11 @@ considered 'empty'. See an [example](#optional-job-parameters). -## Function Contexts +## Function contexts Every place in the Generic Extractor configuration in which a function may be used may allow different arguments of the function. This is referred to as a **function context**. Many contexts share access to **configuration attributes**. -### Configuration Attributes +### Configuration attributes The configuration attributes are accessible in specific function contexts and they represent the entire [`config`](/components/extractors/generic-extractor/configuration/config/) section of the Generic Extractor configuration. There is some processing involved: @@ -462,20 +464,20 @@ will be converted to the following function context: See [example [EX119]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/119-function-nested-config). -### Base URL Context +### Base URL context The Base URL function context is used when setting the [`baseURL` for API](/components/extractors/generic-extractor/configuration/api/#base-url), and it contains [configuration attributes](/components/extractors/generic-extractor/functions/#function-contexts). See an [example](#api-base-url). -### Headers Context +### Headers context The Headers function context is used when setting the [`http.headers` for API](/components/extractors/generic-extractor/configuration/api/#headers) or the [`http.headers` in config](/components/extractors/generic-extractor/configuration/config/#http), and it contains [configuration attributes](/components/extractors/generic-extractor/functions/#function-contexts). See an [example](#headers). -### Parameters Context +### Parameters context The Parameters function context is used when setting job [request parameters — `params`](/components/extractors/generic-extractor/configuration/config/jobs/#request-parameters). It contains [configuration attributes](/components/extractors/generic-extractor/functions/#function-contexts) plus the times of the current (`currentStart`) and previous (`previousStart`) run of Generic Extractor. @@ -522,7 +524,7 @@ See an [example of using parameters context](#job-parameters). The `time` values are used in [incremental processing](/components/extractors/generic-extractor/incremental/). -### Placeholder Context +### Placeholder context The Placeholder function context refers to configuration of [placeholders in child jobs](/components/extractors/generic-extractor/configuration/config/jobs/children/#placeholders). When using function to process a placeholder value, the placeholder must be specified as an object with the `path` property. Therefore instead of writing: @@ -559,7 +561,7 @@ of the placeholder. See an [example](#job-placeholders). -### User Data Context +### User data context The User Data function context is used when setting the [`userData`](/components/extractors/generic-extractor/configuration/config/#user-data). The parameters context contains [configuration attributes](/components/extractors/generic-extractor/functions/#function-contexts) plus the times of the current (`currentStart`) and previous (`previousStart`) run of Generic Extractor. The User Data Context is therefore @@ -567,7 +569,7 @@ same as the [Parameters Context](#parameters-context). See an [example](#user-data). -### Login Authentication Context +### Login authentication context The Login Authentication function context is used in the [login authentication](/components/extractors/generic-extractor/configuration/api/authentication/login/) method. Functions are supported in both [`loginRequest`](/components/extractors/generic-extractor/configuration/api/authentication/login/#configuration-parameters) @@ -609,7 +611,7 @@ See an [example](/components/extractors/generic-extractor/configuration/api/auth [complicated example](/components/extractors/generic-extractor/configuration/api/authentication/login/#login-authentication-with-login-and-api-request) of using functions in both login request and API request. -### Query Authentication Context +### Query authentication context The Query Authentication function context is used in the [query authentication](/components/extractors/generic-extractor/configuration/api/authentication/query/) method. The Query Authentication Context contains [configuration attributes](/components/extractors/generic-extractor/functions/#function-contexts) plus @@ -687,7 +689,7 @@ leads to the following function context: See the [basic example](#api-default-parameters) and a [more complicated example](#api-query-authentication). -### OAuth 2.0 Authentication Context +### OAuth 2.0 authentication context The OAuth Authentication Context is used for the [`oauth20`](/components/extractors/generic-extractor/configuration/api/authentication/oauth20/) authentication method (it is not applicable to `oauth10`) and contains the following: @@ -786,7 +788,7 @@ the application is published) are added to the `authorization` section. For usage, see [OAuth examples](/components/extractors/generic-extractor/configuration/api/authentication/oauth20/). -### OAuth 2.0 Login Authentication Context +### OAuth 2.0 login authentication context The OAuth Login Authentication Context is used for the [`oauth20.login`](/components/extractors/generic-extractor/configuration/api/authentication/oauth20-login/) authentication method (it is not applicable to `oauth20`). The OAuth Login Authentication context contains @@ -853,7 +855,7 @@ For usage, see [OAuth Login examples](/components/extractors/generic-extractor/c ## Examples -### API Base URL +### API base URL When [publishing your Generic Extractor configuration](/components/extractors/generic-extractor/publish/), chances are you want the end-user to provide a part of the API configuration. Due to the limitations of [how templates work](/components/extractors/generic-extractor/publish/#configuration-considerations), the parameter @@ -911,7 +913,7 @@ final API URL (`http://example.com/api/1.0/`): See [example [EX087] with concat](https://github.com/keboola/generic-extractor/tree/master/doc/examples/087-function-baseurl) or an alternative [example [EX088] with sprintf](https://github.com/keboola/generic-extractor/tree/master/doc/examples/088-function-baseurl-sprintf). -### API Default Parameters +### API default parameters Suppose you have an API which expects a `tokenHash` parameter to be sent with every request. The token hash is supposed to be generated by the SHA-256 hashing algorithm from a token and secret you obtain. @@ -1009,7 +1011,7 @@ the single `users` job. See [example [EX098]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/098-function-hmac). -### API Query Authentication +### API query authentication Suppose you have an API with only a single endpoint `/items` to which you have to pass a `type` parameter to list resources of a given type. On top of that, the API requires an `apiToken` parameter and a `signature` parameter (a hash of the token and type) to be sent with every request. @@ -1082,7 +1084,7 @@ token and resource type (`"query": "type"` is taken from the `jobs.params.type` See [example [EX101]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/101-function-query-auth). -### Job Placeholders +### Job placeholders Let's say you have an API with an endpoint `/users`, returning a list of users, and an endpoint `/user/{userId}`, returning details of a specific user with a given ID. The list response looks like this: @@ -1144,7 +1146,7 @@ See [example [EX085]](https://github.com/keboola/generic-extractor/tree/master/d or a not-so-useful [example [EX086]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/086-function-job-placeholders-reference) (using reference). -### Job Parameters +### Job parameters Let's say you have an API which requires you to send a hash of a certain value with every request. Specifically, each request must be done with the [HTTP POST method](/components/extractors/generic-extractor/tutorial/rest/#method) with content: @@ -1194,7 +1196,7 @@ See [example [EX089]](https://github.com/keboola/generic-extractor/tree/master/d or an alternative [example [EX090] with SHA1 hash](https://github.com/keboola/generic-extractor/tree/master/doc/examples/090-function-job-parameters-sha1). or an alternative [example [EX136]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/136-post-request-functions) with more deeply nested functions. -### Optional Job Parameters +### Optional job parameters Let's say you have an API which allows you to send the list of columns to be contained in the API response. For example, to list users and include their `id`, `name` and `login` properties, call `/users?showColumns=id,name,login`. Also, you want to enter these values as an array in the `config` section because @@ -1239,7 +1241,7 @@ The following configuration does exactly that: See [example [EX097]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/097-function-ifempty). -### User Data +### User data Assume that you have an API returning a response that does not contain any time information. For example: ```json @@ -1368,7 +1370,7 @@ in the [`config` section](/components/extractors/generic-extractor/configuration See [example [EX093]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/093-function-api-http-headers) or an [alternative example [EX094] setting headers in the `config` section](https://github.com/keboola/generic-extractor/tree/master/doc/examples/094-function-config-headers). -### Nested Functions +### Nested functions If the API in the [above example](#headers) tries to mimic the [HTTP authentication](/components/extractors/generic-extractor/configuration/api/authentication/basic/), the header has to be sent as a [base64 encoded](https://en.wikipedia.org/wiki/Base64#MIME) value. diff --git a/src/content/docs/components/extractors/generic-extractor/incremental/index.md b/src/content/docs/components/extractors/generic-extractor/incremental/index.md index a1303d701..a3cd8bb6f 100644 --- a/src/content/docs/components/extractors/generic-extractor/incremental/index.md +++ b/src/content/docs/components/extractors/generic-extractor/incremental/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/incremental/ --- + + Extracting data incrementally is universally beneficial — it **speeds up the extraction** and **lowers the load** on both the API and [Keboola Storage](/storage/) (thus saving @@ -39,7 +41,7 @@ loads by using [`previousStart`](/components/extractors/generic-extractor/functi ## Examples -### Previous Start Example +### Previous start example Assume you have an API supporting a parameter `modified_since` which expects a [Unix Timestamp](https://en.wikipedia.org/wiki/Unix_time). The response then contains only the records that were modified after the specified date. The following configuration can be used: @@ -82,7 +84,7 @@ The last successful time is stored in the [configuration state](/extend/common-i If for some reason you need to reset it, [update the configuration via API](https://api.keboola.com/?service=storage#put-/v2/storage/branch/-branchId-/components/-componentId-/configs/-configurationId-). -### Previous Start Date +### Previous start date If an API similar to the one in the [above example](#previous-start-example) requires the date to be sent as a string, the following jobs configuration (which uses the [`date` function](/components/extractors/generic-extractor/functions/#date)) can be used: @@ -121,7 +123,7 @@ Otherwise the configuration behaves the same way as the [previous example](#prev See [example [EX108]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/108-incremental-load-date). -### Incremental Load From To +### Incremental load from to Another option is an API which requires the `from` and `to` parameters. The following configuration generates the `from` date as the date of the last extraction (using the [`time.previousStart` value](/components/extractors/generic-extractor/functions/#parameters-context)). It also generates the `to` date as the date @@ -164,7 +166,7 @@ This configuration will send a request similar to this one: See [example [EX109]](https://github.com/keboola/generic-extractor/tree/master/doc/examples/109-incremental-load-from-to). -### Incremental Relative Load +### Incremental relative load Suppose you have an API supporting the `from` and `to` parameters as in [the above example](#incremental-load-from-to) and want to extract the last day data. It can be done using the following configuration: diff --git a/src/content/docs/components/extractors/generic-extractor/index.md b/src/content/docs/components/extractors/generic-extractor/index.md index 7d56c2c9a..8b390b02f 100644 --- a/src/content/docs/components/extractors/generic-extractor/index.md +++ b/src/content/docs/components/extractors/generic-extractor/index.md @@ -6,6 +6,8 @@ redirect_from: - /components/extractors/other/generic/ --- + + Generic Extractor is a [Keboola component](/overview/) that acts like a customizable [HTTP REST](/components/extractors/generic-extractor/tutorial/rest/) client. It can be configured to extract data @@ -21,7 +23,7 @@ an entirely new extractor for Keboola in **less than an hour**. To get started quickly, follow our [Generic Extractor tutorial](/components/extractors/generic-extractor/tutorial). -## Generic Extractor Requirements +## Generic Extractor requirements Generic Extractor allows you to extract data from an API into Keboola only by configuring it. No programming skills or additional tools are required. You just need to do two easy things before you start: @@ -29,7 +31,7 @@ No programming skills or additional tools are required. You just need to do two - Have the documentation of your chosen API at hand. The API should be [RESTful](/components/extractors/generic-extractor/tutorial/rest/) and, more or less, follow the HTTP specification. -## Configuration & Development +## Configuration & development Again, if you are new to Generic Extractor, we strongly suggest you go through the [Generic Extractor tutorial](/components/extractors/generic-extractor/tutorial/). It outlines the basic principles and the most important features. @@ -41,7 +43,7 @@ Features such as cURL import, request tests, output mapping generator, or dynami If you intend to develop a more complicated configuration, check out how to [run Generic Extractor locally](/components/extractors/generic-extractor/running/). The documentation includes [several examples](https://github.com/keboola/generic-extractor/tree/master/doc) that [can also be run locally](/components/extractors/generic-extractor/running/#running-examples). -## Publishing Generic Extractor Configuration +## Publishing Generic Extractor configuration Each Generic Extractor configuration can be [published](/components/extractors/generic-extractor/publish/) as a new standalone component. However, for registration, configurations must be [converted to templates](/components/extractors/generic-extractor/publish/#publishing). @@ -53,7 +55,7 @@ do not limit the configuration. You can always switch to JSON Also, templates can be used only with published components based on Generic Extractor configurations. -## Template Mode +## Template mode Generic Extractor is used as the base for many data source connectors. These components allow you to select pre-defined configurations -- templates -- without the need to configure Generic Extractor manually. @@ -67,7 +69,7 @@ The code will be pre-filled for you based on that template. When finished editing, save the configuration. You can always switch back to the templates, but you'll lose your customizations. -## Generic Extractor Source +## Generic Extractor source As with other Keboola components, the Generic Extractor connector is available on [GitHub](https://github.com/keboola/generic-extractor/). Apart from the main repository, it uses some vital libraries (which partially define its capabilities): diff --git a/src/content/docs/components/extractors/generic-extractor/map/index.md b/src/content/docs/components/extractors/generic-extractor/map/index.md index fb6995bb3..8ecc86d7d 100644 --- a/src/content/docs/components/extractors/generic-extractor/map/index.md +++ b/src/content/docs/components/extractors/generic-extractor/map/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/map/ --- + + *To configure your first Generic Extractor, follow our [tutorial](/components/extractors/generic-extractor/tutorial/).* Use the following sample configuration to navigate among various **configuration options**: diff --git a/src/content/docs/components/extractors/generic-extractor/publish/index.md b/src/content/docs/components/extractors/generic-extractor/publish/index.md index 03dfb1371..344158e32 100644 --- a/src/content/docs/components/extractors/generic-extractor/publish/index.md +++ b/src/content/docs/components/extractors/generic-extractor/publish/index.md @@ -6,11 +6,13 @@ redirect_from: - /extend/generic-extractor/registration/ --- + + It is possible to publish a Generic Extractor configuration as a completely separate component. This enables sharing the API extractor between various projects and simplifies its further configuration. -## Configuration Considerations +## Configuration considerations Before converting your configuration to a universally available component, consider what values in the configuration should be provided by the end-user (typically authentication values). Then design a [configuration schema](/extend/component/ui-options/configuration-schema/) for setting diff --git a/src/content/docs/components/extractors/generic-extractor/running/index.md b/src/content/docs/components/extractors/generic-extractor/running/index.md index 7696c6358..0b3101c99 100644 --- a/src/content/docs/components/extractors/generic-extractor/running/index.md +++ b/src/content/docs/components/extractors/generic-extractor/running/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/running/ --- + + Generic Extractor is normally run from within the Keboola user interface. It can be found in the **Extractors** section and all you need to do is provide its configuration JSON. No other settings are necessary. @@ -12,7 +14,7 @@ and all you need to do is provide its configuration JSON. No other settings are Because creating the configuration JSON can be a non-trivial task, there are some things which can help you in developing the configuration. -## Debug Mode +## Debug mode Debug mode can be turned on by setting `"debug": true` in the `config` section of the configuration, e.g.: ```json @@ -36,13 +38,13 @@ why something is skipped, etc. visible in the events. Also, debug mode considerably slows the extraction. Therefore it should never be turned on in production configurations. -## Running Locally +## Running locally If you are working on a complicated configuration, or developing a new component based on Generic Extractor, running every configuration from the Keboola UI may be slow and tedious. You may run Generic Extractor locally, provided that you have access to Docker. The following is **not necessary** to run or configure Generic Extractor in Keboola. -### Run Built Version +### Run built version Create an empty directory somewhere and in it create a `config.json` file with a configuration you want to execute. For example: @@ -119,7 +121,7 @@ When you store such configuration in the Keboola UI, it will automatically be en The above configuration then **cannot** be run locally. Read more about [encryption](/extend/encryption/). -### Building and Running the Image +### Building and running the image To build the container from source: - Clone this repository: `git clone https://github.com/keboola/generic-extractor.git`. @@ -137,7 +139,7 @@ To run the built container: Before running the extractor again, it is recommended to clear the `out` directory by running `docker compose run --rm extractor rm -rf data/out`. -## Running Examples +## Running examples [All examples](https://github.com/keboola/generic-extractor/tree/master/doc) referenced in this documentation are actually runnable against the proper API. Because it is difficult to find the specific API for the case (and gain access to it), you can test these configurations against a [mock server](https://github.com/keboola/ex-generic-mock-server). diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md index fd23c1cfd..7896e75a5 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/basic/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/basic/ --- + + Before configuring Generic Extractor, you should have a basic understanding of [REST API](/components/extractors/generic-extractor/tutorial/rest/) and @@ -21,7 +23,7 @@ and comprises [several sections](/components/extractors/generic-extractor/config A [user interface](/components/extractors/generic-extractor/configuration/#user-interface) is available that can help you with the configuration and generate the JSON configuration for you. -### Base Configuration +### Base configuration The first configuration part is a `Base Configuration` section where you can set the Base URL and Authentication method of the API you connect to. @@ -72,7 +74,7 @@ configuration section: The `password` property is prefixed with the hash mark `#`, meaning the value will be [encrypted](/extend/encryption/) once you save the configuration. -### Endpoint Section +### Endpoint section Once you set up the Base Configuration, you can set up the actual endpoint to be queried. Start by clicking the **+ New Endpoint** button: @@ -169,7 +171,7 @@ represented in a single column of the `campaigns` table. Generic Extractor, ther value with a generated key, for example, `campaigns_75d5b14d79d034cd07a9d95d5f0ca5bd`, and automatically creates a new table that has the column `JSON_parentId` with that value so that you can join the tables together. -### Final JSON Configuration +### Final JSON configuration The main parts of the configuration and their nesting are shown in the following schema: ![Schema - Generic Extractor configuration](/components/extractors/generic-extractor/generic-intro.png) @@ -216,3 +218,5 @@ of doing much more; see other parts of this tutorial for an explanation of pagin (resources) to be extracted. - [Mapping](/components/extractors/generic-extractor/tutorial/mapping/) — describes how the JSON response is converted into CSV files that will be imported into Storage. + +**Next:** [Pagination →](/components/extractors/generic-extractor/tutorial/pagination/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/index.md index 945d4e0e2..2aca02d68 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/ --- + + In this tutorial, we will guide you through configuring Generic Extractor for a new API. In our case, MailChimp — an email marketing service. @@ -34,7 +36,7 @@ data**: [generate your API Key](https://mailchimp.com/help/about-api-keys/#Find-or-Generate-Your-API-Key). It will look like this: `c40xxxxxxxxxxxxxxxxxxxxxxxxxxxxx-us13`. -## Get Started +## Get started Let's take a closer look at the [MailChimp API](https://mailchimp.com/developer/) now. There are plenty of documentation guides available. To explore the API and review what information is in each resource, use, for example, the [Playground](https://us1.api.mailchimp.com/playground/). @@ -69,7 +71,7 @@ requests and responses. The response body is in [JSON](/components/extractors/ge ... ``` -## Next Steps +## Next steps Now you have everything you need to actually start extracting the data. Continue with your Generic Extractor configuration here: @@ -79,4 +81,6 @@ configuration here: - [Jobs](/components/extractors/generic-extractor/tutorial/jobs/) — describes the API endpoints (resources) to be extracted. - [Mapping](/components/extractors/generic-extractor/tutorial/mapping/) — describes how the JSON - response is converted into CSV files that will be imported into Storage. \ No newline at end of file + response is converted into CSV files that will be imported into Storage. + +**Next:** [REST HTTP API introduction →](/components/extractors/generic-extractor/tutorial/rest/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md index 2aae6e239..3bfa237be 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/jobs/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/jobs/ --- + + On your way through the Generic Extractor tutorial, you have learned about @@ -24,7 +26,7 @@ Moreover, each campaign has three **sub-resources**: and `/campaigns/{campaign_id}/send-checklist`. The `{campaign_id}` expression represents a placeholder that a specific campaign ID should replace. To retrieve the sub-resource, use child jobs. -## Child Jobs +## Child jobs In the [previous part](/components/extractors/generic-extractor/tutorial/pagination/#running) of the tutorial, you created this job @@ -146,7 +148,7 @@ request. So, to join the two tables together in SQL, you would use the join cond However, you have to remember what table the `parent_id` column refers to. -## Multiple Jobs +## Multiple jobs You have probably noticed that the `jobs` and `children` properties are arrays. It means that you can retrieve multiple endpoints in a single configuration. Let's pick the campaign `content` sub-resource too: @@ -238,3 +240,5 @@ MailChimp API and giving us a lot of trouble, is best to be ignored. The answer You might also have noticed some duplicate records in the table `in.c-ge-tutorial.campaigns__campaign_id__content` along the way. You'll look into this as well. + +**Next:** [Mapping →](/components/extractors/generic-extractor/tutorial/mapping/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md index 13f3e7dd1..8d7703bc1 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/json/index.md @@ -5,12 +5,14 @@ redirect_from: - /extend/generic-extractor/tutorial/json/ --- + + [JSON (JavaScript Object Notation)](http://www.json.org/) is an easy-to-work-with format for describing structured data. Before you start working with JSON, familiarize yourself with basic programming jargon. It is also recommended to have a text editor with JSON support (you can also use an [online editor](http://www.jsoneditoronline.org/)). -## Object Representation +## Object representation To describe structured data, JSON uses **objects** and **arrays**. ### Objects @@ -65,7 +67,7 @@ The terminology varies a lot and other expressions are also commonly used: - Property — also a field / key / index - Array — also a collection / list / vector / ordinal array / sequence -## Data Values +## Data values Each property value always has one of the following data types: - String — text @@ -145,3 +147,5 @@ The order of items in an object is not important. It is also worth noting that ` ## Summary This page contains a little introduction to JSON documents. We intentionally avoided many details, but you should now understand what JSON is, and how to write some stuff in it. + +**Next:** [Basic configuration →](/components/extractors/generic-extractor/tutorial/basic/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md index d7df96615..9e6e33994 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/mapping/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/mapping/ --- + + In the previous part of the tutorial, you [extracted the content of a MailChimp campaign](/components/extractors/generic-extractor/tutorial/jobs/). Now, it's time to clean up the response. @@ -215,7 +217,7 @@ Note that the `destination` value is arbitrary but must be a valid column name. The data type name (`content`) must match the value of the `dataType` property as defined in some jobs. -## Parent Reference +## Parent reference The above mapping works, but it is missing the campaign ID, and you cannot match the content to some campaign records. Therefore, you must extract the campaign ID from the context (i.e., the job parameter). This can be done using a special `user` mapping. @@ -329,9 +331,9 @@ the extracted data later in [Transformations](/transformations/). However, if you intend to use your configuration regularly or want to make it into a component, setting up a mapping is recommended. -## Tips and Tricks +## Tips and tricks -### Key Containing a Dot Character +### Key containing a dot character The key of the mapping supports dot notation to traverse into children. So, if the key contains a dot, you need to change the delimiter. See the following example: @@ -350,3 +352,5 @@ The key of the mapping supports dot notation to traverse into children. So, if t ``` As you changed the delimiter from the default `.` to `/`, it's no longer parsed as two separate keys `created` and `date`, but rather just a single key `created.date`. + +**Next:** [Configuration reference →](/components/extractors/generic-extractor/configuration/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md index 34084748e..a5f47a450 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/pagination/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/pagination/ --- + + Pagination breaks a result with a large number of items into separate pages and is used very commonly in many API calls. @@ -165,3 +167,5 @@ getting incomplete data. The next two parts of our tutorial deal with setting up (resources) to be extracted. - [Mapping](/components/extractors/generic-extractor/tutorial/mapping/) — describes how the JSON response is converted into CSV files that will be imported into Storage. + +**Next:** [Jobs →](/components/extractors/generic-extractor/tutorial/jobs/) diff --git a/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md b/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md index 1225930dc..83f96594d 100644 --- a/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md +++ b/src/content/docs/components/extractors/generic-extractor/tutorial/rest/index.md @@ -5,6 +5,8 @@ redirect_from: - /extend/generic-extractor/tutorial/rest/ --- + + An [API (Application Programming Interface)](https://en.wikipedia.org/wiki/Application_programming_interface) is an [interface](https://en.wikipedia.org/wiki/Interface_(computing)) to an application, or a **service** @@ -18,7 +20,7 @@ Used by web browsers and other API clients, it defines how two parties (client a - Client creates an HTTP **request** and sends it to the server over the network. - Server processes the request, creates a **response**, and sends it to the client over the network. -## HTTP Request +## HTTP request An HTTP request is composed of: - URL @@ -93,14 +95,14 @@ example is the `X-StorageAPIToken` header used with Keboola [Storage API](/stora The `POST`, `PUT` and `PATCH` requests can send parameters the same way as the `GET` requests in the URL. But they can also send them in the request **body**. These are sometimes called **POST data/postdata**. -## HTTP Response +## HTTP response An HTTP response is composed of: - Response Headers — same as the request headers (only sent by the server) - Response Body — actual content of the resource - Status Code — status of the request -#### HTTP Status +#### HTTP status The HTTP Status and [status code](https://en.wikipedia.org/wiki/List_of_HTTP_status_codes) represent a standardized way of describing the response state. For example, the status `200 OK` (200 is the status code) is associated with a successful response. There are many HTTP Statuses, but the following rules apply: @@ -134,3 +136,4 @@ to get responses from virtually any HTTP REST API. Since the REST rules are not is not possible to ensure that Generic Extractor will be capable of reading 100% of APIs, even when declared as RESTful by someone. +**Next:** [JSON introduction →](/components/extractors/generic-extractor/tutorial/json/)