Skip to main content

Web Scraper API

Web Scraper is a cloud web-scraping automation platform for extracting structured data from websites at scale. Its REST API lets you manage sitemaps, launch and monitor scraping jobs, inspect data quality, and download scraped results. Nexla connects to the Web Scraper Cloud API so you can ingest this data into your flows and push new sitemap and job configurations back to your account.

Web Scraper API icon

Power end-to-end data operations for your Web Scraper API API with Nexla. Our bi-directional Web Scraper API connector is purpose-built for Web Scraper API, making it simple to ingest data, sync it across systems, and deliver it anywhere — all with no coding required. Nexla turns API-sourced data into ready-to-use, reusable data products and makes it easy to send data to Web Scraper API or any other destination. With comprehensive monitoring, lineage tracking, and access controls, Nexla keeps your Web Scraper API workflows fast, secure, and fully governed.

Features

Type: API

SourceDestination

  • Seamless API Integration: Connect to any endpoint as source or destination without coding, with automatic data product creation
  • Visual Composition & Chaining: Build complex integrations using visual templates, chain API calls, and compose workflows with data validation and filtering
  • API Proxy: Expose curated slices of your data securely with a secure and customizable API proxy that validates and transforms data on the fly
  • Request optimization with intelligent batching, retry, and caching to minimize API calls and costs

Prerequisites

Before creating a Web Scraper credential, you need an API token from your Web Scraper Cloud account. The Web Scraper Cloud REST API is served from https://api.webscraper.io and authenticates every request with an API token, which can be sent either as a query parameter named api_token or as a Bearer token in the Authorization header. Nexla sends the token as the api_token query parameter.

To obtain your API token, follow these steps:

  1. Sign in to your Web Scraper Cloud account at cloud.webscraper.io.

  2. Navigate to the API section of your account (available at cloud.webscraper.io/api).

  3. Copy the existing API token displayed on the page, or generate a new token if one has not been created yet.

  4. Store the token securely, as you will need it to configure your Nexla credential. Treat the token as sensitive information; anyone with the token can access your Web Scraper account data.

For detailed information about the API, endpoints, and rate limits, see the Web Scraper Cloud API documentation.

Authenticate

Credentials required

Authenticate using your Web Scraper API key. Obtain your key from your account settings and use it with Bearer token authentication.

FieldRequiredSecretDescription
API KeyYesYesYour Web Scraper API key for authentication
Base URLYesNoThe base URL for the Web Scraper API.

Create a credential in Nexla

  1. After selecting the data source/destination type, click the Add Credential tile to open the Add New Credential overlay.

  2. Enter a name for the credential in the Credential Name field and a short, meaningful description in the Credential Description field.

  3. Enter your Web Scraper API token in the API Key field. This is the token you obtained in Prerequisites and must be kept confidential.

  4. Enter the API base URL in the Base URL field. Unless directed otherwise, use the default value https://api.webscraper.io.

  5. Click the Save button at the bottom of the overlay. The newly added credential will now appear in a tile on the Authenticate screen during data source/destination creation.

    If your API token is compromised, generate a new token in the API section of your Web Scraper Cloud account and update your Nexla credential. For detailed information about API authentication and available endpoints, see the Web Scraper Cloud API documentation.

Use as a data source

To create a new data flow, navigate to the Integrate section, and click the New Data Flow button. Select the Web Scraper API connector tile, then select the credential that will be used to connect to your Web Scraper account, and click Next; or, create a new Web Scraper API credential for use in this flow.

Endpoint templates

Nexla provides pre-built templates that can be used to rapidly configure data sources to ingest data from common Web Scraper API endpoints. Select the endpoint from which this source will fetch data from the Endpoint pulldown menu. Available endpoint templates are listed in the expandable boxes below.

[Rest API] List Sitemaps

Returns a list of all sitemaps in the Web Scraper account.

  • This endpoint is paginated. Nexla automatically iterates through the pages to retrieve the full list of sitemaps.

[Rest API] Get Account Information

Retrieves the current user account information and details.

[Rest API] List Scraping Jobs

Returns a list of all scraping jobs in the Web Scraper account.

  • This endpoint is paginated. Nexla automatically iterates through the pages to retrieve the full list of scraping jobs.

[Rest API] Get Scraping Job

Retrieves details of a specific scraping job by its ID.

  • Provide the Scraping Job ID of the scraping job whose details you want to retrieve.

[Rest API] Get Scraping Job Data Quality

Retrieves data quality metrics and statistics for a specific scraping job.

  • Provide the Scraping Job ID of the scraping job whose data quality metrics you want to retrieve.

[Rest API] Get Scraping Job JSON Data

Downloads the scraped data from a completed scraping job in JSON format.

  • Provide the Scraping Job ID of the completed scraping job whose JSON data you want to download.

[Rest API] List Scraping Job Problematic URLs

Returns a paginated list of problematic URLs encountered during a scraping job.

  • Provide the Scraping Job ID of the scraping job whose problematic URLs you want to list. This endpoint is paginated, and Nexla automatically iterates through the pages.

[Rest API] Get Sitemap Details

Retrieves detailed information about a specific sitemap.

  • Provide the Sitemap ID of the sitemap whose details you want to retrieve.

[Rest API] Get Sitemap Scheduler

Retrieves the scheduler configuration for a specific sitemap.

  • Provide the Sitemap ID of the sitemap whose scheduler configuration you want to retrieve.

[Rest API] Get Usage

Returns rate-limit usage and quota information for the current account.

Once the selected endpoint template has been configured, click the Test button to the right of the endpoint selection menu to retrieve a sample of the data that will be fetched. Sample data will be displayed in the Endpoint Test Result panel on the right, allowing you to verify that the source is configured correctly before saving.

Manual configuration

Web Scraper API data sources can also be manually configured to ingest data from any valid Web Scraper API endpoint, including endpoints not covered by the pre-built templates, chained API calls, or custom request parameters. Select the Advanced tab at the top of the configuration screen, and follow the instructions in Connect to Any API to configure the API method, endpoint URL, date/time and lookup macros, path to data, metadata, and request headers.

Once all of the relevant settings have been configured, click the Create button in the upper right corner of the screen to save and create the new Web Scraper API data source. Nexla will now begin ingesting data from the configured endpoint and will organize any data that it finds into one or more Nexsets.

Use as a destination

Click the + icon on the Nexset that will be sent to the Web Scraper API destination, and select the Send to Destination option from the menu. Select the Web Scraper API connector from the list of available destination connectors, then select the credential that will be used to connect to your Web Scraper account, and click Next; or, create a new Web Scraper API credential for use in this flow.

Endpoint templates

Nexla provides pre-built templates that can be used to rapidly configure destinations to send data to common Web Scraper API endpoints. Select the endpoint to which data will be sent from the Endpoint pulldown menu. Then, click on the template in the list below to expand it, and follow the instructions to configure additional endpoint settings.

[Rest API] Create Sitemap

Create a new sitemap for scraping a website.

  • Each record from your Nexset is sent as a JSON request body to create a new sitemap.

[Rest API] Update Sitemap

Update an existing sitemap configuration.

  • Provide the Sitemap ID of the sitemap to update. Each record from your Nexset is sent as a JSON request body with the updated configuration.

[Rest API] Delete Sitemap

Delete an existing sitemap.

  • Provide the Sitemap ID of the sitemap to delete.

[Rest API] Create Scraping Job

Create a new scraping job to scrape a sitemap.

  • Each record from your Nexset is sent as a JSON request body to launch a new scraping job.

[Rest API] Delete Scraping Job

Delete an existing scraping job.

  • Provide the Scraping Job ID of the scraping job to delete.

[Rest API] Enable Sitemap Scheduler

Activates the scheduler for a specific sitemap with the provided scheduler configuration.

  • Provide the Sitemap ID of the sitemap whose scheduler to enable. Each record from your Nexset is sent as a JSON request body with the scheduler configuration.

[Rest API] Disable Sitemap Scheduler

Deactivates the scheduler for a specific sitemap.

  • Provide the Sitemap ID of the sitemap whose scheduler to disable.

Manual configuration

Web Scraper API destinations can also be manually configured to send data to any valid Web Scraper API endpoint. Select the Advanced tab at the top of the configuration screen, and follow the instructions in Connect to Any API to configure the API method, data format, endpoint URL, request headers, attribute exclusions, record batching, and response webhooks.

Save & activate

Once all endpoint settings have been configured, click the Done button in the upper right corner of the screen to save and create the destination. To send the data to the configured Web Scraper API endpoint, open the destination resource menu, and select Activate.

The Nexset data will not be sent to the Web Scraper API endpoint until the destination is activated. Destinations can be activated immediately or at a later time, providing full control over data movement.