Turn a website into a REST API
This walkthrough takes you from a public collection URL to a bounded API request in cURL or Python. You will review the source, validate fields, and check whether returned records satisfy the scope you selected.
On this page
The table-lamp collection and JSON below are fictional teaching examples, not a supported source or a live extraction. Use a website you are authorized to access and the connection details supplied by your validated build. Source access and machine access can vary; stop if either is unavailable.
1. Choose one collection#
Open the API generator with a category or results URL. For example, https://example.com/lighting/table-lamps/ describes a narrower collection than a store homepage.
State the records you need: “Product name, price, currency, availability, and product URL for table lamps in this category.” Check that inspection finds individual products and evidence for your requested fields. Navigation links, advertisements, login pages, and challenge pages are not product records.
See choosing a source if inspection finds a different collection or cannot access the page.
2. Review fields and validate the API#
Select the collection and review its observed fields, types, and missing values. A price without a currency may be unsuitable for comparison. A field found on one product page may be missing from other records.
Build and validate the selected API, then review its saved schema and sample records. A preview is not a working endpoint. Use fields and schemas to decide whether the output fits your application.
3. Copy the connection details#
Use the API origin, source identifier, definition version, credential, and request example from the saved API detail page. If machine access is unavailable, complete validation in the workspace before attempting these requests.
Set these variables in your server environment or secret manager; keep the credential out of browser code, source control, and shared logs:
| Environment variable | Value from your workspace |
|---|---|
CLEANEDWEB_API_BASE |
Supplied HTTPS API origin, without a path |
CLEANEDWEB_SOURCE_ID |
Saved source identifier |
CLEANEDWEB_DEFINITION_VERSION |
Saved 64-character lowercase definition version |
CLEANEDWEB_API_KEY |
Scoped API credential |
CLEANEDWEB_IDEMPOTENCY_KEY |
Unique key saved with this intended request |
Save the request key alongside the source, version, endpoint, and inputs before sending. Use a UUID or 1–200 letters, numbers, ., _, :, or -. Keep that value across a retry of the same request. A new intentional execution uses a new key. Read idempotency and recovery before automating requests.
The example contract below is the same one used in the API quickstart. Follow your generated request example if its route or inputs differ.
4. Make one bounded request with cURL#
This POST executes the source and may incur usage. For a collection API supporting these inputs, it requests at most one collection page and disables detail-page enrichment. Search APIs can require a constrained query instead; follow their generated request. Do not automatically repeat a failed or timed-out request with a new key.
umask 077
: "${CLEANEDWEB_IDEMPOTENCY_KEY:?Set the saved key for this request}"
curl --fail-with-body --max-time 60 \
"${CLEANEDWEB_API_BASE}/v1/sources/${CLEANEDWEB_SOURCE_ID}/run" \
-H "Authorization: Bearer ${CLEANEDWEB_API_KEY}" \
-H "Idempotency-Key: ${CLEANEDWEB_IDEMPOTENCY_KEY}" \
-H 'Content-Type: application/json' \
--data "{\"expected_version\":\"${CLEANEDWEB_DEFINITION_VERSION}\",\"input\":{\"max_pages\":1,\"detail_limit\":0}}" \
--output curl-response.json
Keep curl-response.json private if the source contains sensitive data. Check the HTTP outcome and the response summary before using the records. A timeout leaves the execution state uncertain; it does not establish that nothing ran.
5. Use website data in Python#
Download the Python client example. It uses Python 3.9 or later with no third-party packages. This example is for collection APIs that support max_pages and detail_limit. Set the same five environment variables, including the saved request key, then run:
umask 077
python3 website-api.py > python-response.json
Use a separate private directory for each intended run. If you already ran the cURL example with this key and the same body, the Python request recovers the same execution rather than collecting again.
The client sends one request with the saved idempotency key, limits, and version check. It refuses redirects, preserves the entire JSON response, and exits with code 2 when the summary needs review or either returned definition version differs from the requested version. Exit code 1 indicates a request, configuration, or response error. On HTTP errors (including 401, 429, and 504), python-response.json contains the HTTP status, allowlisted diagnostic body fields, and available content type, retry timing, and request/trace headers. Credentials are redacted; unrelated fields and headers are omitted. Non-JSON or oversized error bodies are marked without copying their contents. Keep this file private and share only the needed diagnostics with support. There are no automatic retries. Inspect the saved result and workspace before deciding what to do next.
6. Read records together with the summary#
This shortened response illustrates a bounded result:
{
"records": [{ "name": "Arc table lamp", "price": 48, "currency": "USD" }],
"summary": {
"status": "bounded",
"within_scope_passed": true,
"catalog_complete": false
}
}
The request passed within its selected scope. It did not collect the entire catalog. This shortened example omits version and delivery fields. In the full response, require summary.definition_version and metadata.definition_version to equal the version you requested. Check the schema, limits, missing fields, and summary before downstream use. A nonempty array or HTTP 200 is insufficient evidence of completeness.
If the summary reports partial, fails its scope check, or is absent, preserve the output for review. Do not silently replace that outcome with an empty dataset. See errors and retries.
7. Expand only after checking the first result#
Increase collection and detail limits only when the generated contract supports them and you understand the intended usage. max_pages is a collection budget, not a page number; repeating a request does not mean “fetch the next page.” Read pagination and output before expanding.
Keep credentials scoped, preserve missing values, and retain delivery.runId, delivery.chargeReceiptId, and traceId when returned. If the saved definition changes, review the new schema instead of removing expected_version. The source-change guide explains how to handle changed fields and partial runs. Use the Run API reference when building your integration.