CleanedWeb
3 min read
View Markdown ↗

Core concepts

A CleanedWeb integration connects a website collection to an application through a reviewed, versioned API contract. These are the objects you work with along the way.

On this page

Sources and collections#

A source identifies the website data your API collects. Its collection describes the set of records you want: products in a category, jobs matching a search, or articles in a section. One website can contain many collections with different fields and access requirements.

Start with a specific public URL and a sentence describing the scope. A category filter, location, or date range is part of that scope. Choosing a source does not establish that every page on the website is accessible or supported.

Records, fields, and schemas#

A record represents one item. Fields describe it: a product name, price, currency, or source URL. The output schema defines the fields and types your accepted API returns. It is specific to that API.

Keep these distinctions when consuming data:

Value Meaning for your application
A populated field A value was returned; validate its type and meaning.
A missing or null field The value is unavailable under the schema; do not manufacture a default fact.
A detail field Its availability may depend on optional detail-page collection.
An example record A sample for review; it does not establish current source content or full coverage.

Review fields and schemas before making a field mandatory downstream.

Builds and definition versions#

Inspection discovers possible collections and fields. A build validates a selected collection and produces its saved contract. A preview, an unfinished build, and a validated API are distinct stages.

A definition version identifies immutable collection rules and their schema. Keep that version with your integration configuration. expected_version checks that the active definition still matches the one you reviewed. A version identifies the rules, not a frozen copy of the website: source content can change between runs.

See build and validate and versions and changes.

Requests, runs, and results#

A request asks the API to execute with specific inputs. A run is that execution. Its result contains records and an execution summary; customer deliveries also carry identifiers you can use to inspect the saved result and charging receipt.

An idempotency key identifies one intended execution. Keep it with the request before sending. Replaying a settled request with the same identity returns the saved result; a new deliberate collection uses a new key. See idempotency and recovery.

Bounded and partial output#

A request can collect only a selected part of a website. max_pages, where supported, is a collection budget rather than a universal page index. detail_limit bounds detail enrichment.

A bounded result can pass its selected scope without collecting the entire catalog. A partial result signals that some expected work did not complete. Always evaluate summary.status, summary.within_scope_passed, the requested limits, and catalog_complete with the records.

Continue with the quickstart for a first request or the run reference for the shared contract.

CleanedWeb documentationGet help with this guide ↗
Search across the documentation · Esc to close