Store API results
Store records together with the evidence needed to interpret them. A records array alone does not tell you which definition produced it, what scope was requested, or whether the execution completed that scope.
On this page
Start with a bounded request from the quickstart. Preserve the full response privately before transforming it into your application's tables or documents.
Keep a run envelope#
Use a separate run record to describe each attempted execution. The following fields are a suggested application storage design, not a CleanedWeb response schema:
| Stored value | Purpose |
|---|---|
| Application attempt ID and idempotency key | Correlate the request, replay, and ingestion logs |
| Source ID and definition version | Identify the reviewed API contract |
| Execution endpoint, requested input, and limits | Preserve request identity and the intended collection scope |
| Start and observation times | Distinguish source observations from ingestion time |
| Run ID, charge receipt ID, and trace ID | Locate execution, delivery, and usage evidence when available |
| HTTP outcome and execution summary | Preserve partial, failed, or uncertain outcomes |
| Raw response location | Revisit original values and diagnostics |
| Ingestion state | Record whether downstream publication passed your checks |
Do not put API keys, authorization headers, or secret-bearing URLs in this envelope. Apply access controls and retention appropriate to the data you collect.
Validate before publishing#
Check the returned definition version, expected response shape, record count, summary, and required fields. Validate against the saved schema rather than assumptions copied from an unrelated source.
The run API returns definition identity in summary.definition_version and metadata.definition_version. Retain delivery.runId, delivery.chargeReceiptId, and delivery.recordCount when supplied. See idempotency before repeating an execution request.
Keep raw values and record any transformations separately. Preserve missing values according to the contract; a missing price is not zero, and a missing availability value is not evidence that an item is unavailable.
Stage the response before updating customer-facing tables. Commit the run's accepted ingestion state only after the associated record writes succeed. If ingestion fails after a response was stored, recover from that saved response instead of executing the source again just to repeat a database write.
Choose a stable deduplication key#
Use the source's stable identifier where available, combined with your source ID. If you use a record URL, verify its identity behavior before relying on it: URL parameters, redirects, and changed slugs can affect uniqueness.
Avoid array position or a mutable title as the sole identity. Deduplicate within a response and across repeat ingestion of the same response. Keep observation history when changes over time matter; an upsert into the latest-value table alone cannot reconstruct previous values.
Preserve incomplete observations#
An item missing from a bounded or partial result has not necessarily been deleted upstream. Never use such a result to remove all database rows absent from the latest response. Establish complete coverage of the same collection before applying any deletion policy, and document how your application confirms disappearance.
Check pagination and output for the difference between meeting a requested bound and collecting the full catalog. Hold partial, failed, or unverified outcomes for review without overwriting the last accepted dataset with an empty replacement.
If the request times out, preserve its attempt metadata and reconcile the execution before another POST. Follow errors and retries and source changes. Use monitoring to track both collection outcomes and downstream ingestion failures.