# Schedule API runs

Run recurring collection from a scheduler you operate, such as a server cron job, workflow runner, or task queue. This guide describes a customer-managed integration; it does not rely on a native CleanedWeb scheduling endpoint.

Complete the [quickstart](/docs/quickstart/) and verify your storage path before enabling a schedule. Decide how fresh the data must be, how much collection each run may perform, and who investigates an incomplete or uncertain outcome.

## Save a bounded job configuration

Keep the API origin, source identifier, reviewed definition version, and accepted inputs in your job configuration. Load the credential from your server's secret manager at execution time. Check its scope and expiry, and rotate it before the schedule loses access.

Define operational limits in your scheduler:

| Setting | What to decide |
| --- | --- |
| Frequency | The observation interval your application actually needs |
| Time zone | An explicit zone, including behavior around daylight saving changes |
| Collection budget | Supported page and detail limits for one execution |
| Maximum runtime | When the worker stops waiting and records an uncertain outcome |
| Concurrency | How overlapping attempts for the same source are prevented |
| Recovery owner | Who reviews unresolved execution or ingestion failures |

Use [pagination and output](/docs/pagination/) to set the collection budget. `max_pages` limits work within an execution; repeated requests do not mean “continue from the next page.” Avoid accumulating missed intervals into an unbounded catch-up batch.

## Execute one attempt at a time

Use a durable job record and a lock or queue partition for each source. Begin conservatively with one active execution per source, and account for any other consumers sharing it. Do not assume a universal concurrency allowance; follow the limits available for the API and workspace.

Before dispatch, save an application attempt ID, scheduled time, definition version, intended input, and an `Idempotency-Key`. Customer-workspace runs require this header. Persist one key for each intended execution; use a new key for the next scheduled observation. Issue the request once, save the response and returned `delivery.runId` when present, then check its summary before publishing records downstream.

Prevent the next scheduled interval from starting overlapping work when the current attempt is still running or unresolved. If a worker exits or its lock expires, reconcile the prior execution before dispatching its replacement.

## Separate retries from reconciliation

Disable blind POST retries in the HTTP client, job runner, and infrastructure. A client timeout or lost response does not prove that the source execution stopped or that no usage occurred.

Record the outcome as unknown and inspect available workspace run evidence. Retain the same idempotency key and identical request when reconciling that execution; a settled replay uses the existing result. Do not generate a replacement key to get around an unresolved attempt. Read [idempotency](/docs/idempotency/) for replay behavior and [errors and retries](/docs/errors/) before deciding whether to resubmit. Honor returned retry timing for rejected requests.

When collection succeeded but ingestion failed, replay the saved response into your storage path. See [storing results](/docs/storing-results/) for deduplication and publication checks.

## Review the schedule in operation

Monitor the last accepted observation, partial results, authentication failures, version conflicts, and unresolved attempts. An active scheduler or recent heartbeat does not establish that usable records were delivered.

Pause affected jobs when a definition change requires review. Validate the replacement contract and a bounded response before resuming the normal interval. Use [monitoring](/docs/monitoring/), [definition versions](/docs/versioning/), and the [production checklist](/docs/production-checklist/) to keep that process repeatable.
