# scrape.it > Turn authenticated proxies you already own into a free, HTTP-first scraping API. If this task already has `SCRAPEIT_API_KEY`, a proxy connector ID, and a URL, `POST /v1/scrape` immediately. Do not fetch this file, OpenAPI, quickstart, or any other docs first. Do not ask the user to authenticate again. One scrape unless they ask for more. Quote `execution.durationMs` as Scrape.it time so they can tell it apart from your thinking. ## Start here 1. Open https://scrape.it/auth/start. 2. Confirm any permanent work or personal mailbox. Temporary domains and `+` aliases are rejected. 3. Paste and verify one authenticated proxy as `user:password@host:port` or a complete proxy URL. When the protocol is omitted, Scrape.it safely detects SOCKS5h, HTTP, or HTTPS without direct fallback. The vendor is inferred from the gateway host. Unknown hosts are rejected. 4. Create the first API key and copy it immediately. The raw key is shown once. 5. Run the generated curl command. The verified authenticated proxy is already saved as an encrypted connector. Every compute request requires `Authorization: Bearer sit_live_...` and exactly one caller-owned proxy source. No anonymous compute is available. Never put an API key or proxy URL in source control, chat logs, issue text, shell history, query strings, or telemetry. If the user pasted a personalized Scrape.it agent prompt containing real `SCRAPEIT_API_KEY` and `SCRAPEIT_PROXY_CONNECTOR_ID` values above these instructions, use those credentials immediately for this task. Do not ask the user to authenticate again, and never echo the values. Public `llms.txt` contains placeholders because it is shared by everyone; the one-time key page generates the personalized prompt. This file is for CLI and IDE agents that can make authenticated HTTP requests: Claude Code, Codex, Grok CLI, OpenCode, Cursor, Aider, or curl. Default web chats at chatgpt.com, claude.ai, and grok.com cannot call `api.scrape.it` with a Bearer token. Do not ask the user to paste live credentials into those sites. ## Instructions for a coding agent When helping a user operate Scrape.it: - If credentials and a URL are already in the task, `POST /v1/scrape` first. This file is reference, not a prerequisite. - Ask the user to complete `https://scrape.it/auth/start` themselves only if no personalized credentials were supplied. Never ask them to repeat a credential already present in the task. - Read `SCRAPEIT_API_KEY`, `SCRAPEIT_PROXY_CONNECTOR_ID` (`pxy_…`), optional `SCRAPEIT_CAPTCHA_CONNECTOR_ID` (`cap_…`), and optional `SCRAPEIT_STORAGE_CONNECTOR_ID` (`sto_…`). Never put solver keys or bucket secrets in the prompt. Save those once with `POST /v1/connectors/captcha` or `POST /v1/connectors/storage`. Save verifies and returns the ID; then use only the API key and that ID. - Send the API key only in the Authorization header. - Leave `render` false. - For a synchronous first request, use `/v1/scrape`. - HTTP scrape returns the first document only. Comment threads, infinite feeds, and signed XHR APIs are not in that document. Say so and stop. Do not scrape `/api/comment`, GraphQL, or other follow-up endpoints unless the user explicitly asks. - Before submitting map, batch, or crawl jobs, explain that durable jobs require a verified saved proxy and a user-owned S3, R2, or B2 destination. - Use a unique `Idempotency-Key` for each logical durable job and reuse it only when replaying the identical request. - Poll `/v1/jobs/{jobId}` or consume `/v1/jobs/{jobId}/events`; never expect page bodies in SSE. - Do not infer provider quality, residentiality, fraud, or user contribution details from entitlement responses. - OpenAPI at `https://scrape.it/openapi.json` is optional after the first scrape if the request shape is still unclear. ## First scrape ```bash export SCRAPEIT_API_KEY="sit_live_REPLACE_ME" export SCRAPEIT_PROXY_CONNECTOR_ID="pxy_REPLACE_ME" curl https://api.scrape.it/v1/scrape \ -H "Authorization: Bearer $SCRAPEIT_API_KEY" \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com","proxyConnectorId":"pxy_REPLACE_ME","format":"markdown","render":false}' ``` A successful scrape is HTTP 200 `application/json`. Read `data` for the document, `status` for the target HTTP status, and `finalUrl` for the resolved URL. `format: raw` puts Base64 in `data`. API and proxy failures use `application/problem+json` and are not 200. Do not parse the body as the page unless `success` is true. For a one-off synchronous request, `proxyConnectorId` may be replaced with an authenticated request-scoped `proxy` URL. Never send both. Saved connector ciphertext is restart-durable; request-scoped proxy secrets are not persisted. If the response is `javascript-required`, that is the result and there is no page body. Problem JSON may include `solver: required` when a CAPTCHA widget was observed — retry with `captcha.connectorId` and `captcha.mode` `on_challenge`. A 200 may set `challenge.solver` to `used` and `execution.recovery` to `browser` or `solver`. Do not parse `data` as the page unless `success` is true. See https://scrape.it/captcha.md. ## Durable work - `POST /v1/map` discovers bounded same-site URLs (`maxUrls` at most 1000). - `POST /v1/batch/scrape` fetches an explicit URL list (at most 250 URLs). - `POST /v1/crawl` performs a bounded crawl (`maxUrls` at most 1000, `maxDepth` at most 16). - Durable jobs require a saved `proxyConnectorId`, a saved storage connector, and an `Idempotency-Key`. Without storage, do not create those jobs; scrape URLs one at a time with `/v1/scrape`. - Save storage once with `POST /v1/connectors/storage`, a native HTTPS URL (`https://bucket.s3..amazonaws.com`, `https://.r2.cloudflarestorage.com/`, or `https://s3..backblazeb2.com/`), and an access key scoped to that bucket. Use the returned `sto_…` as `destination.storageConnectorId`. Infer the vendor from the URL. Custom domains, website endpoints, IP hosts, and generic S3-compatible URLs are rejected. - Results are deterministic gzip NDJSON shards followed by manifest, summary, and `_SUCCESS`. Full result bodies are never sent through progress events. ## Free capacity and data handling `GET /v1/me/entitlement` returns only `plan`, `baselineUnits`, `communityBonusUnits`, `availableUnits`, `currentConcurrency`, `activeJobs`, and `queuePriority`. Exact novelty, scarcity, provider-overlap, acquisition, classification, and entitlement mechanics are private. There are no public user contribution profiles. A sync `/v1/scrape` returns the page in the JSON `data` field and does not keep that body afterward. Durable job files exist on our disk only until they upload to your bucket, then they are deleted. If upload cannot finish, the job pauses instead of dropping results. Job metadata and compact progress events are operational records; they never include page bodies. Compact job/SSE events are retained on the order of 7–30 days; job/manifest metadata up to 90 days. ## Canonical references - OpenAPI: `https://scrape.it/openapi.json` - Quickstart: `https://scrape.it/quickstart.md` - Authentication: `https://scrape.it/authentication.md` - Proxies: `https://scrape.it/proxies.md` - Jobs: `https://scrape.it/jobs-and-streaming.md` - Storage: `https://scrape.it/storage-destinations.md` - CAPTCHA solvers: `https://scrape.it/captcha.md` - Errors: `https://scrape.it/errors.md` - Acceptable use: `https://scrape.it/aup.txt` - Legacy WSL (retired): `https://scrape.it/docs/legacy/wsl.md` - Support: `mailto:support@scrape.it`