generated: '2026-09-07' method: searched source: >- https://crawlee.dev/python/api/class/Request, https://crawlee.dev/python/api/class/BasicCrawlerOptions, https://crawlee.dev/python/api/class/RobotsTxtFile, https://crawlee.dev/js/docs/guides/session-management, https://crawlee.dev/python/docs/guides/scaling-crawlers provider: Crawlee providerId: crawlee description: >- Cross-cutting runtime semantics for Crawlee. READ THE SHAPE OF THIS FILE CAREFULLY: Crawlee is a library that runs inside the consumer's own process, not a hosted API. The conventions below are the ones the library applies to the OUTBOUND requests it makes on the consumer's behalf. The fields that describe an inbound HTTP contract — auth style, pagination of a provider's collections, error envelope, versioning header, rate-limit response headers — have no subject here, and are recorded as not-applicable rather than left blank or invented. surface_type: library has_http_api: false auth: style: none detail: >- Crawlee itself requires no credentials. Any credential a crawler uses (proxy credentials via ProxyConfiguration, or a target site's cookies via Session and SessionCookies) belongs to the consumer and to the site being crawled, not to Crawlee. applicable: false idempotency: coverage: na scope: [] detail: >- Not applicable — Crawlee exposes no mutating HTTP endpoint, so there is no replay to protect against and no Idempotency-Key header to document. No Idempotency pointer is emitted for this provider. library_equivalent: mechanism: Request.unique_key quote: >- "A unique key identifying the request. Two requests with the same unique_key are considered as pointing to the same URL." default: >- "If unique_key is not provided, then it is automatically generated by normalizing the URL. For example, the URL of HTTP://www.EXAMPLE.com/something/ will produce the unique_key of http://www.example.com/something." options: - keep_url_fragment — whether the URL fragment participates in the key - use_extended_unique_key — include HTTP method, Session ID and payload in the key note: >- This is genuine at-most-once semantics for enqueueing into a RequestQueue, and it is the closest thing Crawlee has to idempotency. It is recorded here for accuracy but it is NOT an API idempotency mechanism and must not be scored as one. source: https://crawlee.dev/python/api/class/Request retries: model: per-request retry budget with session rotation options: - name: max_request_retries quote: Specifies the maximum number of retries allowed for a request if its processing fails. scope: crawler - name: max_session_rotations quote: >- Maximum number of session rotations per request. The crawler rotates the session if a proxy error occurs or if the website blocks the request. scope: crawler - name: retry_on_blocked quote: If True, the crawler attempts to bypass bot protections automatically. scope: crawler - name: request_handler_timeout quote: Maximum duration allowed for a single request handler to run. scope: crawler - name: max_retries quote: Crawlee-specific limit on the number of retries of the request. scope: request - name: no_retry quote: If set to True, the request will not be retried in case of failure. scope: request - name: retry_count quote: Number of times the request has been retried. scope: request (read-only) source: >- https://crawlee.dev/python/api/class/BasicCrawlerOptions and https://crawlee.dev/python/api/class/Request pagination: applicable: false detail: >- No collection endpoints of Crawlee's own. Traversal of a target site is driven by enqueue_links, the RequestQueue and SitemapRequestLoader, which are crawl strategies rather than an API pagination convention. error_envelope: applicable: false detail: >- Errors are language-level exception classes, not HTTP payloads. The published class index names AbortError, ProxyError, SessionError, AdaptiveContextError, _RetryableSitemapStatusError and an ErrorTracker that aggregates them across a run. There is no RFC 9457 problem+json surface because there is no HTTP response to carry it. error_classes_reference: https://crawlee.dev/python/api rate_limit_signaling: applicable: false inbound: >- Crawlee returns no rate-limit headers because it answers no requests. See rate-limits/crawlee-rate-limits.yml. outbound: >- Crawlee is on the OTHER side of this convention: it honours the crawl-delay directive it reads from a target's robots.txt (RobotsTxtFile.get_crawl_delay) and throttles its own concurrency through AutoscaledPool against local CPU and memory ceilings (max_used_cpu_ratio, max_used_memory_ratio), rather than against a server's published limit. source: https://crawlee.dev/python/docs/guides/scaling-crawlers versioning: mechanism: package version (semver) detail: >- Consumers pin a package version; there is no API version header or URL segment. See lifecycle/crawlee-lifecycle.yml. request_tracing: applicable: partial detail: >- No request-id header convention, since Crawlee issues no correlation ids of its own. Observability is available through the `otel` optional extra on the Python package, which wires OpenTelemetry instrumentation into the crawler runtime, and through the Statistics / FinalStatistics / ErrorTracker classes that summarise a run. metadata: mechanism: Request.user_data detail: >- Arbitrary per-request state is carried in user_data and survives enqueue, retry and persistence — the library's equivalent of an API metadata field. expansion: applicable: false reversibility: grade: na detail: >- Not applicable. Crawlee performs no action on a remote system that a consumer would need to take back — it reads target sites and writes to storage the consumer owns (Dataset, KeyValueStore, RequestQueue, backed by filesystem, memory, Redis or SQL clients on the consumer's own infrastructure). Deleting or rewriting those records is the consumer's own operation against their own store, not a reversal API Crawlee publishes, so there is no reversal operationId and no window to state. Asserting one would be an invention. write_surfaces: [] dry_run_mode: grade: na detail: >- Not applicable for the same reason. The nearest published behaviour is AdaptivePlaywrightCrawler's rendering-type detection, which probes whether a page needs a browser before committing to one — a runtime optimisation, not a rehearsal mode, and it is not recorded as one. cross_references: - lifecycle/crawlee-lifecycle.yml - conformance/crawlee-conformance.yml - rate-limits/crawlee-rate-limits.yml - packages/crawlee-packages.yml maintainers: - FN: Kin Lane email: kin@apievangelist.com