{"id":985,"date":"2026-10-02T22:16:11","date_gmt":"2026-10-02T19:16:11","guid":{"rendered":"https:\/\/mexela.com\/blog\/scrapy-proxy-setup\/"},"modified":"2026-10-02T22:16:11","modified_gmt":"2026-10-02T19:16:11","slug":"scrapy-proxy-setup","status":"publish","type":"post","link":"https:\/\/mexela.com\/blog\/scrapy-proxy-setup\/","title":{"rendered":"Scrapy Proxy Setup: Authentication, Middleware, and Controlled Retries"},"content":{"rendered":"<p class=\"mexela-answer\">Configure a Scrapy proxy through the built-in <code>HttpProxyMiddleware<\/code>: either let the crawler process inherit supported proxy environment variables or set the <code>proxy<\/code> value in an individual request&#8217;s <code>meta<\/code> dictionary. Choose one owner, load authentication from protected configuration, set conservative concurrency and download timeouts, retry only transient failures with a finite budget, and verify the observed exit route before starting a crawl.<\/p>\n<p class=\"mexela-scope\"><strong>Scope:<\/strong> this guide covers Scrapy&#8217;s downloader proxy middleware, environment versus per-request precedence, credential handling, download-handler limits, retry classification, responsible crawl controls, and route evidence. It does not grant permission to crawl a site, bypass access controls, defeat anti-abuse systems, or guarantee support in every third-party download handler. The <a href=\"\/blog\/proxy-setup-developer-guides\/\">Proxy Setup and Developer Guides hub<\/a> links to adjacent client guides.<\/p>\n<h2 id=\"middleware\">Understand what HttpProxyMiddleware changes<\/h2>\n<p>The maintained <a href=\"https:\/\/docs.scrapy.org\/en\/latest\/topics\/downloader-middleware.html#module-scrapy.downloadermiddlewares.httpproxy\" rel=\"noopener\">Scrapy HttpProxyMiddleware documentation<\/a> says the middleware selects the HTTP proxy by setting the <code>proxy<\/code> metadata value on request objects. It supports environment-owned routes and an explicit per-request value. That makes the middleware a routing component in the downloader chain, not a general identity or policy bypass.<\/p>\n<p>The middleware is enabled by default when <code>HTTPPROXY_ENABLED<\/code> remains true. Before adding custom middleware, inspect effective project settings and the downloader handler. Reimplementing the built-in behavior can accidentally overwrite request metadata, expose secrets in logs, or run in the wrong middleware order. Keep a custom component only when the routing decision genuinely needs application logic and document its position in the chain.<\/p>\n<p>Write the intended route before changing settings: spider, downloader middleware, selected proxy, then authorized destination. Record Scrapy and Python versions, the download handler, a non-secret proxy label, authentication mode, location requirement, concurrency cap, timeout, and verification URL. A response status alone does not identify which network path produced it.<\/p>\n<h2 id=\"ownership\">Choose environment variables or request meta proxy ownership<\/h2>\n<p>Scrapy documents the lowercase <code>http_proxy<\/code>, <code>https_proxy<\/code>, and <code>no_proxy<\/code> environment variables for this middleware. Environment ownership is convenient when one container, worker, or scheduled job should use the same egress policy. It also means the route can differ between an interactive shell, a service manager, a queue worker, and a deployed container.<\/p>\n<p>A per-request <code>meta['proxy']<\/code> value takes precedence over the HTTP or HTTPS environment variables and ignores the environment no-proxy value. That precedence is powerful and easy to misuse. Adding a request meta proxy to one callback can silently remove a direct exception that operators believed still applied. Test the proxied and bypass paths independently whenever ownership changes.<\/p>\n<p>Prefer environment ownership when deployment controls one stable route for the entire crawler. Prefer per-request metadata when a small, explicit subset of authorized requests needs a distinct endpoint. Do not mix both merely as fallback layers. When a proxy is mandatory, missing or malformed configuration should stop startup rather than send traffic directly.<\/p>\n<h2 id=\"request-meta\">Set a per-request proxy without committing credentials<\/h2>\n<p>The request meta proxy value is a URI. Construct it from validated runtime configuration, keep the secret outside source control, and never log the complete value. The example uses symbolic placeholders read from an approved settings object; it cannot run until the project provides real values through its secret boundary.<\/p>\n<pre><code>import scrapy\n\nclass RouteCheckSpider(scrapy.Spider):\n    name = \"route_check\"\n    custom_settings = {\n        \"CONCURRENT_REQUESTS\": 2,\n        \"DOWNLOAD_TIMEOUT\": 15,\n        \"DOWNLOAD_VERIFY_CERTIFICATES\": True,\n        \"RETRY_ENABLED\": False,\n    }\n\n    async def start(self):\n        proxy_uri = self.settings.get(\"APP_PROXY_URI\")\n        if not proxy_uri:\n            raise RuntimeError(\"APP_PROXY_URI is required\")\n        yield scrapy.Request(\n            \"https:\/\/route-check.example.invalid\/json\",\n            meta={\"proxy\": proxy_uri},\n            callback=self.parse_route,\n            errback=self.record_failure,\n        )\n\n    def parse_route(self, response):\n        payload = response.json()\n        observed = payload.get(\"observedAddress\")\n        if not observed:\n            raise RuntimeError(\"Missing route evidence\")\n        expected_exit = self.settings.get(\"APP_EXPECTED_EXIT\")\n        if not expected_exit:\n            raise RuntimeError(\"APP_EXPECTED_EXIT is required\")\n        if observed != expected_exit:\n            raise RuntimeError(\"Observed route did not match the approved exit\")\n        yield {\"route_ok\": True, \"status\": response.status}\n\n    def record_failure(self, failure):\n        self.logger.error(\"Route check failed: %s\", failure.type.__name__)<\/code><\/pre>\n<p>Set <code>APP_EXPECTED_EXIT<\/code> to the approved exit reported by your controlled acceptance check. The example emits a passing route item only when the returned address matches that value. Its errback logs the exception class without exporting a URI, password, or raw failure message. Keep detailed redacted diagnostics in the protected operational record.<\/p>\n<p>Keep <code>APP_PROXY_URI<\/code> out of Scrapy settings dumps, job arguments, request fingerprints, exported items, and support bundles. If the provider uses username and password authentication, assemble the runtime URI inside the narrowest configuration boundary and redact URI user information from exception messages. For source-IP authentication, verify the crawler host&#8217;s actual outbound address before allowlisting it.<\/p>\n<p>The <a href=\"\/blog\/python-requests-proxy\/\">Python Requests proxy guide<\/a> is useful for a single-request control outside Scrapy. A passing Requests check proves the gateway is reachable from that process, but Scrapy still needs its own handler and middleware verification.<\/p>\n<h2 id=\"authentication\">Separate proxy authentication from destination authentication<\/h2>\n<p>Proxy authentication controls access to the gateway. Destination authentication controls access to the site or API. Do not reuse credentials between those boundaries. A proxy URI may support user information, but embedding a password makes accidental disclosure more likely through project settings, debug logs, exceptions, shell history, process inspection, or monitoring exports.<\/p>\n<p>An HTTP 407 normally indicates a proxy authentication challenge or rejection. Destination 401 responses concern destination credentials. A 403 may reflect destination authorization, policy, or crawl controls; a 429 usually indicates pacing. Preserve the response source and request context before changing credentials. The <a href=\"\/blog\/proxy-authentication-username-password-vs-ip-auth\/\">proxy authentication guide<\/a> compares credential and source-IP access models.<\/p>\n<p>Scrapy exposes an authentication encoding setting for the proxy middleware, with a documented default of Latin-1. Do not change it speculatively. If a username cannot be represented or authentication fails only for particular characters, confirm the provider&#8217;s required encoding in a controlled test. Secret rotation should refresh worker processes and repeat one route check before resuming the crawl.<\/p>\n<h2 id=\"handlers\">Check download-handler protocol limits<\/h2>\n<p>Support for request proxy metadata ultimately belongs to the active download handler. Scrapy&#8217;s documentation warns that not every third-party handler implements it and specifically notes that the H2 handler does not currently support the meta key. A project can therefore have correct middleware settings and still lack proxy behavior in the selected transport.<\/p>\n<p>Proxy URI scheme and destination scheme are separate. An HTTP proxy can carry an HTTPS destination through tunneling when the handler supports the combination. The current Scrapy documentation also describes narrower HTTPS-proxy support in the HTTP\/1.1 handler and SOCKS support in a particular httpx handler rather than every built-in handler. Pin the handler and test the exact protocol pair used in production.<\/p>\n<p>For the current Scrapy handler, explicitly set <code>DOWNLOAD_VERIFY_CERTIFICATES<\/code> to <code>True<\/code>; the documented default does not enforce certificate verification. The <a href=\"https:\/\/docs.scrapy.org\/en\/latest\/topics\/settings.html#download-verify-certificates\" rel=\"noopener\">Scrapy certificate verification setting<\/a> depends on download-handler support, so verify the exact handler and version. The acceptance-only example disables automatic retries to retain the first failure before classification.<\/p>\n<p>Keep TLS verification enabled. A certificate error is evidence about destination trust, hostname validation, a tunnel, or an approved inspection layer. Disabling verification hides that boundary and can expose crawl credentials and response data. Record the Python, OpenSSL, Twisted, Scrapy, and handler versions for a reproducible support case.<\/p>\n<h2 id=\"policy\">Set conservative crawl permissions and concurrency<\/h2>\n<p>A proxy does not grant permission to crawl. Review the destination&#8217;s terms, documented API rules, applicable law, account agreement, and robots policy before sending requests. Use the site&#8217;s API or data export when available. Identify the crawler honestly when required, honor access restrictions, and provide a contact route for an authorized operational crawl.<\/p>\n<p>Start with one request at a time. Increase concurrency only after observing destination capacity, response quality, latency, and published limits. Configure per-domain concurrency and download delay where appropriate so a large global worker pool cannot overwhelm one host. Auto-throttling may help adapt load, but it does not replace a hard cap or permission boundary.<\/p>\n<p>A stable proxy can make sessions and source identity more predictable, but it must not be used to evade a block or multiply traffic beyond the intended rate. When a destination rejects the crawler, pause and investigate. Rotating endpoints after every rejection can amplify harm and destroy the evidence needed to distinguish authentication, policy, and transient network failures.<\/p>\n<h2 id=\"retries\">Use controlled retries by failure class<\/h2>\n<p>Controlled retries require a finite attempt count, an overall time budget, backoff, jitter, and a list of eligible failures. Retry a read-only request after a brief connection reset, selected gateway timeout, or temporary service response only when the destination contract permits it. Retain the first failure and record the final reason.<\/p>\n<p>Do not retry proxy 407 responses, certificate failures, malformed proxy configuration, destination 401 or 403 responses, robots denials, or parsing defects as if they were transient. Treat 429 according to the destination&#8217;s retry guidance and slow down the whole relevant concurrency group. A retry should not switch identity or route automatically unless that behavior is explicitly authorized and tested.<\/p>\n<p>Scrapy&#8217;s retry middleware can reschedule eligible requests, so calculate worst-case traffic as original requests multiplied by attempts. Combine that with redirect limits and any spider-level rescheduling. Use idempotent GET or HEAD checks during acceptance; do not repeat state-changing requests without a documented idempotency mechanism.<\/p>\n<h2 id=\"verification\">Run a bounded route verification crawl<\/h2>\n<ol>\n<li>Record Python, Scrapy, Twisted, TLS library, download handler, and effective settings without secrets.<\/li>\n<li>Choose an endpoint you own or are authorized to use that reports the source address it observes.<\/li>\n<li>When policy permits, run one direct request and retain the observed source and duration.<\/li>\n<li>Enable exactly one proxy owner: environment variables or request metadata.<\/li>\n<li>Run one proxied request with low concurrency, a finite timeout, and no automatic route rotation.<\/li>\n<li>Confirm the observed source matches the expected assignment and destination TLS remains valid.<\/li>\n<li>Repeat once from the production worker context.<\/li>\n<li>Test one approved destination page at a conservative rate.<\/li>\n<\/ol>\n<p class=\"mexela-expected\"><strong>Expected observation:<\/strong> HttpProxyMiddleware applies the selected route, the active handler connects within the timeout, the controlled endpoint reports the assigned proxy exit instead of the direct control, and a second request from the real worker uses the same ownership rule without exposing credentials.<\/p>\n<p>Store a small structured record: timestamp, spider and release version, redacted route label, destination category, status class, elapsed time, retry count, and expected-versus-observed route result. Do not store the complete proxy URI, authorization header, cookies, response body, or customer data.<\/p>\n<p>If the project also sends requests through PHP, compare the <a href=\"\/blog\/php-guzzle-proxy\/\">Guzzle proxy configuration<\/a> separately. Its request options and timeout settings do not replace Scrapy middleware ownership or crawler retry policy.<\/p>\n<h2 id=\"troubleshooting\">Diagnose the crawler pipeline by boundary<\/h2>\n<table>\n<thead>\n<tr>\n<th>Observation<\/th>\n<th>Likely boundary<\/th>\n<th>First check<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Proxy host cannot resolve<\/td>\n<td>Configuration or DNS<\/td>\n<td>Validate the redacted host and worker resolver<\/td>\n<\/tr>\n<tr>\n<td>Connection refused or timed out<\/td>\n<td>Gateway reachability<\/td>\n<td>Check scheme, port, firewall, and download timeout<\/td>\n<\/tr>\n<tr>\n<td>HTTP 407<\/td>\n<td>Proxy authentication<\/td>\n<td>Check gateway account and secret rotation<\/td>\n<\/tr>\n<tr>\n<td>Unexpected direct route<\/td>\n<td>Ownership or bypass<\/td>\n<td>Inspect request meta, environment, and no-proxy intent<\/td>\n<\/tr>\n<tr>\n<td>Works in Requests, fails in Scrapy<\/td>\n<td>Handler or middleware<\/td>\n<td>Confirm effective handler and proxy-meta support<\/td>\n<\/tr>\n<tr>\n<td>HTTP 403 or 429<\/td>\n<td>Destination policy<\/td>\n<td>Pause, review permission and pacing, then reduce load<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Use an errback to collect exception class, failure stage, request label, retry count, and elapsed time. Sanitize the message before export because it may contain a proxy or destination URI. Compare interactive and worker environments, middleware order, handler selection, DNS, CA bundle, and injected settings. Change one variable at a time.<\/p>\n<p class=\"mexela-limits\"><strong>Operational limits:<\/strong> one successful Scrapy request proves one process, middleware chain, handler, proxy route, destination, and moment. It does not prove that every spider uses the same metadata, that another worker inherited the same environment, or that a destination permits higher volume. A proxy changes routing; it does not create crawl permission or remove destination rate limits.<\/p>\n<h2 id=\"next-step\">Choose capacity after the controlled crawl passes<\/h2>\n<p>Document the proxy protocol, authentication method, location, session stability, concurrency cap, transfer estimate, handler, bypass policy, retry budget, and acceptance URL. Compare those measured requirements with current <a href=\"\/proxy-pricing\/\">Mexela proxy pricing<\/a> and confirm inventory before increasing crawler capacity.<\/p>\n<h2 id=\"faq\">Frequently asked questions<\/h2>\n<div class=\"mexela-faq\">\n<h3>Is HttpProxyMiddleware enabled by default?<\/h3>\n<p>Its enable setting defaults to true, but effective behavior still depends on project settings, middleware configuration, and the active download handler.<\/p>\n<h3>Does request meta proxy override environment variables?<\/h3>\n<p>Yes. Scrapy documents that the per-request value takes precedence over HTTP and HTTPS proxy variables and ignores the environment no-proxy value.<\/p>\n<h3>Can every Scrapy download handler use the proxy meta key?<\/h3>\n<p>No. The handler must implement support. Verify the exact handler and protocol combination used by the project.<\/p>\n<h3>Should a spider retry HTTP 407?<\/h3>\n<p>No. Treat it as a proxy-authentication failure, fix the credential boundary, and rerun one controlled request.<\/p>\n<h3>How many concurrent requests should I start with?<\/h3>\n<p>Begin with one or a very small number, observe the destination&#8217;s rules and behavior, then increase only within an approved cap.<\/p>\n<h3>Does using a proxy make crawling permitted?<\/h3>\n<p>No. Permission, terms, robots policy, API rules, and applicable law remain separate from the network route.<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Configure Scrapy proxy routing through HttpProxyMiddleware, protect credentials, respect bypass intent, classify failures, and verify the route safely.<\/p>\n","protected":false},"author":0,"featured_media":986,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[189],"tags":[],"_links":{"self":[{"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/posts\/985"}],"collection":[{"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/comments?post=985"}],"version-history":[{"count":0,"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/posts\/985\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/media\/986"}],"wp:attachment":[{"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/media?parent=985"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/categories?post=985"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/mexela.com\/blog\/wp-json\/wp\/v2\/tags?post=985"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}