Scrapy 动态代理IP Scrapy 配置动态代理IP的实现
BradyCC 人气:0想了解Scrapy 配置动态代理IP的实现的相关内容吗,BradyCC在本文为您仔细讲解Scrapy 动态代理IP的相关知识和一些Code实例,欢迎阅读和指正,我们先划重点:Scrapy,动态代理IP,Scrapy,,代理IP,下面大家一起来学习吧。
应用 Scrapy框架 ,配置动态IP处理反爬。
# settings 配置中间件 DOWNLOADER_MIDDLEWARES = { 'text.middlewares.TextDownloaderMiddleware': 543, # 'text.middlewares.RandomUserAgentMiddleware': 544, # 'text.middlewares.CheckUserAgentMiddleware': 545, 'text.middlewares.ProxyMiddleware': 546, 'text.middlewares.CheckProxyMiddleware': 547 } # settings 配置可用动态IP PROXIES = [ "http://101.231.104.82:80", "http://39.137.69.6:8080", "http://39.137.69.10:8080", "http://39.137.69.7:80", "http://39.137.77.66:8080", "http://117.191.11.102:80", "http://117.191.11.113:8080", "http://117.191.11.113:80", "http://120.210.219.103:8080", "http://120.210.219.104:80", "http://120.210.219.102:80", "http://119.41.236.180:8010", "http://117.191.11.80:8080" ]
# middlewares 配置中间件 import random class ProxyMiddleware(object): def process_request(self, request, spider): ip = random.choice(spider.settings.get('PROXIES')) print('测试IP:', ip) request.meta['proxy'] = ip class CheckProxyMiddleware(object): def process_response(self, request, response, spider): print('代理IP:', request.meta['proxy']) return response
加载全部内容