Python—requests模塊詳解

昵稱QAb6ICvc 2022-04-25

展開全文

1、模塊說明

requests是使用Apache2 licensed 許可證的HTTP庫。

用python編寫。

比urllib2模塊更簡潔。

Request支持HTTP連接保持和連接池，支持使用cookie保持會話，支持文件上傳，支持自動響應(yīng)內(nèi)容的編碼，支持國際化的URL和POST數(shù)據(jù)自動編碼。

在python內(nèi)置模塊的基礎(chǔ)上進(jìn)行了高度的封裝，從而使得python進(jìn)行網(wǎng)絡(luò)請求時，變得人性化，使用Requests可以輕而易舉的完成瀏覽器可有的任何操作。

現(xiàn)代，國際化，友好。

requests會自動實現(xiàn)持久連接keep-alive

2、基礎(chǔ)入門

1）導(dǎo)入模塊

import requests

2）發(fā)送請求的簡潔

　　示例代碼：獲取一個網(wǎng)頁（個人github）

import requests

r = requests.get('https://github.com/Ranxf')       # 最基本的不帶參數(shù)的get請求r1 = requests.get(url='http://dict.baidu.com/s', params={'wd': 'python'})      # 帶參數(shù)的get請求

我們就可以使用該方式使用以下各種方法

1   requests.get('https://github.com/timeline.json’)                                # GET請求2   requests.post(“http:///post”)                                        # POST請求3   requests.put(“http:///put”)                                          # PUT請求4   requests.delete(“http:///delete”)                                    # DELETE請求5   requests.head(“http:///get”)                                         # HEAD請求6   requests.options(“http:///get” )                                     # OPTIONS請求

3）為url傳遞參數(shù)

>>> url_params = {'key':'value'}       #    字典傳遞參數(shù)，如果值為None的鍵不會被添加到url中>>> r = requests.get('your url',params = url_params)>>> print(r.url)
　　your url?key=value

4）響應(yīng)的內(nèi)容

r.encoding                       #獲取當(dāng)前的編碼r.encoding = 'utf-8'             #設(shè)置編碼r.text                           #以encoding解析返回內(nèi)容。字符串方式的響應(yīng)體，會自動根據(jù)響應(yīng)頭部的字符編碼進(jìn)行解碼。r.content                        #以字節(jié)形式（二進(jìn)制）返回。字節(jié)方式的響應(yīng)體，會自動為你解碼 gzip 和 deflate 壓縮。r.headers                        #以字典對象存儲服務(wù)器響應(yīng)頭，但是這個字典比較特殊，字典鍵不區(qū)分大小寫，若鍵不存在則返回Noner.status_code                     #響應(yīng)狀態(tài)碼r.raw                             #返回原始響應(yīng)體，也就是 urllib 的 response 對象，使用 r.raw.read()   r.ok                              # 查看r.ok的布爾值便可以知道是否登陸成功
 #*特殊方法*#r.json()                         #Requests中內(nèi)置的JSON解碼器，以json形式返回,前提返回的內(nèi)容確保是json格式的，不然解析出錯會拋異常r.raise_for_status()             #失敗請求(非200響應(yīng))拋出異常

post發(fā)送json請求：

1 import requests2 import json3  
4 r = requests.post('https://api.github.com/some/endpoint', data=json.dumps({'some': 'data'}))5 print(r.json())

5）定制頭和cookie信息

header = {'user-agent': 'my-app/0.0.1''}cookie = {'key':'value'}
 r = requests.get/post('your url',headers=header,cookies=cookie)

data = {'some': 'data'}
headers = {'content-type': 'application/json',           'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:22.0) Gecko/20100101 Firefox/22.0'}
 
r = requests.post('https://api.github.com/some/endpoint', data=data, headers=headers)print(r.text)

6）響應(yīng)狀態(tài)碼

使用requests方法后，會返回一個response對象，其存儲了服務(wù)器響應(yīng)的內(nèi)容，如上實例中已經(jīng)提到的 r.text、r.status_code……
獲取文本方式的響應(yīng)體實例：當(dāng)你訪問 r.text 之時，會使用其響應(yīng)的文本編碼進(jìn)行解碼，并且你可以修改其編碼讓 r.text 使用自定義的編碼進(jìn)行解碼。

1 r = requests.get('http://www.')2 print(r.text, '\n{}\n'.format('*'*79), r.encoding)3 r.encoding = 'GBK'4 print(r.text, '\n{}\n'.format('*'*79), r.encoding)

示例代碼：

1 import requests2 
3 r = requests.get('https://github.com/Ranxf')       # 最基本的不帶參數(shù)的get請求4 print(r.status_code)                               # 獲取返回狀態(tài)5 r1 = requests.get(url='http://dict.baidu.com/s', params={'wd': 'python'})      # 帶參數(shù)的get請求6 print(r1.url)7 print(r1.text)        # 打印解碼后的返回數(shù)據(jù)

運行結(jié)果：

/usr/bin/python3.5 /home/rxf/python3_1000/1000/python3_server/python3_requests/demo1.py200http://dict.baidu.com/s?wd=python
…………

Process finished with exit code 0

 r.status_code                      #如果不是200，可以使用 r.raise_for_status() 拋出異常

7）響應(yīng)

r.headers                                  #返回字典類型,頭信息r.requests.headers                         #返回發(fā)送到服務(wù)器的頭信息r.cookies                                  #返回cookier.history                                  #返回重定向信息,當(dāng)然可以在請求是加上allow_redirects = false 阻止重定向

8）超時

r = requests.get('url',timeout=1)           #設(shè)置秒數(shù)超時，僅對于連接有效

9)會話對象，能夠跨請求保持某些參數(shù)

s = requests.Session()
s.auth = ('auth','passwd')
s.headers = {'key':'value'}
r = s.get('url')
r1 = s.get('url1')

10）代理

proxies = {'http':'ip1','https':'ip2' }
requests.get('url',proxies=proxies)

匯總：

# HTTP請求類型# get類型r = requests.get('https://github.com/timeline.json')# post類型r = requests.post("http://m./post")# put類型r = requests.put("http://m./put")# delete類型r = requests.delete("http://m./delete")# head類型r = requests.head("http://m./head")# options類型r = requests.options("http://m./get")# 獲取響應(yīng)內(nèi)容print(r.content) #以字節(jié)的方式去顯示，中文顯示為字符print(r.text) #以文本的方式去顯示#URL傳遞參數(shù)payload = {'keyword': '香港', 'salecityid': '2'}
r = requests.get("http://m./webapp/tourvisa/visa_list", params=payload) 
print（r.url） #示例為http://m./webapp/tourvisa/visa_list?salecityid=2&keyword=香港#獲取/修改網(wǎng)頁編碼r = requests.get('https://github.com/timeline.json')print （r.encoding）#json處理r = requests.get('https://github.com/timeline.json')print（r.json()） # 需要先import json    # 定制請求頭url = 'http://m.'headers = {'User-Agent' : 'Mozilla/5.0 (Linux; Android 4.2.1; en-us; Nexus 4 Build/JOP40D) AppleWebKit/535.19 (KHTML, like Gecko) Chrome/18.0.1025.166 Mobile Safari/535.19'}
r = requests.post(url, headers=headers)print （r.request.headers)#復(fù)雜post請求url = 'http://m.'payload = {'some': 'data'}
r = requests.post(url, data=json.dumps(payload)) #如果傳遞的payload是string而不是dict，需要先調(diào)用dumps方法格式化一下# post多部分編碼文件url = 'http://m.'files = {'file': open('report.xls', 'rb')}
r = requests.post(url, files=files)# 響應(yīng)狀態(tài)碼r = requests.get('http://m.')print(r.status_code)    
# 響應(yīng)頭r = requests.get('http://m.')print (r.headers)print (r.headers['Content-Type'])print (r.headers.get('content-type')) #訪問響應(yīng)頭部分內(nèi)容的兩種方式
    # Cookiesurl = 'http:///some/cookie/setting/url'r = requests.get(url)
r.cookies['example_cookie_name']    #讀取cookies    url = 'http://m./cookies'cookies = dict(cookies_are='working')
r = requests.get(url, cookies=cookies) #發(fā)送cookies#設(shè)置超時時間r = requests.get('http://m.', timeout=0.001)#設(shè)置訪問代理proxies = {           "http": "http://10.10.1.10:3128",           "https": "http://10.10.1.100:4444",
          }
r = requests.get('http://m.', proxies=proxies)#如果代理需要用戶名和密碼，則需要這樣：proxies = {    "http": "http://user:pass@10.10.1.10:3128/",
}

# HTTP請求類型# get類型r = requests.get('https://github.com/timeline.json')# post類型r = requests.post("http://m./post")# put類型r = requests.put("http://m./put")# delete類型r = requests.delete("http://m./delete")# head類型r = requests.head("http://m./head")# options類型r = requests.options("http://m./get")# 獲取響應(yīng)內(nèi)容print(r.content) #以字節(jié)的方式去顯示，中文顯示為字符print(r.text) #以文本的方式去顯示#URL傳遞參數(shù)payload = {'keyword': '香港', 'salecityid': '2'}
r = requests.get("http://m./webapp/tourvisa/visa_list", params=payload) 
print（r.url） #示例為http://m./webapp/tourvisa/visa_list?salecityid=2&keyword=香港#獲取/修改網(wǎng)頁編碼r = requests.get('https://github.com/timeline.json')print （r.encoding）#json處理r = requests.get('https://github.com/timeline.json')print（r.json()） # 需要先import json    # 定制請求頭url = 'http://m.'headers = {'User-Agent' : 'Mozilla/5.0 (Linux; Android 4.2.1; en-us; Nexus 4 Build/JOP40D) AppleWebKit/535.19 (KHTML, like Gecko) Chrome/18.0.1025.166 Mobile Safari/535.19'}
r = requests.post(url, headers=headers)print （r.request.headers)#復(fù)雜post請求url = 'http://m.'payload = {'some': 'data'}
r = requests.post(url, data=json.dumps(payload)) #如果傳遞的payload是string而不是dict，需要先調(diào)用dumps方法格式化一下# post多部分編碼文件url = 'http://m.'files = {'file': open('report.xls', 'rb')}
r = requests.post(url, files=files)# 響應(yīng)狀態(tài)碼r = requests.get('http://m.')print(r.status_code)    
# 響應(yīng)頭r = requests.get('http://m.')print (r.headers)print (r.headers['Content-Type'])print (r.headers.get('content-type')) #訪問響應(yīng)頭部分內(nèi)容的兩種方式
    # Cookiesurl = 'http:///some/cookie/setting/url'r = requests.get(url)
r.cookies['example_cookie_name']    #讀取cookies    url = 'http://m./cookies'cookies = dict(cookies_are='working')
r = requests.get(url, cookies=cookies) #發(fā)送cookies#設(shè)置超時時間r = requests.get('http://m.', timeout=0.001)#設(shè)置訪問代理proxies = {           "http": "http://10.10.1.10:3128",           "https": "http://10.10.1.100:4444",
          }
r = requests.get('http://m.', proxies=proxies)#如果代理需要用戶名和密碼，則需要這樣：proxies = {    "http": "http://user:pass@10.10.1.10:3128/",
}

3、示例代碼

GET請求

1 # 1、無參數(shù)實例
 2   
 3 import requests 4   
 5 ret = requests.get('https://github.com/timeline.json') 6   
 7 print(ret.url) 8 print(ret.text) 9   
10   
11   
12 # 2、有參數(shù)實例13   
14 import requests15   
16 payload = {'key1': 'value1', 'key2': 'value2'}17 ret = requests.get("http:///get", params=payload)18   
19 print(ret.url)20 print(ret.text)

POST請求

# 1、基本POST實例
  import requests
  
payload = {'key1': 'value1', 'key2': 'value2'}
ret = requests.post("http:///post", data=payload)  
print(ret.text)  
  
# 2、發(fā)送請求頭和數(shù)據(jù)實例
  import requestsimport json
  
url = 'https://api.github.com/some/endpoint'payload = {'some': 'data'}
headers = {'content-type': 'application/json'}
  
ret = requests.post(url, data=json.dumps(payload), headers=headers)  
print(ret.text)print(ret.cookies)

請求參數(shù)

參數(shù)示例代碼

json請求：

#! /usr/bin/python3import requestsimport jsonclass url_request():    def __init__(self):        ''' init '''if __name__ == '__main__':
    heard = {'Content-Type': 'application/json'}
    payload = {'CountryName': '中國',               'ProvinceName': '四川省',               'L1CityName': 'chengdu',               'L2CityName': 'yibing',               'TownName': '',               'Longitude': '107.33393',               'Latitude': '33.157131',               'Language': 'CN'}
    r = requests.post("http://www./CityLocation/json/LBSLocateCity", heards=heard, data=payload)
    data = r.json()    if r.status_code!=200:        print('LBSLocateCity API Error' + str(r.status_code))    print(data['CityEntities'][0]['CityID'])  # 打印返回json中的某個key的value
    print(data['ResponseStatus']['Ack'])    print(json.dump(data, indent=4, sort_keys=True, ensure_ascii=False))  # 樹形打印json，ensure_ascii必須設(shè)為False否則中文會顯示為unicode

Xml請求：

#! /usr/bin/python3import requestsclass url_request():    def __init__(self):        """init"""if __name__ == '__main__':
    heards = {'Content-type': 'text/xml'}
    XML = '<?xml version="1.0" encoding="utf-8"?><soap:Envelope xmlns:xsi="http://www./2001/XMLSchema-instance" xmlns:xsd="http://www./2001/XMLSchema" xmlns:soap="http://schemas./soap/envelope/"><soap:Body><Request xmlns="http:///"><jme><JobClassFullName>WeChatJSTicket.JobWS.Job.JobRefreshTicket,WeChatJSTicket.JobWS</JobClassFullName><Action>RUN</Action><Param>1</Param><HostIP>127.0.0.1</HostIP><JobInfo>1</JobInfo><NeedParallel>false</NeedParallel></jme></Request></soap:Body></soap:Envelope>'
    url = 'http://jobws.push.mobile.xx/RefreshWeiXInTokenJob/RefreshService.asmx'
    r = requests.post(url=url, heards=heards, data=XML)
    data = r.text    print(data)

狀態(tài)異常處理

import requests

URL = 'http://ip.taobao.com/service/getIpInfo.php'  # 淘寶IP地址庫APItry:
    r = requests.get(URL, params={'ip': '8.8.8.8'}, timeout=1)
    r.raise_for_status()  # 如果響應(yīng)狀態(tài)碼不是 200，就主動拋出異常except requests.RequestException as e:    print(e)else:
    result = r.json()    print(type(result), result, sep='\n')

上傳文件

使用request模塊，也可以上傳文件，文件的類型會自動進(jìn)行處理：

import requests
 
url = 'http://127.0.0.1:8080/upload'files = {'file': open('/home/rxf/test.jpg', 'rb')}#files = {'file': ('report.jpg', open('/home/lyb/sjzl.mpg', 'rb'))}     #顯式的設(shè)置文件名 r = requests.post(url, files=files)print(r.text)

request更加方便的是，可以把字符串當(dāng)作文件進(jìn)行上傳：

import requests
 
url = 'http://127.0.0.1:8080/upload'files = {'file': ('test.txt', b'Hello Requests.')}     #必需顯式的設(shè)置文件名 r = requests.post(url, files=files)print(r.text)

6) 身份驗證

基本身份認(rèn)證(HTTP Basic Auth)

import requestsfrom requests.auth import HTTPBasicAuth
 
r = requests.get('https:///hidden-basic-auth/user/passwd', auth=HTTPBasicAuth('user', 'passwd'))# r = requests.get('https:///hidden-basic-auth/user/passwd', auth=('user', 'passwd'))    # 簡寫print(r.json())

另一種非常流行的HTTP身份認(rèn)證形式是摘要式身份認(rèn)證，Requests對它的支持也是開箱即可用的:

requests.get(URL, auth=HTTPDigestAuth('user', 'pass')

Cookies與會話對象

如果某個響應(yīng)中包含一些Cookie，你可以快速訪問它們：

import requests
 
r = requests.get('http://www./')print(r.cookies['NID'])print(tuple(r.cookies))

要想發(fā)送你的cookies到服務(wù)器，可以使用 cookies 參數(shù)：

import requests
 
url = 'http:///cookies'cookies = {'testCookies_1': 'Hello_Python3', 'testCookies_2': 'Hello_Requests'}# 在Cookie Version 0中規(guī)定空格、方括號、圓括號、等于號、逗號、雙引號、斜杠、問號、@，冒號，分號等特殊符號都不能作為Cookie的內(nèi)容。r = requests.get(url, cookies=cookies)print(r.json())

會話對象讓你能夠跨請求保持某些參數(shù)，最方便的是在同一個Session實例發(fā)出的所有請求之間保持cookies，且這些都是自動處理的，甚是方便。
下面就來一個真正的實例，如下是快盤簽到腳本：

import requests
 
headers = {'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',           'Accept-Encoding': 'gzip, deflate, compress',           'Accept-Language': 'en-us;q=0.5,en;q=0.3',           'Cache-Control': 'max-age=0',           'Connection': 'keep-alive',           'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:22.0) Gecko/20100101 Firefox/22.0'}
 
s = requests.Session()
s.headers.update(headers)# s.auth = ('superuser', '123')s.get('https://www./account_login.htm')
 
_URL = 'http://www./index.php's.post(_URL, params={'ac':'account', 'op':'login'},
       data={'username':'****@foxmail.com', 'userpwd':'********', 'isajax':'yes'})
r = s.get(_URL, params={'ac':'zone', 'op':'taskdetail'})print(r.json())
s.get(_URL, params={'ac':'common', 'op':'usersign'})

requests模塊抓取網(wǎng)頁源碼并保存到文件示例

這是一個基本的文件保存操作，但這里有幾個值得注意的問題：

1.安裝requests包，命令行輸入pip install requests即可自動安裝。很多人推薦使用requests，自帶的urllib.request也可以抓取網(wǎng)頁源碼

2.open方法encoding參數(shù)設(shè)為utf-8，否則保存的文件會出現(xiàn)亂碼。

3.如果直接在cmd中輸出抓取的內(nèi)容，會提示各種編碼錯誤，所以保存到文件查看。

4.with open方法是更好的寫法，可以自動操作完畢后釋放資源

#! /urs/bin/python3import requests'''requests模塊抓取網(wǎng)頁源碼并保存到文件示例'''html = requests.get("http://www.baidu.com")
with open('test.txt', 'w', encoding='utf-8') as f:
    f.write(html.text)    
'''讀取一個txt文件，每次讀取一行，并保存到另一個txt文件中的示例'''ff = open('testt.txt', 'w', encoding='utf-8')
with open('test.txt', encoding="utf-8") as f:    for line in f:
        ff.write(line)
        ff.close()

因為在命令行中打印每次讀取一行的數(shù)據(jù)，中文會出現(xiàn)編碼錯誤，所以每次讀取一行并保存到另一個文件，這樣來測試讀取是否正常。（注意open的時候制定encoding編碼方式）

自動登陸"示例：

#!/usr/bin/env python# -*- coding:utf-8 -*-import requests# ############## 方式一 ##############"""# ## 1、首先登陸任何頁面，獲取cookie
i1 = requests.get(url="http://dig./help/service")
i1_cookies = i1.cookies.get_dict()

# ## 2、用戶登陸，攜帶上一次的cookie，后臺對cookie中的 gpsd 進(jìn)行授權(quán)
i2 = requests.post(
    url="http://dig./login",
    data={
        'phone': "8615131255089",
        'password': "xxooxxoo",
        'oneMonth': ""
    },
    cookies=i1_cookies
)

# ## 3、點贊（只需要攜帶已經(jīng)被授權(quán)的gpsd即可）
gpsd = i1_cookies['gpsd']
i3 = requests.post(
    url="http://dig./link/vote?linksId=8589523",
    cookies={'gpsd': gpsd}
)

print(i3.text)"""# ############## 方式二 ##############"""import requests

session = requests.Session()
i1 = session.get(url="http://dig./help/service")
i2 = session.post(
    url="http://dig./login",
    data={
        'phone': "8615131255089",
        'password': "xxooxxoo",
        'oneMonth': ""
    }
)
i3 = session.post(
    url="http://dig./link/vote?linksId=8589523"
)
print(i3.text)"""

#!/usr/bin/env python# -*- coding:utf-8 -*-import requestsfrom bs4 import BeautifulSoup# ############## 方式一 ###############
# # 1. 訪問登陸頁面，獲取 authenticity_token# i1 = requests.get('https://github.com/login')# soup1 = BeautifulSoup(i1.text, features='lxml')# tag = soup1.find(name='input', attrs={'name': 'authenticity_token'})# authenticity_token = tag.get('value')# c1 = i1.cookies.get_dict()# i1.close()#
# # 1. 攜帶authenticity_token和用戶名密碼等信息，發(fā)送用戶驗證# form_data = {# "authenticity_token": authenticity_token,#     "utf8": "",#     "commit": "Sign in",#     "login": "wupeiqi@live.com",#     'password': 'xxoo'# }#
# i2 = requests.post('https://github.com/session', data=form_data, cookies=c1)# c2 = i2.cookies.get_dict()# c1.update(c2)# i3 = requests.get('https://github.com/settings/repositories', cookies=c1)#
# soup3 = BeautifulSoup(i3.text, features='lxml')# list_group = soup3.find(name='div', class_='listgroup')#
# from bs4.element import Tag#
# for child in list_group.children:#     if isinstance(child, Tag):#         project_tag = child.find(name='a', class_='mr-1')#         size_tag = child.find(name='small')#         temp = "項目:%s(%s); 項目路徑:%s" % (project_tag.get('href'), size_tag.string, project_tag.string, )#         print(temp)# ############## 方式二 ############### session = requests.Session()# # 1. 訪問登陸頁面，獲取 authenticity_token# i1 = session.get('https://github.com/login')# soup1 = BeautifulSoup(i1.text, features='lxml')# tag = soup1.find(name='input', attrs={'name': 'authenticity_token'})# authenticity_token = tag.get('value')# c1 = i1.cookies.get_dict()# i1.close()#
# # 1. 攜帶authenticity_token和用戶名密碼等信息，發(fā)送用戶驗證# form_data = {#     "authenticity_token": authenticity_token,#     "utf8": "",#     "commit": "Sign in",#     "login": "wupeiqi@live.com",#     'password': 'xxoo'# }#
# i2 = session.post('https://github.com/session', data=form_data)# c2 = i2.cookies.get_dict()# c1.update(c2)# i3 = session.get('https://github.com/settings/repositories')#
# soup3 = BeautifulSoup(i3.text, features='lxml')# list_group = soup3.find(name='div', class_='listgroup')#
# from bs4.element import Tag#
# for child in list_group.children:#     if isinstance(child, Tag):#         project_tag = child.find(name='a', class_='mr-1')#         size_tag = child.find(name='small')#         temp = "項目:%s(%s); 項目路徑:%s" % (project_tag.get('href'), size_tag.string, project_tag.string, )#         print(temp)

知乎

#!/usr/bin/env python# -*- coding:utf-8 -*-import reimport jsonimport base64import rsaimport requestsdef js_encrypt(text):
    b64der = 'MIGfMA0GCSqGSIb3DQEBAQUAA4GNADCBiQKBgQCp0wHYbg/NOPO3nzMD3dndwS0MccuMeXCHgVlGOoYyFwLdS24Im2e7YyhB0wrUsyYf0/nhzCzBK8ZC9eCWqd0aHbdgOQT6CuFQBMjbyGYvlVYU2ZP7kG9Ft6YV6oc9ambuO7nPZh+bvXH0zDKfi02prknrScAKC0XhadTHT3Al0QIDAQAB'
    der = base64.standard_b64decode(b64der)

    pk = rsa.PublicKey.load_pkcs1_openssl_der(der)
    v1 = rsa.encrypt(bytes(text, 'utf8'), pk)
    value = base64.encodebytes(v1).replace(b'\n', b'')
    value = value.decode('utf8')    return value


session = requests.Session()

i1 = session.get('https://passport.cnblogs.com/user/signin')
rep = re.compile("'VerificationToken': '(.*)'")
v = re.search(rep, i1.text)
verification_token = v.group(1)

form_data = {    'input1': js_encrypt('wptawy'),    'input2': js_encrypt('asdfasdf'),    'remember': False
}

i2 = session.post(url='https://passport.cnblogs.com/user/signin',
                  data=json.dumps(form_data),
                  headers={                      'Content-Type': 'application/json; charset=UTF-8',                      'X-Requested-With': 'XMLHttpRequest',                      'VerificationToken': verification_token}
                  )

i3 = session.get(url='https://i.cnblogs.com/EditDiary.aspx')print(i3.text)

#!/usr/bin/env python# -*- coding:utf-8 -*-import requests# 第一步：訪問登陸頁,拿到X_Anti_Forge_Token，X_Anti_Forge_Code# 1、請求url:https://passport.lagou.com/login/login.html# 2、請求方法:GET# 3、請求頭:#    User-agentr1 = requests.get('https://passport.lagou.com/login/login.html',
                 headers={                     'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36',
                 },
                 )

X_Anti_Forge_Token = re.findall("X_Anti_Forge_Token = '(.*?)'", r1.text, re.S)[0]
X_Anti_Forge_Code = re.findall("X_Anti_Forge_Code = '(.*?)'", r1.text, re.S)[0]print(X_Anti_Forge_Token, X_Anti_Forge_Code)# print(r1.cookies.get_dict())# 第二步：登陸# 1、請求url:https://passport.lagou.com/login/login.json# 2、請求方法:POST# 3、請求頭:#    cookie#    User-agent#    Referer:https://passport.lagou.com/login/login.html#    X-Anit-Forge-Code:53165984#    X-Anit-Forge-Token:3b6a2f62-80f0-428b-8efb-ef72fc100d78#    X-Requested-With:XMLHttpRequest# 4、請求體：# isValidate:true# username:15131252215# password:ab18d270d7126ea65915c50288c22c0d# request_form_verifyCode:''# submit:''r2 = requests.post(    'https://passport.lagou.com/login/login.json',
    headers={        'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36',        'Referer': 'https://passport.lagou.com/login/login.html',        'X-Anit-Forge-Code': X_Anti_Forge_Code,        'X-Anit-Forge-Token': X_Anti_Forge_Token,        'X-Requested-With': 'XMLHttpRequest'
    },
    data={        "isValidate": True,        'username': '15131255089',        'password': 'ab18d270d7126ea65915c50288c22c0d',        'request_form_verifyCode': '',        'submit': ''
    },
    cookies=r1.cookies.get_dict()
)print(r2.text)

參考：

http://cn./zh_CN/latest/user/quickstart.html#id4

http://www./en/master/

http://docs./en/latest/user/quickstart/

https://www.cnblogs.com/tangdongchu/p/4229049.html#t0

http://www.cnblogs.com/wupeiqi/articles/6283017.html

分類: 爬蟲學(xué)習(xí)