Light

chenjiandongx / bili-spider Goto Github PK

View Code? Open in Web Editor NEW

595.0 26.0 183.0 62.16 MB

📺 B 站全站视频信息爬虫

Home Page: https://github.com/chenjiandongx/bili-video-rank

License: MIT License

Python 100.00%

bilibili spider

bili-spider's Introduction

B 站全站视频信息爬虫

B 站我想大家都熟悉吧，其实 B 站的爬虫网上一搜一大堆。不过 纸上得来终觉浅，绝知此事要躬行，我码故我在。最终爬取到数据总量为 1300 万 条。

开发环境为：Windows10 + python3

准备工作

首先打开 B 站，随便在首页找一个视频点击进去。常规操作，打开开发者工具。这次是目标是通过爬取 B 站提供的 api 来获取视频信息，不去解析网页，解析网页的速度太慢了而且容易被封 ip。

勾选 JS 选项，F5 刷新

找到了 api 的地址

复制下来，去除没必要的内容，得到 https://api.bilibili.com/x/web-interface/archive/stat?aid=15906633 ，用浏览器打开，会得到如下的 json 数据

动手写码

好了，到这里代码就可以码起来了，通过 request 不断的迭代获取数据，为了让爬虫更高效，可以利用多线程。

核心代码

result = []
req = requests.get(url, headers=headers, timeout=6).json()
time.sleep(0.6)     # 延迟，避免太快 ip 被封
try:
    data = req['data']
    video = (
        total,
        data['aid'],        # 视频编号
        data['view'],       # 播放量
        data['danmaku'],    # 弹幕数
        data['reply'],      # 评论数
        data['favorite'],   # 收藏数
        data['coin'],       # 硬币数
        data['share']       # 分享数
    )
    with lock:
        result.append(video)
        if total % 100 == 0:
            print(total)
        total += 1
except:
    pass

迭代爬取

urls = ["http://api.bilibili.com/archive_stat/stat?aid={}".format(i)
        for i in range(10000)]
with futures.ThreadPoolExecutor(32) as executor:    # 多线程
    executor.map(run, urls)

爬取后数据存放进了 MySQL 数据库，总共爬取到了 1300w+ 条数据

前 750w 条数据在这里 bili.zip

bili-spider's People

Contributors

Stargazers

Watchers

Forkers

rebortboss xiaopeng163 zzz233 timedcy jmagenif fwind1 yaochao eysdo samirchen flyme6 hidesoon glien-king eeeeeeeeeeeeeeeeeeeieeeeeeeeeeeeeeeeee lxw4939 huahuijay ysyyy starte hilly0420 erbuer overad ysj123688 zhianlin sunhuang163 peterdocter houzhong archon98 jnulyd cuitxubin hengthu socratesclub zhaohuiqiang wangler2333 cloudpai wangako xrxr19990625 zhoujiawei1993 aohanhongzhi virlier arryboom lxt-python skythebug amapoftheworld qing2qi smartcrane2 wangdh1995 visionwxc yipeng0428 tinasunlin dreamren doublefish20170305 saoinformaticsteam qjhqqqqq wealthe jaygith king1348 freeman66 stoneby lzl726 gaoyuqi sunshinelist liangdong-xjtu yaoxingqi silence28 kusanagibo gonciaper tuskumo233 hhy5277 xieyh730 limuitech lingxuehan lusonpan62678 longtan01 lstarby mimorinian danielislearning xianfengting david56038 rribons kamihati foxgeek36 jiangguoqing huwenshu brahamack jieseo guygubaby asforme zhouzhouanya qweadqw alientales xhui28 shezhenbo62 real-shigure emmalui aaronchiu2017 bruce-zhangs jinligen albertontheway zhonggithub tigergo001 buaawyq

bili-spider's Issues

b站新增了BV号

循环最大到多少合适？

for i in range(10000)

爬虫能否得到视频的真实播放地址信息，及flv格式的视频

例如：
https://app.bilibili.com/v2/playurl?appkey=YvirImLGlLANCLvM&build=6680&buvid=1134402fc290d710313b9db99e310eed&cid=39922399&device=phone&otype=json&platform=iphone&qn=16&sign=36e8ec6971b6c089409aea3ee9ecbf64
目前这个接口，还没有搞清楚sign的生成算法。

一个数字也不走是被封IP了嘛

刚才成功爬了一万条，于是把数量改成50万开始爬呀爬，爬到3百多就再也不走了。。再次运行程序一条也不走了

大佬我是不是被封IP了，B站可以照常上 @chenjiandongx

多进程在哪里

……大佬……没看到启动进程的地方啊……

关于视频名称

没有视频名称看起来数据内容有点不直观，作者有考虑过对应AV号视频名称的获取吗？有没有比解析网页得到视频名称更好的方式呢？

对数据感兴趣的可以留下邮箱

如题，顺便给个 star 什么的

可否共享下数据？

Recommend Projects

React

A declarative, efficient, and flexible JavaScript library for building user interfaces.
Vue.js

🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
Typescript

TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
TensorFlow

An Open Source Machine Learning Framework for Everyone
Django

The Web framework for perfectionists with deadlines.
Laravel

A PHP framework for web artisans
D3

Bring data to life with SVG, Canvas and HTML. 📊📈🎉

Recommend Topics

javascript

JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
web

Some thing interesting about web. New door for the world.
server

A server is a program made to process requests and deliver data to clients.
Machine learning

Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
Visualization

Some thing interesting about visualization, use data art
Game

Some thing interesting about game, make everyone happy.

Recommend Org

Facebook

We are working to build community through open source technology. NB: members must have two-factor auth.
Microsoft

Open source projects and samples from Microsoft.
Google

Google ❤️ Open Source for everyone.
Alibaba

Alibaba Open Source for everyone
D3

Data-Driven Documents codes.
Tencent

China tencent open source team.