xuxueli / xxl-crawler Goto Github PK

View Code? Open in Web Editor NEW

685.0 49.0 297.0 390 KB

A distributed web crawler framework.（分布式爬虫框架XXL-CRAWLER）

Home Page: http://www.xuxueli.com/xxl-crawler

License: Apache License 2.0

Java 100.00%

crawler web spider object-oriented flexible xxl-crawler java distributed

xxl-crawler's Introduction

XXL-CRAWLER

XXL-CRAWLER, a distributed web crawler framework.
-- Home Page --

Introduction

XXL-CRAWLER is a distributed web crawler framework. One line of code develops a distributed crawler. Features such as "multithreaded、asynchronous、dynamic IP proxy、distributed、javascript-rendering".

XXL-CRAWLER 是一个分布式爬虫框架。一行代码开发一个分布式爬虫，拥有"多线程、异步、IP动态代理、分布式、JS渲染"等特性；

Documentation

中文文档

Features

1、简洁：API直观简洁，可快速上手；
2、轻量级：底层实现仅强依赖jsoup，简洁高效；
3、模块化：模块化的结构设计，可轻松扩展
4、面向对象：支持通过注解，方便的映射页面数据到PageVO对象，底层自动完成PageVO对象的数据抽取和封装返回；单个页面支持抽取一个或多个PageVO
5、多线程：线程池方式运行，提高采集效率；
6、分布式支持：通过扩展 "RunData" 模块，并结合Redis或DB共享运行数据可实现分布式。默认提供LocalRunData单机版爬虫。
7、JS渲染：通过扩展 "PageLoader" 模块，支持采集JS动态渲染数据。原生提供 Jsoup(非JS渲染，速度更快)、HtmlUnit(JS渲染)、Selenium+Phantomjs(JS渲染，兼容性高) 等多种实现，支持自由扩展其他实现。
8、失败重试：请求失败后重试，并支持设置重试次数；
9、代理IP：对抗反采集策略规则WAF；
10、动态代理：支持运行时动态调整代理池，以及自定义代理池路由策略；
11、异步：支持同步、异步两种方式运行；
12、扩散全站：支持以现有URL为起点扩散爬取整站；
13、去重：防止重复爬取；
14、URL白名单：支持设置页面白名单正则，过滤URL；
15、自定义请求信息，如：请求参数、Cookie、Header、UserAgent轮询、Referrer等；
16、动态参数：支持运行时动态调整请求参数；
17、超时控制：支持设置爬虫请求的超时时间；
18、主动停顿：爬虫线程处理完页面之后进行主动停顿，避免过于频繁被拦截；

Communication

社区交流

Contributing

Contributions are welcome! Open a pull request to fix a bug, or open an Issue to discuss a new feature or change.

欢迎参与项目贡献！比如提交PR修复一个bug，或者新建 Issue 讨论新特性或者变更。

接入登记

更多接入的公司，欢迎在登记地址登记，登记仅仅为了产品推广。

Copyright and License

This product is open source and free, and will continue to provide free community technical support. Individual or enterprise users are free to access and use.

Licensed under the Apache License, Version 2.0.
Copyright (c) 2015-present, xuxueli.

产品开源免费，并且将持续提供免费的社区技术支持。个人或企业内部可自由的接入和使用。

Donate

No matter how much the amount is enough to express your thought, thank you very much ：） To donate

无论金额多少都足够表达您这份心意，非常感谢：）前往捐赠

xxl-crawler's People

Contributors

Stargazers

Watchers

Forkers

shangana jiachenging apple006 luzhu123 xcf9868 anoous chinesslight ahange zhonghuaxiaodangjia vector4wang tomzhang yankaics jnan88 sides8 tbno1 yoreay xbing1221 changyuxin ysj123688 nbsw zhaotianen hoffmanindustry xuxiaoxiao89 hwlsniper lurker8 heaiso1985 pinggle czou jacksn2014 whattwitter weicq angellee1988 dkgee liuhuiyong djh4230 skyformat99 hzchendou kangzhenkang key2wen spacegithub bbsyaya 1509305429 dalek42 restley dxf1122 guizhixiao andy521 chinarefers z01eternal februay xufebruary norelax lomoye skyworker0725 hiekay minemine678 leejones92 changemeclub yang731022 kevinyzy1 hereyouareeee guyueyuqi xuyunjeff xushuai2 yangyangxiong zhiqinghuang yangth xiebaofu gqfjob wallaceok tianyi19970120 xuegao2015 bossding geekyouth omgteam chali2017 ronanana inferoiny louisjong yehuangcn xiaoyong601 githubcy lisee lovedabai peakve llg-software mohnsnow newlysoft giserh landk1003 magcle nero520 rarshion dngiveu diycp wehle 664138519wj jackjet dream7319 bongdadienanh1

xxl-crawler's Issues

JsoupUtil工具类loadPageSource()方法里Connection没有调用requestBody

JsoupUtil工具类loadPageSource()方法里Connection没有调用requestBody，有的接口要求只能通过Connection.requestBody()传递参数，这种情况下，抓取不到数据。

发送post请求时返回400

你好，我在测试用例中没有找到post请求的模板调用

这是我的调用代码
` Map<String,String> dataMap = new HashMap<>();
dataMap.put("category","**");
dataMap.put("currentPage","1");
dataMap.put("pageSize","30");

    Map<String,String> headerMap = new HashMap<>();
    headerMap.put("Accept-Encoding","gzip");
    headerMap.put("Content-Type","application/json;charset=UTF-8");
    headerMap.put("User-Agent","Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/55.0.2883.87 Safari/537.36");

    XxlCrawler xxlCrawler = new XxlCrawler.Builder()
            .setUrls(url)
            .setAllowSpread(false)
            .setIfPost(true)
            .setHeaderMap(headerMap)
            .setParamMap(dataMap)
            .setPageParser(new PageParser() {
                @Override
                public void parse(Document html, Element pageVoElement, Object pageVo) {
                    XxlJobLogger.log("html:{}",html);
                }
            })
            .build();
    xxlCrawler.start(true);
    return SUCCESS;`

这是报错：
org.jsoup.HttpStatusException: HTTP error fetching URL. Status=400

线程安全问题

LocalRunData 中使用 LinkedBlockingQueue 来记录需要爬取的url, 这是一个线程安全的队列, 还需要加 volatile 关键字吗 ?

CrawlerThread的process方法里判断当前链接是否是白名单链接逻辑有问题

// ------- pagevo ----------
if (!crawler.getRunConf().validWhiteUrl(link)) { // limit unvalid-page parse, only allow spread child
return false;
}

这一段代码返回false，如果用户设置了重试次数，会导致无意义的重试。这里应该返回true

接入xxl-crawler的公司请留下 ”公司名称 + 公司官网地址“，谢谢。

[issue] 多线程情况下，tryFinish()很小的概率会误判当前运行状态

issue description：

多线程情况下，tryFinish()会误判CrawlerThread的运行状态，导致提前stop，以下是运行XxlCrawlerTest，开启3个thread，并打印日志：

概率比较小，大概试10次能出现一次，原因可能如下：
thread-3调用tryFinish()并提前获取了3个CrawlerThread的isRunning状态均为false，刚好此时thread-1调用了crawler.getRunData().getUrl()并将running设为true（但thread-3已经无法知晓），最后thread-3判断runData.getUrlNum()==0为true，由此isEnd为true，导致了误判：

solution：

改写tryFinish()，先判断runData.getUrlNum()==0，再逐一获取CrawlerThread的状态，防止调用crawler.getRunData().getUrl()无法获取running的最新状态：

public void tryFinish(){
    boolean isEnd = runData.getUrlNum()==0;
    boolean isRunning = false;
    for (CrawlerThread crawlerThread: crawlerThreads) {
        if (crawlerThread.isRunning()) {
            isRunning = true;
            break;
        }
    }
    isEnd = isEnd && !isRunning;
    if (isEnd) {
        logger.info(">>>>>>>>>>> xxl crawler is finished.");
        stop();
    }
}

CrawlerThread的running参数加上volatile关键字，保证可见性：

private volatile boolean running;

ajax请求爬取

是否支持ajax(json响应)请求的爬取

【需求】VO嵌套

PageFieldSelect能否使用复杂类型？希望爬取1-N的数据结构

setWhiteUrlRegexs正则传参不起作用

setWhiteUrlRegexs("https://www.kuaidaili.com/free/inha/\\b[1-2]/")
例如这种方式它不会匹配两个url，whiteUrlRegexs.length为1

能否支持获取js执行之后的网页

类似下面那样，或者是js生成的节点。
`

扩散全站功能异常问题.

打开了扩散全站的功能, 但是在 JsoupUtil.findLinks()方法中筛选到的url不全, 标签获得的href是相对路径, 不是决定路径. 使用下面三种方法获得的值全部是相对路径, 校验url不通过导致, 扩散爬取失败, 大佬有遇到过这种情况吗 ?
tips: 使用 JS渲染方式采集数据，"selenisum + phantomjs" 方案

item.absUrl("abs:href");
item.attr("abs:href");
item.attr("href");

爬取的url是 http://www.bootcss.com/

connect timeout超时处理

如何针对对某个url的connect timeout超时做出判断处理，或者重新加入待爬取内容

使用SeleniumPhantomjsPageLoader后，jsoup解析后document对象中的baseUri为空

是否允许基于身份认证的爬虫

实现简单的只有用户名与密码的登陆授权，获取token或session来爬其他的页面。

优化setPageParser避免匿名函数

Document 和 Element对象可以在其他地方取出,使用匿名函数的调用方式并不是很方便

maven引入1.2.2版本，测试07报错

建议使用jdk1.8

建议使用jdk 1.8

[新需求]针对post请求，相同的url，根据参数不同返回不同结果的页面抓取实现

针对post请求，相同的url，根据参数不同返回不同结果的页面抓取实现
是否可考虑在解析页面结果的类中返回当前爬虫对象，这样可以在处理完上一个页面抓取后，向爬虫对象中的url队列添加新的url。增强现在的只能在爬虫初始化的时候添加url（或者只能粗犷的扩散爬取）功能。

com.xuxueli.crawler.thread.CrawlerThread#processPage问题

com.xuxueli.crawler.thread.CrawlerThread#processPage中以下代码应该return false比较合适吧？

if (!crawler.getRunConf().validWhiteUrl(pageRequest.getUrl())) {     // limit unvalid-page parse, only allow spread child, finish here
            return true;
        }