Commit Graph

411 Commits (ef325718219ecb3cef56911dd72ea04c40dd8673)

Author SHA1 Message Date
yihua.huang ef32571821 rewrite Request.equals and hashCode, add Method to check #483 2017-03-11 10:52:39 +08:00
yihua.huang 8b8f535c30 refactor:extract charset detect to utils 2017-03-11 10:43:10 +08:00
yihua.huang a872a6480e fix code sample for github #348 2017-02-25 22:46:29 +08:00
yihua.huang 1d2171805f add test for #228 2017-02-25 22:30:48 +08:00
yihua.huang bbe0b52ddd remove synchronized in QueueScheduler #410 2017-02-25 19:55:45 +08:00
yihua.huang ad69963005 remove synchronize in Page #411 2017-02-25 19:42:12 +08:00
yihua.huang 3a796b9413 remove duplicate code #421 2017-02-25 12:01:12 +08:00
yihua.huang 42f1018010 remove messy code 2017-02-21 14:08:05 +08:00
yihua.huang aaccc93215 new version 2017-01-21 12:04:12 +08:00
yihua.huang 3e633c6871 version 2017-01-21 11:51:14 +08:00
yihua.huang f45e2f118b for release 2017-01-21 11:38:36 +08:00
yihua.huang d60615f503 修复使用startUrls没有设置domain导致使用cookie空指针的问题#438 2017-01-21 11:29:42 +08:00
yihua.huang 407fbb6130 refactor logger#445 2017-01-21 11:05:54 +08:00
Ckex.zha 0dc26c8ca0 optimize code. 2017-01-20 14:03:26 +08:00
Yihua Huang 4f76d62d4f Merge pull request #444 from ckex/develop
绕过安全证书
2017-01-18 23:42:51 +08:00
Ckex.zha e4af05a6f2 绕过安全证书 2017-01-18 17:28:01 +08:00
xbynet@outlook.com c23627bf63 解决post/redirect/post 302跳转问题 2017-01-17 00:07:01 +08:00
yihua.huang d69204b919 0.6.0 2016-12-18 11:45:43 +08:00
yihua.huang 9bdb48b2d0 version 0.6.0 2016-12-18 11:20:28 +08:00
yihua.huang eeb607fd0e 将Spider.processRequest()抛出异常改回原来的逻辑 2016-12-18 11:04:58 +08:00
yihua.huang 97592d6720 Version 0.6.0 2016-12-18 10:58:24 +08:00
yihua.huang 00dfebbceb #424 remove guava dep and add fix docs 2016-12-18 10:45:50 +08:00
yihua.huang c2531c6817 clean dependency 2016-12-18 08:34:46 +08:00
yihua.huang a960a39c44 fix compile error for example change 2016-12-18 08:32:14 +08:00
yihua.huang 7476ceccee more stable test 2016-12-18 08:15:26 +08:00
yihua.huang 5ce3fdfe5a some refactor in log 2016-12-18 08:15:09 +08:00
yihua.huang 98163a3e40 update examples 2016-12-18 07:46:18 +08:00
yihua.huang b090dcd20d sepcific error page for HttpClientDownloaderTest to avoid test error when local port is available 2016-12-18 07:15:06 +08:00
yihua.huang 8f942d6fe2 #419 修复抓取https链接线程无法结束导致进程一直运行的问题 2016-12-18 06:56:01 +08:00
yihua.huang dafd2b77ff fix GithubRepoPageProcessor in example 2016-11-24 08:18:06 +08:00
yihua.huang cfed860fb9 Merge branch 'master' of github.com:code4craft/webmagic 2016-11-22 17:00:27 +08:00
yihua.huang 2189aab652 fix test 2016-11-22 16:58:49 +08:00
Yihua Huang 1491033534 Merge pull request #377 from jerry-sc/monitor-bug
fix the monitor bug which the spider will terminate when a seed url with port
2016-11-19 13:01:30 +08:00
yihua.huang 507556d0aa fix test: ProxyTest.testProxy() do not load exist proxy config 2016-11-19 12:54:39 +08:00
Jerry e56b8c3efc fix the monitor bug which the spider will terminate when a seed url with port 2016-09-22 22:36:18 +08:00
yihua.huang 448e528140 update StringUtils to apache lang3 #314 2016-05-24 13:33:17 +08:00
yihua.huang 3e33959b7a #319 fix javadoc 2016-05-24 13:17:35 +08:00
yihua.huang 8730e3e97a Merge branch 'fix' of git://github.com/kapsterio/webmagic into kapsterio-fix 2016-05-08 20:46:22 +08:00
yihua.huang 2400ff7e1a resovle conflict 2016-05-08 20:31:43 +08:00
yihua.huang b7f3c4bba0 Merge branch 'master' of git://github.com/hepan/webmagic into hepan-master 2016-05-08 20:27:47 +08:00
yihua.huang d8f978fd20 fix test in JsonPathSelectorTest #289 2016-05-08 19:32:03 +08:00
yihua.huang 61c28a0130 refactor on proxypool 2016-05-08 17:53:15 +08:00
yihua.huang b871b210c5 Merge branch 'proxy-strategy' of github.com:EdwardsBean/webmagic into EdwardsBean-proxy-strategy 2016-05-08 17:53:02 +08:00
yihua.huang b5413368de update ut 2016-05-08 16:23:41 +08:00
Jon 83c27ebbc4 增加IP代理认证功能 2016-05-08 16:17:58 +08:00
yihua.huang ca072c5575 fix URL regex in GithubRepoPageProcessor #305 2016-05-08 12:09:45 +08:00
hepan 89c6e52863 代理增加用户名密码认证 2016-04-13 15:16:57 +08:00
Linker Lin 047cb8ff8f updated versions to 0.5.4-SNAPSHOT 2016-04-01 14:51:59 +08:00
zhangheng09 6b179c3d55 这个改动的原因基于两点:1)代理归还给代理池的时机应该是执行完http请求后就要尽早归还 2)http代理应该是HttpClientDownloader该考虑的事,不应该有Spider来处理,Spider并不知道它的downloader是个HttpClientDownloader 2016-03-12 20:09:41 +08:00
zhangheng09 5f106c9c69 当page为null时,意味着非正常的响应状态,应该抛出异常,否则SpiderListener的onSuccess方法和onError方法都会执行 2016-03-12 20:03:27 +08:00