Yihua Huang
1491033534
Merge pull request #377 from jerry-sc/monitor-bug
...
fix the monitor bug which the spider will terminate when a seed url with port
2016-11-19 13:01:30 +08:00
yihua.huang
507556d0aa
fix test: ProxyTest.testProxy() do not load exist proxy config
2016-11-19 12:54:39 +08:00
Jerry
e56b8c3efc
fix the monitor bug which the spider will terminate when a seed url with port
2016-09-22 22:36:18 +08:00
yihua.huang
448e528140
update StringUtils to apache lang3 #314
2016-05-24 13:33:17 +08:00
yihua.huang
3e33959b7a
#319 fix javadoc
2016-05-24 13:17:35 +08:00
yihua.huang
8730e3e97a
Merge branch 'fix' of git://github.com/kapsterio/webmagic into kapsterio-fix
2016-05-08 20:46:22 +08:00
yihua.huang
2400ff7e1a
resovle conflict
2016-05-08 20:31:43 +08:00
yihua.huang
b7f3c4bba0
Merge branch 'master' of git://github.com/hepan/webmagic into hepan-master
2016-05-08 20:27:47 +08:00
yihua.huang
d8f978fd20
fix test in JsonPathSelectorTest #289
2016-05-08 19:32:03 +08:00
yihua.huang
61c28a0130
refactor on proxypool
2016-05-08 17:53:15 +08:00
yihua.huang
b871b210c5
Merge branch 'proxy-strategy' of github.com:EdwardsBean/webmagic into EdwardsBean-proxy-strategy
2016-05-08 17:53:02 +08:00
yihua.huang
b5413368de
update ut
2016-05-08 16:23:41 +08:00
Jon
83c27ebbc4
增加IP代理认证功能
2016-05-08 16:17:58 +08:00
yihua.huang
ca072c5575
fix URL regex in GithubRepoPageProcessor #305
2016-05-08 12:09:45 +08:00
hepan
89c6e52863
代理增加用户名密码认证
2016-04-13 15:16:57 +08:00
Linker Lin
047cb8ff8f
updated versions to 0.5.4-SNAPSHOT
2016-04-01 14:51:59 +08:00
zhangheng09
6b179c3d55
这个改动的原因基于两点:1)代理归还给代理池的时机应该是执行完http请求后就要尽早归还 2)http代理应该是HttpClientDownloader该考虑的事,不应该有Spider来处理,Spider并不知道它的downloader是个HttpClientDownloader
2016-03-12 20:09:41 +08:00
zhangheng09
5f106c9c69
当page为null时,意味着非正常的响应状态,应该抛出异常,否则SpiderListener的onSuccess方法和onError方法都会执行
2016-03-12 20:03:27 +08:00
yihua.huang
c0b8e8f8ae
remove .classpath .project
2016-01-22 14:58:22 +08:00
yihua.huang
a8e6de4b90
Merge branch 'master' of git.oschina.net:flashsword20/webmagic
2016-01-22 10:16:58 +08:00
yihua.huang
0fd4623f0a
Merge branch 'osc'
2016-01-21 19:33:30 +08:00
yihua.huang
ce5495ecd5
remove useless files
2016-01-21 19:31:50 +08:00
yihua.huang
8265c7dade
remove submodules for relase
2016-01-21 19:25:13 +08:00
yihua.huang
7edfa26f90
complete javadoc
2016-01-21 18:34:07 +08:00
yihua.huang
8b90b91e33
complete some javadoc
2016-01-21 18:14:10 +08:00
yihua.huang
2b556cf053
update verison to 0.5.3-SNAPSHOT
2016-01-21 18:05:56 +08:00
yihua.huang
9c5716a543
complete javadoc
2016-01-21 18:05:12 +08:00
yihua.huang
db3cbf6ca5
update version to 0.5.3-SNAPSHOT
2016-01-21 17:58:36 +08:00
yihua.huang
81ce1ffc5f
fix ignore
2016-01-21 12:36:49 +08:00
yihua.huang
93764fa2c9
ignore some test
2016-01-21 12:28:32 +08:00
yihua.huang
5706bb90af
update xsoup to 0.3.1
2016-01-20 12:59:11 +08:00
yihua.huang
7586e3d75c
add some test for github repo downloader
2016-01-19 08:05:53 +08:00
x1ny
90e14b31b0
修正FileCacheQueueScheduler导致程序不能正常结束和未关闭流
...
FileCacheQueueScheduler中开启了一个线程周期运行来保存数据但在爬虫结束后没有关闭导致程序无法结束,以及没有关闭io流。
解决方法:
让FileCacheQueueScheduler实现Closable接口,在close方法中关闭线程以及流。
在Spider的close方法中添加对scheduler的关闭操作。
2015-11-12 23:10:20 +08:00
yihua.huang
56e0cd513a
compile error fix
2015-04-15 23:21:06 +08:00
yihua.huang
c5740b1840
change assert #200
2015-04-15 08:32:08 +08:00
yihua.huang
67eb632f4d
test for issue #200
2015-04-15 08:31:45 +08:00
高军
590561a6e4
修正site.setHttpProxy()不起作用的bug
2015-03-09 15:54:15 +08:00
edwardsbean
19474e4716
add SimpleProxyPool and IProxyPool
2015-02-28 17:50:10 +08:00
edwardsbean
4978665633
add retry sleep time
2015-01-21 13:30:02 +08:00
yihua.huang
8ffc1a7093
add NPE check for POST method
2015-01-13 14:10:00 +08:00
zhugw
bc666e927d
Update Site.java
...
setCycleRetryTimes的javadoc是这么说的:Set cycleRetryTimes times when download fail, 0 by default. Only work in RedisScheduler.
而通过查看源码发现似乎并没有做限制,即只能用于RedisScheduler. 故想问一下该javadoc是否过时了?
2014-09-12 12:42:57 +08:00
yihua.huang
147401ce5e
remove duplicate setPath in ProxyPool
2014-09-09 22:58:44 +08:00
yihua.huang
e7668e01b8
fix SourceRegion error and add some tests on it #144
2014-08-21 14:29:06 +08:00
yihua.huang
4446669c24
fix test
2014-08-18 10:54:24 +08:00
yihua.huang
9866297ec4
Disable jsoup entity escape by Default. Set Html.DISABLE_HTML_ENTITY_ESCAPE to false to enable it. #149
2014-08-14 08:04:56 +08:00
yihua.huang
4e6e946dd7
more friendly exception message in PlainText #144
2014-08-13 10:02:16 +08:00
yihua.huang
af9939622b
move thread package out of selector (because it is add by mistake at the beginning)
2014-06-25 18:19:50 +08:00
yihua.huang
eae37c868b
new sample
2014-06-10 17:38:54 +08:00
yihua.huang
b3a282e58d
some fix for tests #130
2014-06-10 00:05:30 +08:00
yihua.huang
074d767f45
Merge branch 'proxy' of github.com:yxssfxwzy/webmagic into yxssfxwzy-proxy
2014-06-09 23:51:36 +08:00