Partial content web crawling using HTTP/2 and Go
Partial content web crawling using HTTP/2 and Go
Hi, I wrote a low-level HTTP/2 web crawler in Go, which can scrape partial content to save traffic. Tl;dr e.g. the HTML of a YouTube video contains the video description, views, likes etc. in its first 600KB, the remaining 900KB are of no use for me, but I have to pay my proxies by the gigabyte. My crawler receives packet per packet, and if I got everything I needed I reset the request, and only pay-for-what-i-crawled. This is also potentially useful for large-scale crawling operations, where duplicates matter. You could compute a simHash on the fly, and reset on-the-fly before crawling the entire document (again).
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Nuster – A HTTP/2 web cache accelerator
Redux-rx-http, a pragmatic HTTP layer for redux using redux-observable
Handled 6.87M HTTP Req/sec with Rapidoid on Top of Java NIO
Lint for HTTP
Zefner, Knight that will attack your HTTP's endpoint
Symfony HTTP Reponder: ADR Implemented
Of-Watchdog – OpenFaaS Watchdog for HTTP
Bhttp Binary HTTP (RFC 9292) for Go
Oblivious HTTP for Go
Bowtie, a (hopefully) idiomatic HTTP middleware for Go