Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Hah, you know, most crawlers were fine. The only one that actively DDOSed websites was fucking Yandex. It doesn't respect robots.txt and it will actively fight against any rate limits by spawning connections on new IPs the moment one is blocked


Search engine crawlers are a subset of crawlers. Sometimes you’re dealing with aggressive screen-scraping from competitors or various marketing tools.

I’ve had to deal with these things easily bringing sites down.


Same, and when more than half your links are dynamic search results, it can pile on and really bring things to a crawl. I worked on a fairly popular auto classifieds website, and more than 90% of traffic was various scrapers, and some were definitely a burden. Worse, is that it doesn't show up in analytics as it's not running client-js... Ironically equally bad was when bing started scraping with JS and it skewed google analytics.

If all we had to deal with were the users, wouldn't need nearly the spend on the site. Started manually blocking some of the worst offenders.


? I personally experienced many unbehaving crawlers/bots not respecting robots.txt and behaving like another user agent.

Feel free to prove me wrong and disrupt cloudflare by only handling that use-case


Found an old reference of mine what I did to block crawlers/bots

https://news.ycombinator.com/item?id=34101988

> Had a lot of spammers with Russian language. Implemented expanding xml-bombs, Google Captcha, hidden input fields and a couple of other things against bots. But the block on the russian language was most effective ( and since I was dogfooding it, I didn't see the harm at the time. But it's out of scope at this very moment, yes).




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: