Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #229244 > unrolled thread
| Started by | Gene Heskett <gheskett@shentel.net> |
|---|---|
| First post | 2020-12-03 14:00 +0100 |
| Last post | 2020-12-07 04:00 +0100 |
| Articles | 20 on this page of 42 — 20 participants |
Back to article view | Back to linux.debian.user
swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-03 14:00 +0100
Re: swamp rat bots Q john doe <johndoe65534@mail.com> - 2020-12-03 14:10 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-03 15:10 +0100
Re: swamp rat bots Q Greg Wooledge <wooledg@eeg.ccf.org> - 2020-12-03 15:20 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-03 20:00 +0100
Re: swamp rat bots Q Håkon Alstadheim <hakon@alstadheim.priv.no> - 2020-12-03 22:00 +0100
Re: swamp rat bots Q Keith Christian <keith1christian@gmail.com> - 2020-12-03 22:30 +0100
Re: swamp rat bots Q john doe <johndoe65534@mail.com> - 2020-12-03 16:40 +0100
Re: swamp rat bots Q Andy Smith <andy@strugglers.net> - 2020-12-04 03:20 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 06:10 +0100
Re: swamp rat bots Q Andy Smith <andy@strugglers.net> - 2020-12-04 07:10 +0100
Re: swamp rat bots Q john doe <johndoe65534@mail.com> - 2020-12-04 07:40 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 08:00 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 07:50 +0100
Re: swamp rat bots Q john doe <johndoe65534@mail.com> - 2020-12-04 08:20 +0100
Re: swamp rat bots Q Andy Smith <andy@strugglers.net> - 2020-12-04 08:50 +0100
Re: swamp rat bots Q "Jeremy Nicoll" <jn.ml.dbn.25@letterboxes.org> - 2020-12-04 11:50 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 14:50 +0100
Re: swamp rat bots Q Andrei POPESCU <andreimpopescu@gmail.com> - 2020-12-04 12:00 +0100
Re: swamp rat bots Q "hdv@gmail" <hdv.jadev@gmail.com> - 2020-12-04 10:00 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 14:50 +0100
Re: swamp rat bots Q Reco <recoverym4n@enotuniq.net> - 2020-12-04 18:40 +0100
Re: swamp rat bots Q grumpy@mailfence.com - 2020-12-04 19:10 +0100
Re: swamp rat bots Q Reco <recoverym4n@enotuniq.net> - 2020-12-04 19:20 +0100
Re: swamp rat bots Q Carl Fink <carlf@panix.com> - 2020-12-04 19:30 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 21:00 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 21:00 +0100
Re: swamp rat bots Q Tixy <tixy@yxit.co.uk> - 2020-12-04 22:20 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-04 23:10 +0100
Re: swamp rat bots Q Tixy <tixy@yxit.co.uk> - 2020-12-04 23:40 +0100
Re: swamp rat bots Q Gene Heskett <gheskett@shentel.net> - 2020-12-05 01:00 +0100
Re: swamp rat bots Q Andrei POPESCU <andreimpopescu@gmail.com> - 2020-12-05 11:00 +0100
Re: swamp rat bots Q elvis <elvis@dogonfire.com> - 2020-12-05 01:50 +0100
Re: swamp rat bots Q grumpy@mailfence.com - 2020-12-05 02:00 +0100
Web-bot tarpit aka spider trap (was: swamp rat bots Q) Nicolas George <george@nsup.org> - 2020-12-04 15:10 +0100
Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) Gene Heskett <gheskett@shentel.net> - 2020-12-04 17:10 +0100
Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) <tomas@tuxteam.de> - 2020-12-04 22:10 +0100
Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) "Martin McCormick" <martin.m@suddenlink.net> - 2020-12-06 20:20 +0100
Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) David Christensen <dpchrist@holgerdanske.com> - 2020-12-06 21:30 +0100
Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) Charles Curley <charlescurley@charlescurley.com> - 2020-12-06 23:10 +0100
Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) Gene Heskett <gheskett@shentel.net> - 2020-12-07 02:30 +0100
Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) Doug McGarrett <dmcgarrett@optonline.net> - 2020-12-07 04:00 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Gene Heskett <gheskett@shentel.net> |
|---|---|
| Date | 2020-12-04 14:50 +0100 |
| Message-ID | <BijDb-5RW-3@gated-at.bofh.it> |
| In reply to | #229285 |
On Friday 04 December 2020 03:49:39 hdv@gmail wrote: > On 2020-12-03 13:35, Gene Heskett wrote: > > I've had it with a certain bot that that ignore my robots.txt and > > proceeds to mirror my site, several times a day, burning up my > > upload bandwidth. They've moved it to 5 different addresses since > > midnight. > > > > I want to nail the door shut on the first attempted access by these > > AH's. > > > > Does anyone have a ready made script that can watch my httpd "other" > > log, and if a certain name is at the end of the line, grabs the ipv4 > > src address as arg3 of the line, and applies it to iptables DROP > > rules? > > > > Or do I have to invent a new wheel for this? > > > > Basic rules that simplify it somewhat. > > > > 1. this is ipv4 only country and not likely to change in the future > > decade. > > > > 2. the list of offending bot names will probably never go beyond 50, > > if that many. 5 would be realistic. > > > > 3. the src address in the log is at a fixed offset, obtainable with > > the bash MID$ but the dns return will need some acrobatics involving > > the bash RIGHT$ function. > > > > 4. it should track the number of hits, and after so many in a /24 > > block, autoswitch to a /16 block in order to keep the rules file > > from exploding. > > > > Any help will be much appreciated. PM's in this case welcome as I > > can't see broadcasting our armament against these MF'ers being > > broadcast on a public list. > > > > Thanks all. > > > > Cheers, Gene Heskett > > Let me offer you an alternative option. (Most) bots work by analysing > the referrals on each page in your website. Right? So, why not add a > link to a page that normal users will never visit (e.g. because they > do not see the link and thus will never click on it), but will show up > in a bot's analysis? That way you can monitor your logs for entries > containing that page. Every entity requesting that specific URL is > blocked. > Now that idea has some merit. Some of the bots are asking for stuff I've deleted years ago. And which cannot be obtained by links that now exist So I am wondering how they do it as they do not now exist. If I tracked and recorded those the list would be a long one. But I asked specifically how to enable it for one bot, and I've asked that question several times, getting smoke and mirror answers you all assume are helpfull, but which are useless to a new user installing the now 7 years old and long out of date package that in effect has no "how it works" docs. I asked 3 questions in a previous day or so timeline, and no one has actually attempted to actually answer even one of them. Here is one line from that log: and that I just blocked: coyote.coyote.den:80 192.99.6.226 - - [04/Dec/2020:07:18:20 -0500] "GET /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 673 "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; http://mj12bot.com/)" That named file does exist, but since its many years old & outdated, related to a computer that is itself nearly 38 years old, and is out of date as its been 4 or 5 years since I did an hg pull and rebuilt it. My own version of that machine has died of dried out electrolytic capacitors and no longer boots its baby unix os. I am a CET, have been since 1972 so I could fix it, but there comes a time when its time to let it go. I am now building metal carving machines run by LinuxCNC. > HTH > > HdV Cheers, Gene Heskett -- "There are four boxes to be used in defense of liberty: soap, ballot, jury, and ammo. Please use in that order." -Ed Howdershelt (Author) If we desire respect for the law, we must first make the law respectable. - Louis D. Brandeis Genes Web page <http://geneslinuxbox.net:6309/gene>
[toc] | [prev] | [next] | [standalone]
| From | Reco <recoverym4n@enotuniq.net> |
|---|---|
| Date | 2020-12-04 18:40 +0100 |
| Message-ID | <BindM-8au-15@gated-at.bofh.it> |
| In reply to | #229303 |
Hi. On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > But I asked specifically how to enable it for one bot, and I've asked > that question several times, getting smoke and mirror answers you all > assume are helpfull, but which are useless to a new user installing the > now 7 years old and long out of date package that in effect has no "how > it works" docs. I asked 3 questions in a previous day or so timeline, > and no one has actually attempted to actually answer even one of them. > Here is one line from that log: and that I just blocked: > > coyote.coyote.den:80 192.99.6.226 - - > [04/Dec/2020:07:18:20 -0500] "GET /gene/toolshed/c3/build/win32/prep/?C=S;O=D > HTTP/1.1" 200 673 "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; > http://mj12bot.com/)" Taken directly from the link. Bot Type Good crawler (always identifies itself) IP Range Distributed, Worldwide Obeys Robots.txt *Yes* Obeys Crawl Delay Yes Data served at Majestic.com I kindly suggest to all debian-user members to reflect on this, and to stop this pointless discussion. Reco
[toc] | [prev] | [next] | [standalone]
| From | grumpy@mailfence.com |
|---|---|
| Date | 2020-12-04 19:10 +0100 |
| Message-ID | <BinGO-9j-3@gated-at.bofh.it> |
| In reply to | #229320 |
On Fri, 4 Dec 2020, Reco wrote: > Hi. > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: >> But I asked specifically how to enable it for one bot, and I've asked >> that question several times, getting smoke and mirror answers you all >> assume are helpfull, but which are useless to a new user installing the >> now 7 years old and long out of date package that in effect has no "how >> it works" docs. I asked 3 questions in a previous day or so timeline, >> and no one has actually attempted to actually answer even one of them. >> Here is one line from that log: and that I just blocked: >> >> coyote.coyote.den:80 192.99.6.226 - - >> [04/Dec/2020:07:18:20 -0500] "GET /gene/toolshed/c3/build/win32/prep/?C=S;O=D >> HTTP/1.1" 200 673 "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; >> http://mj12bot.com/)" > > Taken directly from the link. > > Bot Type Good crawler (always identifies itself) > IP Range Distributed, Worldwide > Obeys Robots.txt *Yes* > Obeys Crawl Delay Yes > Data served at Majestic.com > > I kindly suggest to all debian-user members to reflect on this, and to > stop this pointless discussion. > > Reco > many times i have needed help and it takes a while for all the suggestions to coalesce i kindly suggest that your opinion is pointless
[toc] | [prev] | [next] | [standalone]
| From | Reco <recoverym4n@enotuniq.net> |
|---|---|
| Date | 2020-12-04 19:20 +0100 |
| Message-ID | <BinQt-cS-7@gated-at.bofh.it> |
| In reply to | #229321 |
On Fri, Dec 04, 2020 at 12:06:41PM -0600, grumpy@mailfence.com wrote: > On Fri, 4 Dec 2020, Reco wrote: > > > Hi. > > > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > > > But I asked specifically how to enable it for one bot, and I've asked > > > that question several times, getting smoke and mirror answers you all > > > assume are helpfull, but which are useless to a new user installing the > > > now 7 years old and long out of date package that in effect has no "how > > > it works" docs. I asked 3 questions in a previous day or so timeline, > > > and no one has actually attempted to actually answer even one of them. > > > Here is one line from that log: and that I just blocked: > > > > > > coyote.coyote.den:80 192.99.6.226 - - > > > [04/Dec/2020:07:18:20 -0500] "GET /gene/toolshed/c3/build/win32/prep/?C=S;O=D > > > HTTP/1.1" 200 673 "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; > > > http://mj12bot.com/)" > > > > Taken directly from the link. > > > > Bot Type Good crawler (always identifies itself) > > IP Range Distributed, Worldwide > > Obeys Robots.txt *Yes* > > Obeys Crawl Delay Yes > > Data served at Majestic.com > > > > I kindly suggest to all debian-user members to reflect on this, and to > > stop this pointless discussion. > > many times i have needed help and it takes a while for all the > suggestions to coalesce In this particular case it's third Gene's attempt (counting this year only and only at d-u) to solve this "problem". Suffice to say that previous attempts have not coalesced anything, and this one will not too. > i kindly suggest that your opinion is pointless Wow, I must've hit a nerve there. But I'm not offended, so prove me wrong. What's *your* suggestion for solving this problem? Reco
[toc] | [prev] | [next] | [standalone]
| From | Carl Fink <carlf@panix.com> |
|---|---|
| Date | 2020-12-04 19:30 +0100 |
| Message-ID | <Bio09-gD-1@gated-at.bofh.it> |
| In reply to | #229322 |
On 12/4/2020 1:16 PM, Reco wrote: > What's *your* suggestion for solving this problem? Speaking only for myself: maybe someone could point Gene to a new build of fail2ban? A howto for setting it up? A better tool for the job? Saying "Just read the man page" is literally the cliche unhelpful answer techies give to easy questions. -- Carl Fink
[toc] | [prev] | [next] | [standalone]
| From | Gene Heskett <gheskett@shentel.net> |
|---|---|
| Date | 2020-12-04 21:00 +0100 |
| Message-ID | <Bippg-13b-21@gated-at.bofh.it> |
| In reply to | #229322 |
On Friday 04 December 2020 13:16:23 Reco wrote: > On Fri, Dec 04, 2020 at 12:06:41PM -0600, grumpy@mailfence.com wrote: > > On Fri, 4 Dec 2020, Reco wrote: > > > Hi. > > > > > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > > > > But I asked specifically how to enable it for one bot, and I've > > > > asked that question several times, getting smoke and mirror > > > > answers you all assume are helpfull, but which are useless to a > > > > new user installing the now 7 years old and long out of date > > > > package that in effect has no "how it works" docs. I asked 3 > > > > questions in a previous day or so timeline, and no one has > > > > actually attempted to actually answer even one of them. Here is > > > > one line from that log: and that I just blocked: > > > > > > > > coyote.coyote.den:80 192.99.6.226 - - > > > > [04/Dec/2020:07:18:20 -0500] "GET > > > > /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 673 > > > > "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; > > > > http://mj12bot.com/)" > > > > > > Taken directly from the link. > > > > > > Bot Type Good crawler (always identifies itself) > > > IP Range Distributed, Worldwide > > > Obeys Robots.txt *Yes* > > > Obeys Crawl Delay Yes > > > Data served at Majestic.com > > > > > > I kindly suggest to all debian-user members to reflect on this, > > > and to stop this pointless discussion. > > > > many times i have needed help and it takes a while for all the > > suggestions to coalesce > > In this particular case it's third Gene's attempt (counting this year > only and only at d-u) to solve this "problem". > Suffice to say that previous attempts have not coalesced anything, and > this one will not too. > > > i kindly suggest that your opinion is pointless > > Wow, I must've hit a nerve there. > But I'm not offended, so prove me wrong. > > What's *your* suggestion for solving this problem? > > Reco How about starting by answering the 3 questions I have posted instead of calling this a square dance in round dance territory? Cheers, Gene Heskett -- "There are four boxes to be used in defense of liberty: soap, ballot, jury, and ammo. Please use in that order." -Ed Howdershelt (Author) If we desire respect for the law, we must first make the law respectable. - Louis D. Brandeis Genes Web page <http://geneslinuxbox.net:6309/gene>
[toc] | [prev] | [next] | [standalone]
| From | Gene Heskett <gheskett@shentel.net> |
|---|---|
| Date | 2020-12-04 21:00 +0100 |
| Message-ID | <Bippf-13b-5@gated-at.bofh.it> |
| In reply to | #229320 |
On Friday 04 December 2020 12:39:24 Reco wrote: > Hi. > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > > But I asked specifically how to enable it for one bot, and I've > > asked that question several times, getting smoke and mirror answers > > you all assume are helpfull, but which are useless to a new user > > installing the now 7 years old and long out of date package that in > > effect has no "how it works" docs. I asked 3 questions in a previous > > day or so timeline, and no one has actually attempted to actually > > answer even one of them. Here is one line from that log: and that I > > just blocked: > > > > coyote.coyote.den:80 192.99.6.226 - - > > [04/Dec/2020:07:18:20 -0500] "GET > > /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 673 "-" > > "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; http://mj12bot.com/)" > > Taken directly from the link. > > Bot Type Good crawler (always identifies itself) > IP Range Distributed, Worldwide > Obeys Robots.txt *Yes* Sorry, they do not, they've read it and ignored it 428 times in the life of that log which I zeroed out around 1 July of this year. They've also used up my upload bandwidth 37760 times in the life of that log. I have a 10 megabit service, means 2.5 going up. That hogging of my upload bandwidth can only be defined as a DDOS attack. Of the 112 DROP rules I currently have, at least 80 were generated by their activities against my site. > Obeys Crawl Delay Yes What the heck is that? > Data served at Majestic.com > > I kindly suggest to all debian-user members to reflect on this, and to > stop this pointless discussion. > > Reco At this point it sounds like you are defending them, but until they read and obey robots.txt, I will continue to block them the instant I can ID that its them. Cheers, Gene Heskett -- "There are four boxes to be used in defense of liberty: soap, ballot, jury, and ammo. Please use in that order." -Ed Howdershelt (Author) If we desire respect for the law, we must first make the law respectable. - Louis D. Brandeis Genes Web page <http://geneslinuxbox.net:6309/gene>
[toc] | [prev] | [next] | [standalone]
| From | Tixy <tixy@yxit.co.uk> |
|---|---|
| Date | 2020-12-04 22:20 +0100 |
| Message-ID | <BiqEG-24y-17@gated-at.bofh.it> |
| In reply to | #229327 |
On Fri, 2020-12-04 at 14:51 -0500, Gene Heskett wrote: > On Friday 04 December 2020 12:39:24 Reco wrote: > > > Hi. > > > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > > > But I asked specifically how to enable it for one bot, and I've > > > asked that question several times, getting smoke and mirror answers > > > you all assume are helpfull, but which are useless to a new user > > > installing the now 7 years old and long out of date package that in > > > effect has no "how it works" docs. I asked 3 questions in a previous > > > day or so timeline, and no one has actually attempted to actually > > > answer even one of them. Here is one line from that log: and that I > > > just blocked: > > > > > > coyote.coyote.den:80 192.99.6.226 - - > > > [04/Dec/2020:07:18:20 -0500] "GET > > > /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 673 "-" > > > "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; http://mj12bot.com/)" > > > > Taken directly from the link. > > > > Bot Type Good crawler (always identifies itself) > > IP Range Distributed, Worldwide > > Obeys Robots.txt *Yes* > > Sorry, they do not, they've read it and ignored it 428 times in the life > of that log which I zeroed out around 1 July of this year. Why would they read it if they we're going to just ignore it, perhaps your robots.txt is broken? Hint, it is, in 2 or 3 different ways I can see (if it's http://geneslinuxbox.net:6309/robots.txt we're talking about). That file doesn't have any syntactically correct entry in there for blocking that bot. I don't know why you seem set on blaming malice on part of a bot whose front web page has sections like: How can I block MJ12bot? How can I slow down MJ12bot? What commands in robots.txt does MJ12bot support? Why did my robots.txt block not work on MJ12bot? The URL for that page is in the user-agent string from the log snippet you posted above. -- Tixy
[toc] | [prev] | [next] | [standalone]
| From | Gene Heskett <gheskett@shentel.net> |
|---|---|
| Date | 2020-12-04 23:10 +0100 |
| Message-ID | <Birr4-2Qp-23@gated-at.bofh.it> |
| In reply to | #229334 |
On Friday 04 December 2020 16:14:29 Tixy wrote: > On Fri, 2020-12-04 at 14:51 -0500, Gene Heskett wrote: > > On Friday 04 December 2020 12:39:24 Reco wrote: > > > Hi. > > > > > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > > > > But I asked specifically how to enable it for one bot, and I've > > > > asked that question several times, getting smoke and mirror > > > > answers you all assume are helpfull, but which are useless to a > > > > new user installing the now 7 years old and long out of date > > > > package that in effect has no "how it works" docs. I asked 3 > > > > questions in a previous day or so timeline, and no one has > > > > actually attempted to actually answer even one of them. Here is > > > > one line from that log: and that I just blocked: > > > > > > > > coyote.coyote.den:80 192.99.6.226 - - > > > > [04/Dec/2020:07:18:20 -0500] "GET > > > > /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 673 > > > > "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; > > > > http://mj12bot.com/)" > > > > > > Taken directly from the link. > > > > > > Bot Type Good crawler (always identifies itself) > > > IP Range Distributed, Worldwide > > > Obeys Robots.txt *Yes* > > > > Sorry, they do not, they've read it and ignored it 428 times in the > > life of that log which I zeroed out around 1 July of this year. > > Why would they read it if they we're going to just ignore it, perhaps > your robots.txt is broken? Hint, it is, in 2 or 3 different ways I can > see (if it's http://geneslinuxbox.net:6309/robots.txt we're talking > about). That file doesn't have any syntactically correct entry in > there for blocking that bot. And what might that be like, I'll fix it right now > I don't know why you seem set on blaming malice on part of a bot whose > front web page has sections like: The evidence I have collected so far indicates they don't care who they ddos with their repeated suckage. So I never considered allowing their site anything like direct access by going to it. Their actions speak MUCH louder than the words below. > How can I block MJ12bot? > How can I slow down MJ12bot? > What commands in robots.txt does MJ12bot support? > Why did my robots.txt block not work on MJ12bot? > > The URL for that page is in the user-agent string from the log snippet > you posted above. Cheers, Gene Heskett -- "There are four boxes to be used in defense of liberty: soap, ballot, jury, and ammo. Please use in that order." -Ed Howdershelt (Author) If we desire respect for the law, we must first make the law respectable. - Louis D. Brandeis Genes Web page <http://geneslinuxbox.net:6309/gene>
[toc] | [prev] | [next] | [standalone]
| From | Tixy <tixy@yxit.co.uk> |
|---|---|
| Date | 2020-12-04 23:40 +0100 |
| Message-ID | <BirU5-35U-3@gated-at.bofh.it> |
| In reply to | #229335 |
On Fri, 2020-12-04 at 17:06 -0500, Gene Heskett wrote: > On Friday 04 December 2020 16:14:29 Tixy wrote: > > > On Fri, 2020-12-04 at 14:51 -0500, Gene Heskett wrote: > > > On Friday 04 December 2020 12:39:24 Reco wrote: > > > > Hi. > > > > > > > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > > > > > But I asked specifically how to enable it for one bot, and > I've > > > > > asked that question several times, getting smoke and mirror > > > > > answers you all assume are helpfull, but which are useless to > a > > > > > new user installing the now 7 years old and long out of date > > > > > package that in effect has no "how it works" docs. I asked 3 > > > > > questions in a previous day or so timeline, and no one has > > > > > actually attempted to actually answer even one of them. Here > is > > > > > one line from that log: and that I just blocked: > > > > > > > > > > coyote.coyote.den:80 192.99.6.226 - - > > > > > [04/Dec/2020:07:18:20 -0500] "GET > > > > > /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 673 > > > > > "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; > > > > > http://mj12bot.com/)" > > > > > > > > Taken directly from the link. > > > > > > > > Bot Type Good crawler (always identifies itself) > > > > IP Range Distributed, Worldwide > > > > Obeys Robots.txt *Yes* > > > > > > Sorry, they do not, they've read it and ignored it 428 times in > the > > > life of that log which I zeroed out around 1 July of this year. > > > > Why would they read it if they we're going to just ignore it, > perhaps > > your robots.txt is broken? Hint, it is, in 2 or 3 different ways I > can > > see (if it's http://geneslinuxbox.net:6309/robots.txt we're talking > > about). That file doesn't have any syntactically correct entry in > > there for blocking that bot. > > And what might that be like, I'll fix it right now OK, I'll do your proofreading... At the end of the robots.txt you are missing a colon from a rule that disallows everything for all bots... User-agent * Disallow: / That should be: User-agent: * Disallow: / But if you just want to disable the bot you reckon is a problem, the front page of their site (https://mj12bot.com/) says you want: User-agent: MJ12bot Disallow: / Or you could read their page to see the robots.txt syntax for slowing down crawling, which I assume is relevant to other bots to you may have problems with. The other rules above your disallow everything (which are superfluous if you keep that final rule) also have typos, you have a '0' here... User-0agent: * Disallow: /doc/ And this rule has a space in the URL... User-agent: * Disallow: stress test I'm pretty sure URLs can't have actual space characters in them and that must be a typo on your behalf. Also something I read when looking at this issue a few hours ago (but can't find again) reckoned that Google's bot let you have multiple statements on a line separated by spaces, e.g. Disallow: foo Disallow: bar So it seems likely that having a space in the URL like you have isn't legal, and could possibly upset parsing. -- Tixy
[toc] | [prev] | [next] | [standalone]
| From | Gene Heskett <gheskett@shentel.net> |
|---|---|
| Date | 2020-12-05 01:00 +0100 |
| Message-ID | <Bit9w-3Uu-5@gated-at.bofh.it> |
| In reply to | #229336 |
On Friday 04 December 2020 17:37:02 Tixy wrote: > On Fri, 2020-12-04 at 17:06 -0500, Gene Heskett wrote: > > On Friday 04 December 2020 16:14:29 Tixy wrote: > > > On Fri, 2020-12-04 at 14:51 -0500, Gene Heskett wrote: > > > > On Friday 04 December 2020 12:39:24 Reco wrote: > > > > > Hi. > > > > > > > > > > On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: > > > > > > But I asked specifically how to enable it for one bot, and > > > > I've > > > > > > > > asked that question several times, getting smoke and mirror > > > > > > answers you all assume are helpfull, but which are useless > > > > > > to > > > > a > > > > > > > > new user installing the now 7 years old and long out of date > > > > > > package that in effect has no "how it works" docs. I asked 3 > > > > > > questions in a previous day or so timeline, and no one has > > > > > > actually attempted to actually answer even one of them. Here > > > > is > > > > > > > > one line from that log: and that I just blocked: > > > > > > > > > > > > coyote.coyote.den:80 192.99.6.226 - - > > > > > > [04/Dec/2020:07:18:20 -0500] "GET > > > > > > /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 > > > > > > 673 "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; > > > > > > http://mj12bot.com/)" > > > > > > > > > > Taken directly from the link. > > > > > > > > > > Bot Type Good crawler (always identifies itself) > > > > > IP Range Distributed, Worldwide > > > > > Obeys Robots.txt *Yes* > > > > > > > > Sorry, they do not, they've read it and ignored it 428 times in > > > > the > > > > > > life of that log which I zeroed out around 1 July of this year. > > > > > > Why would they read it if they we're going to just ignore it, > > > > perhaps > > > > > your robots.txt is broken? Hint, it is, in 2 or 3 different ways I > > > > can > > > > > see (if it's http://geneslinuxbox.net:6309/robots.txt we're > > > talking about). That file doesn't have any syntactically correct > > > entry in there for blocking that bot. > > > > And what might that be like, I'll fix it right now > > OK, I'll do your proofreading... > > At the end of the robots.txt you are missing a colon from a rule that > disallows everything for all bots... > > User-agent * > Disallow: / > > That should be: > > User-agent: * > Disallow: / > > But if you just want to disable the bot you reckon is a problem, the > front page of their site (https://mj12bot.com/) says you want: > > User-agent: MJ12bot > Disallow: / > > Or you could read their page to see the robots.txt syntax for slowing > down crawling, which I assume is relevant to other bots to you may > have problems with. > > The other rules above your disallow everything (which are superfluous > if you keep that final rule) also have typos, you have a '0' here... > > User-0agent: * > Disallow: /doc/ > > And this rule has a space in the URL... > > User-agent: * > Disallow: stress test > > I'm pretty sure URLs can't have actual space characters in them and > that must be a typo on your behalf. Also something I read when looking > at this issue a few hours ago (but can't find again) reckoned that > Google's bot let you have multiple statements on a line separated by > spaces, e.g. > Fat fingers syndrome, I've suffered from that for 86 years. Short fat fingers that are fond of pressing 2 keys at once. Thanks for pointing it out nicely. The unreal part is that I have made an excellent living since I was about 14 and quit school to go fix them new-fangled things called Televisions in '48 when the first tv station came on the air in central Iowa. I wound up as the Chief Engineer at a string of tv stations from the early '70's on. But now I'm 86, eating my own cooking & trying to keep a 30% pump with some replacement parts running well enough to wake up the next morning. And building my own CNC machinery to keep me busy and out of the bars. Not much I haven't tried since. My fingerprints have been to 37,000 feet deep in the mohole as I helped build the tv cameras that were on the Navy's Trieste in Feb 1960. They say water isn't compressible, but when the outside pressure is nearly 18,000 psi, it sure is. > Disallow: foo Disallow: bar > > So it seems likely that having a space in the URL like you have isn't > legal, and could possibly upset parsing. Anyway, I've fixed what you pointed out, and am watching the log. Thank you, a lot. Stay safe and well, Tixy. Cheers, Gene Heskett -- "There are four boxes to be used in defense of liberty: soap, ballot, jury, and ammo. Please use in that order." -Ed Howdershelt (Author) If we desire respect for the law, we must first make the law respectable. - Louis D. Brandeis Genes Web page <http://geneslinuxbox.net:6309/gene>
[toc] | [prev] | [next] | [standalone]
| From | Andrei POPESCU <andreimpopescu@gmail.com> |
|---|---|
| Date | 2020-12-05 11:00 +0100 |
| Message-ID | <BiCw9-276-1@gated-at.bofh.it> |
| In reply to | #229339 |
[Multipart message — attachments visible in raw view] — view raw
On Vi, 04 dec 20, 18:55:22, Gene Heskett wrote: > On Friday 04 December 2020 17:37:02 Tixy wrote: > > > > OK, I'll do your proofreading... [...] > Fat fingers syndrome, I've suffered from that for 86 years. Typing errors happen to everyone. Triple-checking the result when something is not working as expected usually catches most, but possibly not all of them. When posting a question to the list it would really save a lot of time for all involved if one would just attach[1] the config files in question. Most of them should be accepted by the list without issues. For larger files (e.g. logs or similar) there is always gzip. [1] copy-paste or similar can hide issues like wrong line-endings. Kind regards, Andrei -- http://wiki.debian.org/FAQsFromDebianUser
[toc] | [prev] | [next] | [standalone]
| From | elvis <elvis@dogonfire.com> |
|---|---|
| Date | 2020-12-05 01:50 +0100 |
| Message-ID | <BitVU-4AA-3@gated-at.bofh.it> |
| In reply to | #229336 |
Just goes to show if you whinge hard enough and pretend to be useless it will irritate someone enough to do your work. You don't even bother to check your robots.txt and complained about some evil bot all year. Priceless. On 5/12/20 8:37 am, Tixy wrote: > On Fri, 2020-12-04 at 17:06 -0500, Gene Heskett wrote: >> On Friday 04 December 2020 16:14:29 Tixy wrote: >> >>> On Fri, 2020-12-04 at 14:51 -0500, Gene Heskett wrote: >>>> On Friday 04 December 2020 12:39:24 Reco wrote: >>>>> Hi. >>>>> >>>>> On Fri, Dec 04, 2020 at 08:39:42AM -0500, Gene Heskett wrote: >>>>>> But I asked specifically how to enable it for one bot, and >> I've >>>>>> asked that question several times, getting smoke and mirror >>>>>> answers you all assume are helpfull, but which are useless to >> a >>>>>> new user installing the now 7 years old and long out of date >>>>>> package that in effect has no "how it works" docs. I asked 3 >>>>>> questions in a previous day or so timeline, and no one has >>>>>> actually attempted to actually answer even one of them. Here >> is >>>>>> one line from that log: and that I just blocked: >>>>>> >>>>>> coyote.coyote.den:80 192.99.6.226 - - >>>>>> [04/Dec/2020:07:18:20 -0500] "GET >>>>>> /gene/toolshed/c3/build/win32/prep/?C=S;O=D HTTP/1.1" 200 673 >>>>>> "-" "Mozilla/5.0 (compatible; MJ12bot/v1.4.8; >>>>>> http://mj12bot.com/)" >>>>> Taken directly from the link. >>>>> >>>>> Bot Type Good crawler (always identifies itself) >>>>> IP Range Distributed, Worldwide >>>>> Obeys Robots.txt *Yes* >>>> Sorry, they do not, they've read it and ignored it 428 times in >> the >>>> life of that log which I zeroed out around 1 July of this year. >>> Why would they read it if they we're going to just ignore it, >> perhaps >>> your robots.txt is broken? Hint, it is, in 2 or 3 different ways I >> can >>> see (if it's http://geneslinuxbox.net:6309/robots.txt we're talking >>> about). That file doesn't have any syntactically correct entry in >>> there for blocking that bot. >> And what might that be like, I'll fix it right now > OK, I'll do your proofreading... > > At the end of the robots.txt you are missing a colon from a rule that > disallows everything for all bots... > > User-agent * > Disallow: / > > That should be: > > User-agent: * > Disallow: / > > But if you just want to disable the bot you reckon is a problem, the > front page of their site (https://mj12bot.com/) says you want: > > User-agent: MJ12bot > Disallow: / > > Or you could read their page to see the robots.txt syntax for slowing > down crawling, which I assume is relevant to other bots to you may have > problems with. > > The other rules above your disallow everything (which are superfluous > if you keep that final rule) also have typos, you have a '0' here... > > User-0agent: * > Disallow: /doc/ > > And this rule has a space in the URL... > > User-agent: * > Disallow: stress test > > I'm pretty sure URLs can't have actual space characters in them and > that must be a typo on your behalf. Also something I read when looking > at this issue a few hours ago (but can't find again) reckoned that > Google's bot let you have multiple statements on a line separated by > spaces, e.g. > > Disallow: foo Disallow: bar > > So it seems likely that having a space in the URL like you have isn't > legal, and could possibly upset parsing. > -- .....I'VE GOT BLISTERS ON MY FINGERS!.....
[toc] | [prev] | [next] | [standalone]
| From | grumpy@mailfence.com |
|---|---|
| Date | 2020-12-05 02:00 +0100 |
| Message-ID | <Biu5z-4E5-1@gated-at.bofh.it> |
| In reply to | #229340 |
On Sat, 5 Dec 2020, elvis wrote: > Just goes to show if you whinge hard enough and pretend to be useless it will > irritate someone enough to do your work. > > > You don't even bother to check your robots.txt and complained about some evil > bot all year. Priceless. now now little girl don't get your panties bunched
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2020-12-04 15:10 +0100 |
| Subject | Web-bot tarpit aka spider trap (was: swamp rat bots Q) |
| Message-ID | <BijWx-6eq-3@gated-at.bofh.it> |
| In reply to | #229285 |
[Multipart message — attachments visible in raw view] — view raw
hdv@gmail (12020-12-04): > Let me offer you an alternative option. (Most) bots work by analysing the > referrals on each page in your website. Right? So, why not add a link to a > page that normal users will never visit (e.g. because they do not see the > link and thus will never click on it), but will show up in a bot's analysis? > That way you can monitor your logs for entries containing that page. Every > entity requesting that specific URL is blocked. This made me think of something. A long time ago, a friend of mine implemented, to trap the badly-behaved robots, something called the Book of Infinity: a set of deterministic pseud-random pages linking to sub-pages ad infinitum, with ever growing URLs. As it happened, it was not actually a good idea, and released a lot of CO₂, and the very badly behaved robots had to be blacklisted from explring it. (At some point, we had the same problem when Googlebot tried to brute-force our online make-your-own-adventure book, but Googlebot heeds robots.txt.) But it could be coupled with techniques inspired by spam tarpits: have the server reply at a crawl to force the bots to waste resources, while keeping the resource consumption on the server strictly bounded. Oh, I just noticed I was not the first one to think of it: Wikipedia tells me it's called a spider trap. https://en.wikipedia.org/wiki/Spider_trap Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | Gene Heskett <gheskett@shentel.net> |
|---|---|
| Date | 2020-12-04 17:10 +0100 |
| Subject | Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) |
| Message-ID | <BilOF-7oP-5@gated-at.bofh.it> |
| In reply to | #229307 |
On Friday 04 December 2020 09:00:05 Nicolas George wrote: > hdv@gmail (12020-12-04): > > Let me offer you an alternative option. (Most) bots work by > > analysing the referrals on each page in your website. Right? So, why > > not add a link to a page that normal users will never visit (e.g. > > because they do not see the link and thus will never click on it), > > but will show up in a bot's analysis? That way you can monitor your > > logs for entries containing that page. Every entity requesting that > > specific URL is blocked. > > This made me think of something. > > A long time ago, a friend of mine implemented, to trap the > badly-behaved robots, something called the Book of Infinity: a set of > deterministic pseud-random pages linking to sub-pages ad infinitum, > with ever growing URLs. > > As it happened, it was not actually a good idea, and released a lot of > CO₂, and the very badly behaved robots had to be blacklisted from > explring it. (At some point, we had the same problem when Googlebot > tried to brute-force our online make-your-own-adventure book, but > Googlebot heeds robots.txt.) > > But it could be coupled with techniques inspired by spam tarpits: have > the server reply at a crawl to force the bots to waste resources, > while keeping the resource consumption on the server strictly bounded. > > Oh, I just noticed I was not the first one to think of it: Wikipedia > tells me it's called a spider trap. > > https://en.wikipedia.org/wiki/Spider_trap > Sounds like a good idea, I'll have to think about it, feed the bots in 256 byte pieces every 5 seconds to keep them from timing out, with 256 bytes from rnd mixed in to make a dos packet? :) Just be sure the crc is good. ;-) > Regards, Take care now, Nicolas. Cheers, Gene Heskett -- "There are four boxes to be used in defense of liberty: soap, ballot, jury, and ammo. Please use in that order." -Ed Howdershelt (Author) If we desire respect for the law, we must first make the law respectable. - Louis D. Brandeis Genes Web page <http://geneslinuxbox.net:6309/gene>
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2020-12-04 22:10 +0100 |
| Subject | Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) |
| Message-ID | <Biqv1-1Yc-37@gated-at.bofh.it> |
| In reply to | #229315 |
[Multipart message — attachments visible in raw view] — view raw
On Fri, Dec 04, 2020 at 11:02:48AM -0500, Gene Heskett wrote: [...] > Sounds like a good idea, I'll have to think about it, feed the bots in > 256 byte pieces every 5 seconds to keep them from timing out, with 256 > bytes from rnd mixed in to make a dos packet? :) Just be sure the crc is > good. ;-) There used to be a firewall thingmajig doing tarpit. Ah, nftables also has an addon for that. That said, it's eating resources on your side too, and chances are that almost every resource, from CPU power to electrical power is cheaper on the other side. The best strategy, therefore, seems to be DROP. This, at least, lets the other side wondering whether an answer is coming for as long as their timeout is -- and even unsure about whether their victim is there at all. Revenge may taste enticing, but isn't always the wisest adviser. Cheers - t
[toc] | [prev] | [next] | [standalone]
| From | "Martin McCormick" <martin.m@suddenlink.net> |
|---|---|
| Date | 2020-12-06 20:20 +0100 |
| Subject | Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) |
| Message-ID | <Bj7JE-6ds-11@gated-at.bofh.it> |
| In reply to | #229332 |
It's nice to see that I am not the only sick puppy out there. At 69 years old, I still don't have much trouble with getting in touch with my inner 12-year-old when it comes to intrusive marketing which is so prevalent these days. I found our old dial-up modem in a box of odds and ends 2 years ago and wondered if it could read callerID tones sent after the first ring. It can so I started on a perl program that initializes the modem for callerID and then compares the strings received with a pair of files, one of which is called scum and contains callerID name packets of folks we don't want to talk to. The other is called good and looks for names of friends or anyone else we like hearing from. It is actually scanned first and causes the program to abort. If anyone's name matches a name in scum, or the caller's ID appears to be blocked or, in one case, matches a whole exchange, (first 3 digits after the area code), I call a subroutine that makes the modem answer for half a second then drops the call. We were bombarded with garbage calls all day long until the US presidential election and it was so satisfying to hear the program kill the call by answering just as the second ring began. I also installed subroutines that looked for the word "SPAM?" just before the name or "ROBO?" also just before the name. We now get very few unwanted calls but occasionally, we'll get a call from lala-land from someone we don't know who lets the phone ring until the answering machine picks up and then fails to leave a message and I tell my wife, "I'll go put them on the juke box." especially if it looks like the name of a business with which we have no relationship. Our phone is very quiet these days except for legitimate calls. Martin McCormick
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2020-12-06 21:30 +0100 |
| Subject | Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) |
| Message-ID | <Bj8Pn-6Pm-1@gated-at.bofh.it> |
| In reply to | #229395 |
On 12/6/20 11:15 AM, Martin McCormick wrote: > It's nice to see that I am not the only sick puppy out there. > > At 69 years old, I still don't have much trouble with > getting in touch with my inner 12-year-old when it comes to > intrusive marketing which is so prevalent these days. > > I found our old dial-up modem in a box of odds and ends 2 > years ago and wondered if it could read callerID tones sent after > the first ring. It can so I started on a perl program that > initializes the modem for callerID and then compares the strings > received with a pair of files, one of which is called scum and > contains callerID name packets of folks we don't want to talk to. > The other is called good and looks for names of friends or anyone > else we like hearing from. It is actually scanned first and > causes the program to abort. > > If anyone's name matches a name in scum, or the caller's > ID appears to be blocked or, in one case, matches a whole > exchange, (first 3 digits after the area code), I call a > subroutine that makes the modem answer for half a second then > drops the call. > > We were bombarded with garbage calls all day long until > the US presidential election and it was so satisfying to hear the > program kill the call by answering just as the second ring began. > > I also installed subroutines that looked for the word > "SPAM?" just before the name or "ROBO?" also just before the > name. > > We now get very few unwanted calls but occasionally, > we'll get a call from lala-land from someone we don't know who > lets the phone ring until the answering machine picks up and then > fails to leave a message and I tell my wife, "I'll go put them on > the juke box." especially if it looks like the name of a business > with which we have no relationship. > > Our phone is very quiet these days except for legitimate > calls. > > Martin McCormick https://www.asterisk.org/ https://crosstalksolutions.com/howto-pwn-telemarketers-with-lenny/ https://www.youtube.com/watch?v=RRhRImp6kKQ David
[toc] | [prev] | [next] | [standalone]
| From | Charles Curley <charlescurley@charlescurley.com> |
|---|---|
| Date | 2020-12-06 23:10 +0100 |
| Subject | Re: Web-bot tarpit aka spider trap (was: swamp rat bots Q) |
| Message-ID | <Bjao9-7SN-3@gated-at.bofh.it> |
| In reply to | #229395 |
On Sun, 06 Dec 2020 13:15:50 -0600 "Martin McCormick" <martin.m@suddenlink.net> wrote: > I found our old dial-up modem in a box of odds and ends 2 > years ago and wondered if it could read callerID tones sent after > the first ring. It can so I started on a perl program that > initializes the modem for callerID and then compares the strings > received with a pair of files, one of which is called scum and > contains callerID name packets of folks we don't want to talk to. > The other is called good and looks for names of friends or anyone > else we like hearing from. It is actually scanned first and > causes the program to abort. You wouldn't care to make this software available, would you? Consider adding an option to send a fax tone instead. :-) -- Does anybody read signatures any more? https://charlescurley.com https://charlescurley.com/blog/
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.debian.user
csiph-web