View Issue Details
| ID | Project | Category | View Status | Date Submitted | Last Update |
|---|---|---|---|---|---|
| 0000709 | unreal | ircd | public | 2003-02-05 01:26 | 2003-11-20 19:43 |
| Reporter | marviiin | Assigned To | syzop | ||
| Priority | none | Severity | minor | Reproducibility | sometimes |
| Status | closed | Resolution | fixed | ||
| Product Version | 3.2-beta14 | ||||
| Summary | 0000709: ident checking is not working correctly [waiting] | ||||
| Description | In connecting to a server with ~100 clients, only about 30% of the time do users running identds experience a correct ident query from the ircd. [UPDATE: please see new comments, its now a problem with 10% "hanging"] | ||||
| Steps To Reproduce | Take a few clients with known working identds, and reconnect to the server ~10 times. Of those connects, on average only 30% of them will have idented correctly. Some people will have a 90% success rate, others more like 5%. | ||||
| Additional Information | The problem seems to be a quirk with the timing(?) of nonblocking sockets (using O_NONBLOCK). The problem is that the WRITE_SOCK(cptr->authfd, authbuf, strlen(authbuf)) in send_authports() in s_auth.c fails, returning -1 and setting errno to P_EAGAIN. Rather than trying to write to the socket again at a later date, the ircd then promptly gives up by falling through to authsenderr:. I've worked around this on my ircd by checking errno for P_EAGAIN, and then returning straight out of send_authports() without touching the client's flags. Thus writing to the authfd socket gets deferred until next time the FDLIST is walked. After a few iterations the write() works, and all carries on as normal. The catch is of course that a few clients which are behind snarky firewalls cause write() to repeatedly error with P_EAGAIN. Whilst in debugmode the ircd times out rapidly with a DNS/AUTH timeout - in normal mode this seems to cause the client to hang indefinitely. So to get round that I added an int hammer; field to the Client structure - and after 4 hammers on the write() to authfd, the thing gracefully gives up and fails. I realise this is a horrible kludge, but I thought it might be of use in fixing the bug properly. Needless to say, the problem presented itself on a completely unaltered ircd 'out of the box'. The only quirk in the Config is choosing not to use the default linux threads (this caused huge CPU utilisation, instability, and the ircd refused to shutdown or die short of a kill -9). The relavent code for send_authports() in s_auth.c for the kludge described above is: int amountwritten; errno = 0; amountwritten = WRITE_SOCK(cptr->authfd, authbuf, strlen(authbuf)); if (amountwritten != strlen(authbuf)) { Debug((DEBUG_SEND, "failed absolutely miserably with WRITE_SOCK: amountwritten = %d, strlen(authbug) = %d. errno is %d", amountwritten, strlen(authbuf), errno)); if (ERRNO == P_EAGAIN && ++(cptr->hammer)<4) { return; } authsenderr: I can obviously provide this as a genuine patch if anyone's interested. | ||||
| 3rd party modules | |||||
|
|
Hehe I didn't even know it (sometimes) worked ;). |
|
|
Hm.. /* * send_authports * * Send the ident server a query giving "theirport , ourport". * The write is only attempted *once* so it is deemed to be a fail if the * entire write doesn't write all the data given. This shouldnt be a * problem since the socket should have a write buffer far greater than * this message to store it in should problems arise. -avalon */ AFAIK this comment is true, so I'm wondering why you are getting the EAGAIN error... interresting... |
|
|
I'll find out (maybe it didnt connect yet, or some other error) ;P |
|
|
The fix was even easier than I thought since timeout handling is already done in s_bsd.c ;P So it should be fixed in CVS now, can you confirm that? (I only have a LAN to test, which is "too fast" to reproduce this) [*corrected bad english*] edited on: 02-05 21:35 |
|
|
Assuming fixed in CVS. Re-report if this was not the case. |
|
|
quote from marviiin (mail): when putting it live on my server, about 10% of the users found themselves timing out (and being disconnected) at Check Ident after waiting for about a minute. (CONNECTTIMEOUT set to 20). So there's probably some other bug in the timeout code somewhere (??). |
|
|
Updated title/marked as new/etc, this is now a "different bug" so someone else can fix it too if he's bored. |
|
|
I have been able to reproduce the "can't connect at all because of the if (ERRNO == P_EAGAIN) return{} in send_authports" fix, by writing a little perl script which loads 150 clients onto the CVS test server, has them randomly 'chat' in a few channels and to each other, and then tries to connect from a machine which has something 'silly' listening on 113. I've been using a telnet server as quick hack to listen locally on the port and then just hang on the socket. With no clients connected, the DNS/AUTH part of the world seems to time out quite reasonably after around CONNECTTIMEOUT seconds - but with the reasonably loaded server, the more violent d/cing timeout occurs - i.e. anyone who's behind a router/firewall that grabs hold of incoming 113 and doesn't let them go is completely screwed. It seems that the DNS/AUTH timeout code in s_bsd.c is for some reason dependent on the load on the server - which sounds pretty dodgy, given that presumably it's meant to time out within a given specified time regardless. Meanwhile, my icky hack of "hammering 4 times on send_authports and then giving up" empirically seems to work pretty reliably :) |
|
|
Our whole auth system is a disaster: doing auth? no Sending [:dev.dyndns.org 001 k000-lh :Welcome to the HokNet IRC Network [email protected]] to k000-lh [..] doing auth? no Sending [:dev.dyndns.org 001 k001-an :Welcome to the HokNet IRC Network [email protected]] to k001-an but some time later: doing auth? yes Sending [:dev.dyndns.org 001 k013-qo :Welcome to the HokNet IRC Network [email protected]] to k013-qo [the doing auth thing is: Debug((DEBUG_NOTICE, "doing auth? %s", DoingAuth(cptr) ? "yes" : "no")); added just after the ch_cl access check in s_bsd.c] I don't want to work on this anymore, lol. |
|
|
This bug is a duplicate of #0000485 which shows the problem was already present in beta12... |
|
|
I tried to reproduce it here but without much success... Can you try to reproduce it with this: http://www.vulnscan.org/tmp/Unreal3.2-xdebug.tar.gz (XX debugging line stuff added). Ofcourse start with src/ircd -x 100 or something so it creates a nice big debug.log. If you succeeded in reproducing it, please send your debug.log ([g]zipped!! ;P) to [email protected] . Thanks ;). edited on: 02-09 23:20 |
|
|
Arathorn/marviiin[same]: can you _also_ (still do the debugging version stuff :P) paste your client class from your configfile? :) |
|
|
*hope you got some time this weekend*.. would be nice if this was fixed in beta15 [no there's no release schedule, but it's slowly getting time] ;). |
|
|
<~10 days before b15 (just a feeling :P). |
|
|
Oh and IF you are testing, can you also try to reproduce it at latest CVS? I've fixed 3 ident bugs since this xdebug version. |
|
|
> With 3 clients connected to a brand spanking new devel server, > the one that was sitting behind a NATting firewall blocking port 113 > had to wait 40 odd seconds to connect. Ok, I tried it, but I can't reproduce something bad: It waits ~30 (29, 30, 31) seconds and then just continues. I also tried connecting from the firewalled hosts while 200 other hosts were connecting, same results. Also, those 40 seconds you are talking about is controlled by CONNECTTIMEOUT, I _really_ recommend something like 20 unless you are running a polish server with routing problems every day... Maybe this will get configurable in configfile ;). So if you can reproduce something like waiting ~1m while your CONNECTTIMEOUT is set at 30 then something is really wrong and I'm interrested ;). |
|
|
Just for the record: this bug will not be fixed in beta15... Maybe in beta16, who knows :P. |
|
|
I wonder if this bug is fixed now. Dunnow if you are still alive ;p |
|
|
lol @ syzop |
|
|
Still alive and kicking - but ended up having to concentrate on other stuff; b15 didn't fix things; we (irc.theonering.net) turned off ident checking, turned off threading & scanning, and never looked back. Both "bad file descriptor" bugs disappeared (as per another bug report somewhere around here) - and the upgrade to b16 the other day went seamlessly. Given that ident checking is pretty useless nowadays, as you originally pointed out, and the fact that we've given up on using it, I simply haven't had time to fiddle around any further with it. I haven't tried it at all with b16 at all, but thanks for thinking of me ;) It's possible that at some point a few months down the road I might have time to waste on trying to duplicate the problems again with b16, but given the nightmare involved in reproducing them, and the extent to which they depended on the server being under 'typical' load (and not wanting to screw around with the poor bastards using the live server)... |
|
|
Ok, good ;). |
| Date Modified | Username | Field | Change |
|---|---|---|---|
| 2003-05-04 21:11 | syzop | Note Added: 0002646 | |
| 2003-05-04 21:33 | cyberCloWn | Note Added: 0002648 | |
| 2003-05-05 00:05 | marviiin | Note Added: 0002652 | |
| 2003-05-05 00:15 | syzop | Status | acknowledged => resolved |
| 2003-05-05 00:15 | syzop | Resolution | reopened => fixed |
| 2003-05-05 00:15 | syzop | Assigned To | => syzop |
| 2003-05-05 00:15 | syzop | Note Added: 0002654 | |
| 2003-11-20 19:43 | syzop | Status | resolved => closed |