View Issue Details

IDProjectCategoryView StatusLast Update
0002075unrealircdpublic2004-12-11 21:21
Reporterbrain2 Assigned To 
PrioritynormalSeverityblockReproducibilityalways
Status closedResolutionopen 
Product Version3.2.1 
Summary0002075: Disconnect of large amount of clients blocks ircd
DescriptionDuring testing we connected 1500 clients to our ircd (as is standard part of our trial process before the first trial link), as this server is much more powerful than most of our other boxes (2.4ghz AMD Athlon XP, 1024mb ram, 100mb/sec burstable connection) we decided to connect more clients than usual to it. After connecting 1550 clients we disconnected them all at once, at which point the ircd blocked for 6 minutes.
Steps To ReproduceConnect 1550+ clients to a machine of equal or less spec, disconnect a large number of them.
Additional InformationI've seen other ircds deal with this without blocking, perhaps threading might be the answer here?
3rd party modulesN/A

Activities

syzop

2004-09-16 23:29

administrator   ~0007668

[no@threading]
6 minutes isn't really normal, so I don't know what went wrong there.
Personally, I haven't really had such problems when doing clonetests (2K, 5K or even 10K).
Just checked disconnecting 2000 clones just to be sure (all in 1 channel == huge amount of traffic)... took 25 seconds, and that's on a P450 (Linux).

syzop

2004-09-16 23:32

administrator   ~0007669

(oh and 4 seconds if not in channel btw)

brain2

2004-09-16 23:43

reporter   ~0007670

1550 clients in the same channel quitting by closing their connections without a QUIT line.

also four non-opers were disconnected with 'software caused connection abort' before the actual disconnects began (we never actually saw the quitmessages, we just came back 6 minutes later when it was responding, to find all the clients disconnected as expected)

On the other hand there's no problem connecting them, even rapidly ;)

syzop

2004-09-16 23:54

administrator   ~0007671

> 1550 clients in the same channel quitting by closing their connections without a QUIT line.

That's what I did indeed, well.. with 2000 clones ;p. (and I did see the clones quiting.. although I forgot their quit reason I assume it was reset by peer)

I must admit I always put the connect snomask off (/mode nick +s -cF), because if I get flooded too much my mirc often disconnects.. I've seen that before with 10K clone tests, so.. blah ;).

Which FreeBSD version are you using btw?
Personally I had quite some problems getting freebsd with many fds (3K.. 5K..) running (granted, that was under vmware w/384M @1.533GHz, but.. Linux can do fine with 10K ;p). I guess it involves quite some tweaking.
I'm pretty sure there are quite some xK servers out there running FreeBSD.
Like, I know of a few 3-4K Linux servers, so...

brain2

2004-09-17 20:56

reporter   ~0007681

freebsd 5.2.1 as in system specs, on i386, recompiled kernel to support quotas and ipfw, with scsi support and ipv6 removed (unneeded)... a lot of the large servers on the top 10 nets use freebsd as their base, however they probably don't try to disconnect so many clients in one go :)

the ircd handles fine while theyre there, even if theyre all spamming random output from 'fortune', its just when they go thats the issue

al5001

2004-09-17 21:37

reporter   ~0007683

Last edited: 2004-09-17 21:40

Did you set up FreeBSD's sysctl properly? By default, kern.maxfilesperproc is 957 on 5.2.1... Which limits your IRCd quite a bit... you'll want to change it to 10000. You'll also want to change kern.maxproc to 10000, and kern.maxfiles to 10000. Good luck.

edit: maxproc probably isn't very important, but that's what I use anyway.

edited on: 2004-09-17 21:40

brain2

2004-09-17 22:41

reporter   ~0007690

kern.maxfiles is 8008 which leaves us around 7200 which is pleanty.

maxfilesperproc isnt 975, as obviously we managed to get the 1500 clients on there and disconnect them, it was only when we told them to disconnect that the issue occured...

brain2

2004-09-17 22:43

reporter   ~0007691

> Oh btw, what exactly is happening during those 6 minutes...
> CPU spike? RAM spike? Both?

dont know... couldnt do anything, machine deadlocked. (no we havent configured logon classes next, thats next on our agenda to limit the max CPU hours and percentile per process)

syzop

2004-09-17 23:21

administrator   ~0007692

I did some more tests, and I now actually turned my DEBUGMODE off (a major factor that I forgot to think about ;p) and compared it a bit with hybrid.. just for fun :).
Clients connect just as fast as on hybrid.
And as for quiting, it now takes 6 seconds, hybrid takes 2.

So... that's a p450 w/384M quiting 2K clients in 6 seconds, and your machine which is roughly 6 times faster needs 360 seconds.. doesn't exactly make sense eh? ;p.

And then unreal killing your whole server...

I dunnow, I kinda suspect something else than unreal then...

Issue History

Date Modified Username Field Change
2004-09-16 22:58 brain2 New Issue
2004-09-16 22:58 brain2 3rd party modules => N/A
2004-09-16 23:29 syzop Note Added: 0007668
2004-09-16 23:32 syzop Note Added: 0007669
2004-09-16 23:43 brain2 Note Added: 0007670
2004-09-16 23:54 syzop Note Added: 0007671
2004-09-17 20:56 brain2 Note Added: 0007681
2004-09-17 21:37 al5001 Note Added: 0007683
2004-09-17 21:40 al5001 Note Edited: 0007683
2004-09-17 22:41 brain2 Note Added: 0007690
2004-09-17 22:43 brain2 Note Added: 0007691
2004-09-17 23:21 syzop Note Added: 0007692
2004-12-11 21:21 syzop Status new => closed