From jonathan.buzzard at strath.ac.uk Sat Feb 7 17:23:42 2026 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Sat, 7 Feb 2026 17:23:42 +0000 Subject: [gpfsug-discuss] Removing GUI nodes In-Reply-To: References: <71b47392-3544-464a-80a4-6623de559f13@fz-juelich.de> Message-ID: I am trying to remove the GUI nodes from our cluster as they are no longer required. I have managed one but the other is proving to be a problem If I try and delete it I get [root at gpfs0 ~]# mmdelnode gui1 mmdelnode: Node gui1. is being used as a performance monitoring collector node. mmdelnode: Command failed. Examine previous error messages to determine cause. OK, lets change that [root at gpfs0 ~]# mmchnode --noperfmon -N gui1 Sat 7 Feb 12:29:48 GMT 2026: mmchnode: Processing node gui1. mmchnode: Propagating the cluster configuration data to all affected nodes. This is an asynchronous process. That looks to have worked, lets try again 15 minutes later [root at gpfs0 ~]# mmdelnode gui1 mmdelnode: Node gui1. is being used as a performance monitoring collector node. mmdelnode: Command failed. Examine previous error messages to determine cause. Still no joy. From recollection the performance monitoring was only added as a requirement for setting up the GUI and was added to the GUI nodes. Do I have to designate another node for performance monitoring or is there a way to get rid of it altogether, given we don't have GUI nodes any more and are unlikely to have them ever again either. JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG From novosirj at rutgers.edu Sat Feb 7 17:30:58 2026 From: novosirj at rutgers.edu (Ryan Novosielski) Date: Sat, 7 Feb 2026 17:30:58 +0000 Subject: [gpfsug-discuss] Removing GUI nodes In-Reply-To: References: <71b47392-3544-464a-80a4-6623de559f13@fz-juelich.de> Message-ID: You might need to change the perfmon config. Or disable perf monitoring. You don’t need the GUI to take advantange of that stuff though, AFAIK. -- #BlackLivesMatter ____ || \\UTGERS, |---------------------------*O*--------------------------- ||_// the State | Ryan Novosielski (he/him) - novosirj at rutgers.edu || \\ University | Sr. Technologist - 973/972.0922 (2x0922) ~*~ RBHS Campus || \\ of NJ | Office of Advanced Research Computing - MSB A555B, Newark `' On Feb 7, 2026, at 12:23, Jonathan Buzzard wrote: I am trying to remove the GUI nodes from our cluster as they are no longer required. I have managed one but the other is proving to be a problem If I try and delete it I get [root at gpfs0 ~]# mmdelnode gui1 mmdelnode: Node gui1. is being used as a performance monitoring collector node. mmdelnode: Command failed. Examine previous error messages to determine cause. OK, lets change that [root at gpfs0 ~]# mmchnode --noperfmon -N gui1 Sat 7 Feb 12:29:48 GMT 2026: mmchnode: Processing node gui1. mmchnode: Propagating the cluster configuration data to all affected nodes. This is an asynchronous process. That looks to have worked, lets try again 15 minutes later [root at gpfs0 ~]# mmdelnode gui1 mmdelnode: Node gui1. is being used as a performance monitoring collector node. mmdelnode: Command failed. Examine previous error messages to determine cause. Still no joy. From recollection the performance monitoring was only added as a requirement for setting up the GUI and was added to the GUI nodes. Do I have to designate another node for performance monitoring or is there a way to get rid of it altogether, given we don't have GUI nodes any more and are unlikely to have them ever again either. JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG _______________________________________________ gpfsug-discuss mailing list gpfsug-discuss at gpfsug.org http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org -------------- next part -------------- An HTML attachment was scrubbed... URL: From janfrode at tanso.net Sat Feb 7 23:00:23 2026 From: janfrode at tanso.net (Jan-Frode Myklebust) Date: Sun, 8 Feb 2026 00:00:23 +0100 Subject: [gpfsug-discuss] Removing GUI nodes In-Reply-To: References: <71b47392-3544-464a-80a4-6623de559f13@fz-juelich.de> Message-ID: mmperfmon config update --collectors othernode1,othernode2 -jf lør. 7. feb. 2026 kl. 18:32 skrev Ryan Novosielski : > You might need to change the perfmon config. Or disable perf monitoring. > > You don’t need the GUI to take advantange of that stuff though, AFAIK. > > -- > #BlackLivesMatter > ____ > || \\UTGERS, |---------------------------*O*--------------------------- > ||_// the State | Ryan Novosielski (he/him) - novosirj at rutgers.edu > || \\ University | Sr. Technologist - 973/972.0922 (2x0922) ~*~ RBHS Campus > || \\ of NJ | Office of Advanced Research Computing - MSB > A555B, Newark > `' > > On Feb 7, 2026, at 12:23, Jonathan Buzzard > wrote: > > > I am trying to remove the GUI nodes from our cluster as they are no longer > required. I have managed one but the other is proving to be a problem > > If I try and delete it I get > > [root at gpfs0 ~]# mmdelnode gui1 > mmdelnode: Node gui1. is being used as a performance monitoring > collector node. > mmdelnode: Command failed. Examine previous error messages to determine > cause. > > OK, lets change that > > [root at gpfs0 ~]# mmchnode --noperfmon -N gui1 > Sat 7 Feb 12:29:48 GMT 2026: mmchnode: Processing node gui1. > mmchnode: Propagating the cluster configuration data to all > affected nodes. This is an asynchronous process. > > > That looks to have worked, lets try again 15 minutes later > > [root at gpfs0 ~]# mmdelnode gui1 > mmdelnode: Node gui1. is being used as a performance monitoring > collector node. > mmdelnode: Command failed. Examine previous error messages to determine > cause. > > Still no joy. From recollection the performance monitoring was only added > as a requirement for setting up the GUI and was added to the GUI nodes. > > Do I have to designate another node for performance monitoring or is there > a way to get rid of it altogether, given we don't have GUI nodes any more > and are unlikely to have them ever again either. > > > JAB. > > -- > Jonathan A. Buzzard Tel: +44141-5483420 > HPC System Administrator, ARCHIE-WeSt. > University of Strathclyde, John Anderson Building, Glasgow. G4 0NG > > _______________________________________________ > gpfsug-discuss mailing list > gpfsug-discuss at gpfsug.org > http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org > > > _______________________________________________ > gpfsug-discuss mailing list > gpfsug-discuss at gpfsug.org > http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org > -------------- next part -------------- An HTML attachment was scrubbed... URL: From jonathan.buzzard at strath.ac.uk Sun Feb 8 11:06:18 2026 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Sun, 8 Feb 2026 11:06:18 +0000 Subject: [gpfsug-discuss] Removing GUI nodes In-Reply-To: References: <71b47392-3544-464a-80a4-6623de559f13@fz-juelich.de> Message-ID: <3548b8ae-5cca-4534-84cc-cec9d6db6c78@strath.ac.uk> On 07/02/2026 23:00, Jan-Frode Myklebust wrote: > mmperfmon config update --collectors othernode1,othernode2 > That just pushes the performance monitoring to other nodes. I am pretty sure prior to installing the GUI nodes I didn't have any "performance monitoring" and having never used it would like to get back to that state of affairs :-) [root at gpfs0 ~]# mmperfmon config delete --all mmperfmon: The performance monitoring configuration cannot be deleted while there are nodes designated for performance monitoring. Please remove the designation with mmchnode --noperfmon. mmperfmon: Command failed. Examine previous error messages to determine cause. Noting that I have already attempted to designate my one remaining performance monitoring node (aka GUI node) as not doing that with mmchnode which does not seem to do anything. Noting that the manpage for mmperfmon say for config delete --all removes the entire performance monitoring configuration from IBM Storage Scale. I have tried upgrading the GUI node to 5.2.3 to match the rest of the system, and running the mmchnode --noperfmon on the node itself with GPFS both started and stopped. None of that made any difference :-( At least I can run mmchconfig release=LATEST now. Interestingly mmlscluster shows the two DSS-G nodes as having a designation of "quorum-manager-perfmon" and the GUI node as having no designation. Noting the two DSS-G nodes don't have the pmcollector or pmsensors RPM's installed. I suspect I am going to need to open a ticket for this one as something wacky is going on. JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG From ncalimet at lenovo.com Sun Feb 8 14:46:21 2026 From: ncalimet at lenovo.com (Nicolas CALIMET) Date: Sun, 8 Feb 2026 14:46:21 +0000 Subject: [gpfsug-discuss] [External] Re: Removing GUI nodes In-Reply-To: <3548b8ae-5cca-4534-84cc-cec9d6db6c78@strath.ac.uk> References: <71b47392-3544-464a-80a4-6623de559f13@fz-juelich.de> <3548b8ae-5cca-4534-84cc-cec9d6db6c78@strath.ac.uk> Message-ID: Hi, You need to remove the perfmon role on all nodes of the cluster that were monitored (incl. GUI nodes) and delete the perfmon config of the cluster (or the other way around, cannot check right now). The gpfs.gss.pmsensor RPM can then be removed from all nodes where it was installed (incl. GUI) and gpfs.gss.pmcollector from GUI servers. HTH ________________________________ From: gpfsug-discuss on behalf of Jonathan Buzzard Sent: Sunday, February 8, 2026 12:06:18 PM To: gpfsug-discuss at gpfsug.org Subject: [External] Re: [gpfsug-discuss] Removing GUI nodes On 07/02/2026 23:00, Jan-Frode Myklebust wrote: > mmperfmon config update --collectors othernode1,othernode2 > That just pushes the performance monitoring to other nodes. I am pretty sure prior to installing the GUI nodes I didn't have any "performance monitoring" and having never used it would like to get back to that state of affairs :-) [root at gpfs0 ~]# mmperfmon config delete --all mmperfmon: The performance monitoring configuration cannot be deleted while there are nodes designated for performance monitoring. Please remove the designation with mmchnode --noperfmon. mmperfmon: Command failed. Examine previous error messages to determine cause. Noting that I have already attempted to designate my one remaining performance monitoring node (aka GUI node) as not doing that with mmchnode which does not seem to do anything. Noting that the manpage for mmperfmon say for config delete --all removes the entire performance monitoring configuration from IBM Storage Scale. I have tried upgrading the GUI node to 5.2.3 to match the rest of the system, and running the mmchnode --noperfmon on the node itself with GPFS both started and stopped. None of that made any difference :-( At least I can run mmchconfig release=LATEST now. Interestingly mmlscluster shows the two DSS-G nodes as having a designation of "quorum-manager-perfmon" and the GUI node as having no designation. Noting the two DSS-G nodes don't have the pmcollector or pmsensors RPM's installed. I suspect I am going to need to open a ticket for this one as something wacky is going on. JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG _______________________________________________ gpfsug-discuss mailing list gpfsug-discuss at gpfsug.org https://apc01.safelinks.protection.outlook.com/?url=http%3A%2F%2Fgpfsug.org%2Fmailman%2Flistinfo%2Fgpfsug-discuss_gpfsug.org&data=05%7C02%7Cncalimet%40lenovo.com%7Cb509d9568ff940ac354b08de670262ce%7C5c7d0b28bdf8410caa934df372b16203%7C0%7C0%7C639061457179974771%7CUnknown%7CTWFpbGZsb3d8eyJFbXB0eU1hcGkiOnRydWUsIlYiOiIwLjAuMDAwMCIsIlAiOiJXaW4zMiIsIkFOIjoiTWFpbCIsIldUIjoyfQ%3D%3D%7C0%7C%7C%7C&sdata=fPtDlH1XDo0Z9BzSX4nHkq2dI9czuS7bWC5Xr4XUdwE%3D&reserved=0 -------------- next part -------------- An HTML attachment was scrubbed... URL: From novosirj at rutgers.edu Sun Feb 8 15:41:57 2026 From: novosirj at rutgers.edu (Ryan Novosielski) Date: Sun, 8 Feb 2026 15:41:57 +0000 Subject: [gpfsug-discuss] Removing GUI nodes In-Reply-To: <3548b8ae-5cca-4534-84cc-cec9d6db6c78@strath.ac.uk> References: <71b47392-3544-464a-80a4-6623de559f13@fz-juelich.de> <3548b8ae-5cca-4534-84cc-cec9d6db6c78@strath.ac.uk> Message-ID: That does not sound wacky. I believe you can enable those roles without actually having the software installed (mine are that way, for the server side – client clusters, we don’t have that turned on). I would think if you were to remove the role from all of the systems, you should be able to do what you wanted to do. That role designates that the node should be monitored for performance, not host the service. This one may be in the docs, although I’ve not looked myself – how to remove/disable the performance monitoring system. Sent from my iPhone > On Feb 8, 2026, at 06:08, Jonathan Buzzard wrote: > > On 07/02/2026 23:00, Jan-Frode Myklebust wrote: >> mmperfmon config update --collectors othernode1,othernode2 > > That just pushes the performance monitoring to other nodes. I am pretty sure prior to installing the GUI nodes I didn't have any "performance monitoring" and having never used it would like to get back to that state of affairs :-) > > [root at gpfs0 ~]# mmperfmon config delete --all > mmperfmon: The performance monitoring configuration cannot be deleted while there are nodes designated for performance monitoring. Please remove the designation with mmchnode --noperfmon. > mmperfmon: Command failed. Examine previous error messages to determine cause. > > Noting that I have already attempted to designate my one remaining performance monitoring node (aka GUI node) as not doing that with mmchnode which does not seem to do anything. > > Noting that the manpage for mmperfmon say for config delete --all > > removes the entire performance monitoring configuration from > IBM Storage Scale. > > I have tried upgrading the GUI node to 5.2.3 to match the rest of the system, and running the mmchnode --noperfmon on the node itself with GPFS both started and stopped. None of that made any difference :-( At least I can run mmchconfig release=LATEST now. > > Interestingly mmlscluster shows the two DSS-G nodes as having a designation of "quorum-manager-perfmon" and the GUI node as having no designation. Noting the two DSS-G nodes don't have the pmcollector or pmsensors RPM's installed. > > I suspect I am going to need to open a ticket for this one as something wacky is going on. > > > JAB. > > -- > Jonathan A. Buzzard Tel: +44141-5483420 > HPC System Administrator, ARCHIE-WeSt. > University of Strathclyde, John Anderson Building, Glasgow. G4 0NG > > _______________________________________________ > gpfsug-discuss mailing list > gpfsug-discuss at gpfsug.org > http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org From Ivan.Patrick.Lambert at ibm.com Mon Feb 9 10:40:48 2026 From: Ivan.Patrick.Lambert at ibm.com (Ivan Patrick Lambert) Date: Mon, 9 Feb 2026 10:40:48 +0000 Subject: [gpfsug-discuss] Removing GUI nodes In-Reply-To: References: <71b47392-3544-464a-80a4-6623de559f13@fz-juelich.de> Message-ID: Hi John! According to the documentation: https://www.ibm.com/docs/en/storage-scale/5.2.3?topic=reference-mmdelnode-command You need to remove the performance monitoring from the node A node cannot be deleted if any of the following are true: … 4. If the node is configured as a performance monitoring collector. In such cases, you need to remove the node from the performance monitoring configuration by using the mmperfmon config update --collectors command before deleting the node. Deleting a collector node causes loss of all the collected perfmon stats data on the collector node. Hope this helps 😊 Kind regards, Ivan Patrick Lambert EMEA Storage Scale / Storage Scale System Engineer From: gpfsug-discuss on behalf of Jonathan Buzzard Date: Saturday, 7 February 2026 at 18:25 To: gpfsug-discuss at gpfsug.org Subject: [EXTERNAL] [gpfsug-discuss] Removing GUI nodes I am trying to remove the GUI nodes from our cluster as they are no longer required. I have managed one but the other is proving to be a problem If I try and delete it I get [root at gpfs0 ~]# mmdelnode gui1 mmdelnode: Node gui1. is being used as a performance monitoring collector node. mmdelnode: Command failed. Examine previous error messages to determine cause. OK, lets change that [root at gpfs0 ~]# mmchnode --noperfmon -N gui1 Sat 7 Feb 12:29:48 GMT 2026: mmchnode: Processing node gui1. mmchnode: Propagating the cluster configuration data to all affected nodes. This is an asynchronous process. That looks to have worked, lets try again 15 minutes later [root at gpfs0 ~]# mmdelnode gui1 mmdelnode: Node gui1. is being used as a performance monitoring collector node. mmdelnode: Command failed. Examine previous error messages to determine cause. Still no joy. From recollection the performance monitoring was only added as a requirement for setting up the GUI and was added to the GUI nodes. Do I have to designate another node for performance monitoring or is there a way to get rid of it altogether, given we don't have GUI nodes any more and are unlikely to have them ever again either. JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG _______________________________________________ gpfsug-discuss mailing list gpfsug-discuss at gpfsug.org https://urldefense.proofpoint.com/v2/url?u=http-3A__gpfsug.org_mailman_listinfo_gpfsug-2Ddiscuss-5Fgpfsug.org&d=DwIGaQ&c=BSDicqBQBDjDI9RkVyTcHQ&r=w42dL0OuuTTrl9amFK_pSq7sTb5DBqOJPJWkVz4nuxo&m=HiGIov0SPXfrKHwomHu-RgBe5L0zrmJc5oX5rek5HckV1nZt0q7fMc8jzVfAeCY5&s=C73fHYf_YldS99mw4P3-Y4rjKORsdVTYVe44FWVsl0g&e= Unless otherwise stated above: IBM Hrvatska d.o.o. za proizvodnju i trgovinu Ulica Josipa Marohnića 1, 10 000 Zagreb, Hrvatska Upisan kod Trgovačkog suda u Zagrebu pod br. 080011422 Temeljni kapital: 104.580,00 EUR – uplaćen u cijelosti Uprava društva: Tomislav Balun, direktor i Nataša Krupljanin, direktorica Predsjednik nadzornog odbora: Igor Pravica Račun kod: RAIFFEISENBANK AUSTRIA d.d. Zagreb, Magazinska cesta 69, 10000 Zagreb, Hrvatska IBAN: HR5424840081100396574 (SWIFT RZBHHR2X); OIB 43331467622 -------------- next part -------------- An HTML attachment was scrubbed... URL: From jonathan.buzzard at strath.ac.uk Sun Feb 15 22:32:30 2026 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Sun, 15 Feb 2026 22:32:30 +0000 Subject: [gpfsug-discuss] swapped_warn event Message-ID: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> I have a question about the swapped_warn event. It would appear the only way to clear this message is to reboot the node or do a swapoff/swapon If I try an mmhealth event resolve swapped_warn I get a message to say swapped_warn is not manually resolvable. Looking at one of the nodes showing this event, there is 1.2GB of swap being used which is not usual on a Linux server. There is however 160GB of free RAM, the server is not actually "swapping" at the moment and the event is not resolving. It does not appear to be a configurable threshold either. So given that a Linux server is likely to "use" swap even if it has *never* actually ran out of RAM and swapped since it was booted. What's the purpose of this event and can I do something about it? JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG From janfrode at tanso.net Mon Feb 16 11:43:38 2026 From: janfrode at tanso.net (Jan-Frode Myklebust) Date: Mon, 16 Feb 2026 12:43:38 +0100 Subject: [gpfsug-discuss] swapped_warn event In-Reply-To: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> References: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> Message-ID: It seems like someone thinks that linux servers should never use any swap. This swapped_warn triggers if you've used more than 50 MB of swap space. I find this silly.. You can tune it using something like "mmchconfig mmhealth-gpfs-swap_alert_threshold_kb=2000000 --force", but I wouldn't want to pollute my config with such settings.. Maybe just disable it using "mmhealth event hide swapped_warn". -jf On Sun, Feb 15, 2026 at 11:33 PM Jonathan Buzzard wrote: > > > I have a question about the swapped_warn event. It would appear the only > way to clear this message is to reboot the node or do a swapoff/swapon > > If I try an mmhealth event resolve swapped_warn I get a message to say > swapped_warn is not manually resolvable. > > Looking at one of the nodes showing this event, there is 1.2GB of swap > being used which is not usual on a Linux server. There is however 160GB > of free RAM, the server is not actually "swapping" at the moment and the > event is not resolving. > > It does not appear to be a configurable threshold either. > > So given that a Linux server is likely to "use" swap even if it has > *never* actually ran out of RAM and swapped since it was booted. What's > the purpose of this event and can I do something about it? > > > JAB. > > -- > Jonathan A. Buzzard Tel: +44141-5483420 > HPC System Administrator, ARCHIE-WeSt. > University of Strathclyde, John Anderson Building, Glasgow. G4 0NG > > > _______________________________________________ > gpfsug-discuss mailing list > gpfsug-discuss at gpfsug.org > http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org From jonathan.buzzard at strath.ac.uk Mon Feb 16 16:34:46 2026 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Mon, 16 Feb 2026 16:34:46 +0000 Subject: [gpfsug-discuss] swapped_warn event In-Reply-To: References: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> Message-ID: On 16/02/2026 11:43, Jan-Frode Myklebust wrote: > > It seems like someone thinks that linux servers should never use any > swap. This swapped_warn triggers if you've used more than 50 MB of > swap space. I find this silly.. That's a polite way of putting it. Anecdotally, if memory starts to get tight Linux will preventatively push pages to swap "just in case", and might never actually use it. Consequently having a warning if more than 50MB of swap space is used is as useful as a chocolate teapot. If you are going to warn about swap being used, then it needs to be because the kernel was actually shuffling memory between disk and RAM not because it pushed some pages to disk just in case. On an HPC system where the compute nodes are frequently near maximum RAM usage all I unsurprisingly have a load of spurious warnings. > > You can tune it using something like "mmchconfig > mmhealth-gpfs-swap_alert_threshold_kb=2000000 --force", but I wouldn't > want to pollute my config with such settings.. Maybe just disable it > using "mmhealth event hide swapped_warn". > I am debating having a script automatically run "swapoff -a ; swapon -a" on the nodes if these warnings are seen :-) JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG From MDIETZ at de.ibm.com Tue Feb 17 08:39:55 2026 From: MDIETZ at de.ibm.com (Mathias Dietz) Date: Tue, 17 Feb 2026 08:39:55 +0000 Subject: [gpfsug-discuss] swapped_warn event In-Reply-To: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> References: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> Message-ID: Hi Jonathan, Thank you for your question regarding the swapped_warn event. In IBM Spectrum Scale, different categories of health events serve different purposes: State-change events (such as degraded, failed, etc.) indicate actual problems or failures. These events typically resolve automatically; only a small number require manual resolution using the mmhealth event resolve command. TIP events provide recommendations or highlight minor issues, best‑practice deviations, or configuration optimizations. These events do not indicate a failure. If desired, TIP events can be permanently hidden using mmhealth event hide, and once hidden they will not appear again. The swapped_warn event falls into the TIP event category. You are correct that a small amount of swap usage is not necessarily problematic, especially when plenty of RAM is available and the system is not actively swapping. However, we have seen multiple real-world cases where even moderate swap usage negatively impacted system responsiveness and overall Scale performance. Because of this, the TIP is intended to help users who want to extract maximum performance from their systems. If swap usage is not a concern in your environment, you can safely hide the TIP and continue operating normally. I hope this clarifies the purpose of the event and the available options. Please let me know if you need any additional information. best regards Mathias Dietz Storage Scale RAS Architect IBM Deutschland Research & Development GmbH Vorsitzender des Aufsichtsrats: Wolfgang Wendt Geschäftsführung: David Faller Sitz der Gesellschaft: Böblingen / Registergericht: Amtsgericht Stuttgart, HRB 243294 ________________________________ From: gpfsug-discuss on behalf of Jonathan Buzzard Sent: Sunday, February 15, 2026 11:32 PM To: gpfsug main discussion list Subject: [EXTERNAL] [gpfsug-discuss] swapped_warn event I have a question about the swapped_warn event. It would appear the only way to clear this message is to reboot the node or do a swapoff/swapon If I try an mmhealth event resolve swapped_warn I get a message to say swapped_warn is not manually resolvable. Looking at one of the nodes showing this event, there is 1.2GB of swap being used which is not usual on a Linux server. There is however 160GB of free RAM, the server is not actually "swapping" at the moment and the event is not resolving. It does not appear to be a configurable threshold either. So given that a Linux server is likely to "use" swap even if it has *never* actually ran out of RAM and swapped since it was booted. What's the purpose of this event and can I do something about it? JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG _______________________________________________ gpfsug-discuss mailing list gpfsug-discuss at gpfsug.org https://urldefense.proofpoint.com/v2/url?u=http-3A__gpfsug.org_mailman_listinfo_gpfsug-2Ddiscuss-5Fgpfsug.org&d=DwIGaQ&c=BSDicqBQBDjDI9RkVyTcHQ&r=-MaHePVLWSGaTuAdHQuYYijcz_c4mSZc_nxIEMnqfSM&m=x4q-C3lfowRl0YOgnFqC5iyVZdof5xjQOfaUpQxdR4vDBgZnWid74cIfrBmA_bd9&s=_d6CPpuWsfeJ_tP5WNiMebFWaoqirWt4JTf7qdqu4Fw&e= -------------- next part -------------- An HTML attachment was scrubbed... URL: From Alec.Effrat at wellsfargo.com Tue Feb 17 19:53:03 2026 From: Alec.Effrat at wellsfargo.com (Effrat, Alec) Date: Tue, 17 Feb 2026 19:53:03 +0000 Subject: [gpfsug-discuss] swapped_warn event In-Reply-To: References: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> Message-ID: Not sure if this is too trivial for this group… but as a performance person on AIX I generally look to see how often is the sweeping hand coming by and how much RAM is it actually collecting… On Linux you want to watch with something like this: watch -n1 'grep -E "pgfree|pgsteal|pgscan|nr_free_pages" /proc/vmstat' To find out how much pressure your server is under you want to look to see how many page scans are being performed (is counter incrementing quickly)… that means the system WANTS memory, versus how large are the page steals how much memory did it get… Typically, you’d divide the steals by the scans to come up with a pressure ratio. If your server is using ALL of its RAM say for cache hits etc, that can be efficient, so long as the ratio is low to how much RAM it demands and how often is it trying to get free memory. If memory is just lazily being occupied the inactive page count will be high, and as soon as a scan demands memory it will free a TON of memory. But if the memory is marked active, and the scan demands memory it won’t be able to free memory, so you’ll end up with many scans and much smaller freed memory. Hope that helps. From: gpfsug-discuss On Behalf Of Jonathan Buzzard Sent: Monday, February 16, 2026 8:35 AM To: gpfsug-discuss at gpfsug.org Subject: Re: [gpfsug-discuss] swapped_warn event On 16/02/2026 11:43, Jan-Frode Myklebust wrote: > > It seems like someone thinks that linux servers should never use any > swap. This swapped_warn triggers if you've used more than 50 MB of > swap space. I find this silly.. That's a polite way of putting it. Anecdotally, if memory starts to get tight Linux will preventatively push pages to swap "just in case", and might never actually use it. Consequently having a warning if more than 50MB of swap space is used is as useful as a chocolate teapot. If you are going to warn about swap being used, then it needs to be because the kernel was actually shuffling memory between disk and RAM not because it pushed some pages to disk just in case. On an HPC system where the compute nodes are frequently near maximum RAM usage all I unsurprisingly have a load of spurious warnings. > > You can tune it using something like "mmchconfig > mmhealth-gpfs-swap_alert_threshold_kb=2000000 --force", but I wouldn't > want to pollute my config with such settings.. Maybe just disable it > using "mmhealth event hide swapped_warn". > I am debating having a script automatically run "swapoff -a ; swapon -a" on the nodes if these warnings are seen :-) JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG _______________________________________________ gpfsug-discuss mailing list gpfsug-discuss at gpfsug.org https://urldefense.com/v3/__http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org__;!!F9svGWnIaVPGSwU!o15h3llP_t-xAwde_PAYRCF1VYCK7qhlPXsK1TCZNCjGQSiX71ae4hsr_lQQzfzDbWX_iltIVIDP05jTcSf95e2pUMfriO6pIymIMw$ -------------- next part -------------- An HTML attachment was scrubbed... URL: From jonathan.buzzard at strath.ac.uk Wed Feb 18 09:11:14 2026 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Wed, 18 Feb 2026 09:11:14 +0000 Subject: [gpfsug-discuss] swapped_warn event In-Reply-To: References: <53ecedbf-f369-4d07-b4d0-0789b0c33e8d@strath.ac.uk> Message-ID: On 17/02/2026 08:39, Mathias Dietz wrote: > Hi Jonathan, > > Thank you for your question regarding the swapped_warn event. > > In IBM Spectrum Scale, different categories of health events serve > different purposes: > > *State-change events (such as degraded, failed, etc.)* indicate actual > problems or failures. These events typically resolve automatically; only > a small number require manual resolution using the mmhealth event > resolve command. > *TIP events* provide recommendations or highlight minor issues, > best‑practice deviations, or configuration optimizations. These events > do not indicate a failure. If desired, TIP events can be permanently > hidden using mmhealth event hide, and once hidden they will not appear > again. > > The *swapped_warn *event falls into the *TIP event category*. > > You are correct that a small amount of swap usage is not necessarily > problematic, especially when plenty of RAM is available and the system > is not actively swapping. However, we have seen multiple real-world > cases where even moderate swap usage negatively impacted system > responsiveness and overall Scale performance. > > Because of this, the TIP is intended to help users who want to extract > maximum performance from their systems. If swap usage is not a concern > in your environment, you can safely hide the TIP and continue operating > normally. > I agree with everything you have said, but the problem is that GPFS is warning about swap usage even though no swap usage is occurring. That there is usage of swap space is not, and has not been, for over quarter of a century now, a valid measure of whether the Linux kernel is paging memory in/out of disk. You can determine if the kernel is actually paging memory to and from disk with the vmstat command. Checking the nodes reporting the swapped_warn tip shows that they have not paged a single memory page since the last reboot, despite swap usage. As mentioned earlier, the Linux kernel will preemptively write memory pages to swap space as a precaution, so swap usage does not reflect what the GPFS developers think it does. I don't want to hide the tip because it would be useful if it actually did what it said on the tin. The problem is that, at the moment, due to faulty assumptions, the tip is as useful as a chocolate teapot. The TL;DR is that GPFS needs to switch to using a valid measure of paging. JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG