From jonathan.buzzard at strath.ac.uk Tue Jun 24 16:13:48 2025 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Tue, 24 Jun 2025 16:13:48 +0100 Subject: [gpfsug-discuss] Node in cluster but not in cluster Message-ID: <4d9ca2a7-080a-4872-9703-5cd4fa74d37e@strath.ac.uk> I was trying to add a node to the cluster but after an hour I ctrl-c it as it was clearly not in a good place. It told me that it didn't make changes ^Cmmdsh: Caught SIG INT - terminating the child processes. mmaddnode: Interrupt received: No changes made. I checked to make sure all the networking is correct, GPFS packages are installed on the node etc. and it is all good so I go to add it again and mmaddnode: Node <##### censored> was not added to the cluster. The node appears to already belong to a GPFS cluster. mmaddnode: mmaddnode quitting. None of the specified nodes are valid. mmaddnode: Command failed. Examine previous error messages to determine cause. Ok, lets try and delete it mmdelnode: Incorrect node <##### censored> specified for command. mmdelnode: No nodes were found that matched the input specification. mmdelnode: Command failed. Examine previous error messages to determine cause. Trying to do mmsdrrestore on the node gives mmsdrrestore: There is no record for this node in file gpfs0:/var/mmfs/gen/mmsdrfs. Either the node is not part of the cluster, or the file is for a different cluster, or not all of the node's adapter interfaces have been activated yet. mmsdrrestore: Command failed. Examine previous error messages to determine cause. The node does not show up in mmlsnode or mmlscluster Anyone and idea what is going on? JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG From Achim.Rehor at de.ibm.com Tue Jun 24 16:59:30 2025 From: Achim.Rehor at de.ibm.com (Achim Rehor) Date: Tue, 24 Jun 2025 15:59:30 +0000 Subject: [gpfsug-discuss] Node in cluster but not in cluster In-Reply-To: <4d9ca2a7-080a-4872-9703-5cd4fa74d37e@strath.ac.uk> References: <4d9ca2a7-080a-4872-9703-5cd4fa74d37e@strath.ac.uk> Message-ID: I think the node was not added to the cluster, but some changes on the node itself took place already (like creating /var/mmfs/gen ... and possibly others) There is/was a chapter/appendix to the GPFS Adminstration Guide, telling how to permanently remove GPFS from a node, that might help on how to stop the mmaddnode command to cope with a node already belonging to a cluster ... -- Mit freundlichen Grüßen / Kind regards Achim Rehor Technical Support Specialist S​pectrum Scale and ESS (SME) Advisory Product Services Professional IBM Systems Storage Support - EMEA Achim.Rehor at de.ibm.com +49-170-4521194 IBM Deutschland GmbH Vorsitzender des Aufsichtsrats: Sebastian Krause Geschäftsführung: Gregor Pillen (Vorsitzender), Nicole Reimer, Gabriele Schwarenthorer, Christine Rupp, Frank Theisen Sitz der Gesellschaft: Ehningen / Registergericht: Amtsgericht Stuttgart, HRB 14562 / WEEE-Reg.-Nr. DE 99369940 -----Original Message----- From: Jonathan Buzzard > Reply-To: gpfsug main discussion list > To: gpfsug-discuss at gpfsug.org > Subject: [EXTERNAL] [gpfsug-discuss] Node in cluster but not in cluster Date: Tue, 24 Jun 2025 16:13:48 +0100 I was trying to add a node to the cluster but after an hour I ctrl-c it as it was clearly not in a good place. It told me that it didn't make changes ^Cmmdsh: Caught SIG INT - terminating the child processes. mmaddnode: Interrupt received: No changes made. I checked to make sure all the networking is correct, GPFS packages are installed on the node etc. and it is all good so I go to add it again and mmaddnode: Node <##### censored> was not added to the cluster. The node appears to already belong to a GPFS cluster. mmaddnode: mmaddnode quitting. None of the specified nodes are valid. mmaddnode: Command failed. Examine previous error messages to determine cause. Ok, lets try and delete it mmdelnode: Incorrect node <##### censored> specified for command. mmdelnode: No nodes were found that matched the input specification. mmdelnode: Command failed. Examine previous error messages to determine cause. Trying to do mmsdrrestore on the node gives mmsdrrestore: There is no record for this node in file gpfs0:/var/mmfs/gen/mmsdrfs. Either the node is not part of the cluster, or the file is for a different cluster, or not all of the node's adapter interfaces have been activated yet. mmsdrrestore: Command failed. Examine previous error messages to determine cause. The node does not show up in mmlsnode or mmlscluster Anyone and idea what is going on? JAB. -------------- next part -------------- An HTML attachment was scrubbed... URL: From lhorrocks-barlow at ocf.co.uk Tue Jun 24 17:29:03 2025 From: lhorrocks-barlow at ocf.co.uk (Laurence Horrocks-Barlow) Date: Tue, 24 Jun 2025 16:29:03 +0000 Subject: [gpfsug-discuss] Node in cluster but not in cluster Message-ID: Hi JAB, Inline with Achim, I’d be tempted to do the following (based on what’s in your email) From the failed node ssh rm -rfv /var/mmf/gen confirm ssh and/or api auth etc – fix if required confirm firewall ports – fix if required halt -p --------------------------------- From a node in the cluster mmdelnode -p --------------------------------- Boot new node up again Check pings etc, and try again. If that doesn’t work; let us know. Laurence Horrocks-Barlow | Technical Director [https://ocf.co.uk/media/0i1lzfjz/ocf-logo-strapline-2022-blk.png] Phone: 0114 257 2200 Address: OCF Limited, Unit 5 Rotunda Business Centre, Thorncliffe Park, Chapeltown, Sheffield S35 2PG Website: www.ocf.co.uk [LinkedIn icon] [Twitter icon] [https://ocf.co.uk/media/imkoyi2j/line.jpg] OCF Limited is a company registered in England and Wales. Registered number 4132533, VAT number GB 780 6803 14. Registered office address: OCF Limited, 5 Rotunda Business Centre, Thorncliffe Park, Chapeltown, Sheffield, S35 2PG. This message is private and confidential. If you have received this message in error, please notify us immediately and remove it from your system. ------------- Scheduled Annual Leave: Click to see my full annual leave calendar "It is well known that a vital ingredient of success is not knowing that what you're attempting can't be done." -- Sir Terry Pratchett I think the node was not added to the cluster, but some changes on the node itself took place already (like creating /var/mmfs/gen ... and possibly others) There is/was a chapter/appendix to the GPFS Adminstration Guide, telling how to permanently remove GPFS from a node, that might help on how to stop the mmaddnode command to cope with a node already belonging to a cluster ... -- Mit freundlichen Grüßen / Kind regards Achim Rehor Technical Support Specialist S​pectrum Scale and ESS (SME) Advisory Product Services Professional IBM Systems Storage Support - EMEA Achim.Rehor at de.ibm.com> +49-170-4521194 IBM Deutschland GmbH Vorsitzender des Aufsichtsrats: Sebastian Krause Geschäftsführung: Gregor Pillen (Vorsitzender), Nicole Reimer, Gabriele Schwarenthorer, Christine Rupp, Frank Theisen Sitz der Gesellschaft: Ehningen / Registergericht: Amtsgericht Stuttgart, HRB 14562 / WEEE-Reg.-Nr. DE 99369940 -----Original Message----- From: Jonathan Buzzard %3e>> Reply-To: gpfsug main discussion list %3e>> To: gpfsug-discuss at gpfsug.org %22%20%3cgpfsug-discuss at gpfsug.org%3e>> Subject: [EXTERNAL] [gpfsug-discuss] Node in cluster but not in cluster Date: Tue, 24 Jun 2025 16:13:48 +0100 I was trying to add a node to the cluster but after an hour I ctrl-c it as it was clearly not in a good place. It told me that it didn't make changes ^Cmmdsh: Caught SIG INT - terminating the child processes. mmaddnode: Interrupt received: No changes made. I checked to make sure all the networking is correct, GPFS packages are installed on the node etc. and it is all good so I go to add it again and mmaddnode: Node <##### censored> was not added to the cluster. The node appears to already belong to a GPFS cluster. mmaddnode: mmaddnode quitting. None of the specified nodes are valid. mmaddnode: Command failed. Examine previous error messages to determine cause. Ok, lets try and delete it mmdelnode: Incorrect node <##### censored> specified for command. mmdelnode: No nodes were found that matched the input specification. mmdelnode: Command failed. Examine previous error messages to determine cause. Trying to do mmsdrrestore on the node gives mmsdrrestore: There is no record for this node in file gpfs0:/var/mmfs/gen/mmsdrfs. Either the node is not part of the cluster, or the file is for a different cluster, or not all of the node's adapter interfaces have been activated yet. mmsdrrestore: Command failed. Examine previous error messages to determine cause. The node does not show up in mmlsnode or mmlscluster Anyone and idea what is going on? JAB. -------------- next part -------------- An HTML attachment was scrubbed... URL: From jonathan.buzzard at strath.ac.uk Tue Jun 24 17:54:27 2025 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Tue, 24 Jun 2025 17:54:27 +0100 Subject: [gpfsug-discuss] Node in cluster but not in cluster In-Reply-To: References: <4d9ca2a7-080a-4872-9703-5cd4fa74d37e@strath.ac.uk> Message-ID: <79c7cc17-2afd-4f12-9499-dfcb08ca97b2@strath.ac.uk> On 24/06/2025 16:59, Achim Rehor wrote: > > I think the node was not added to the cluster, but some changes on the > node itself took place already (like creating /var/mmfs/gen ... and > possibly others) > > There is/was a chapter/appendix to the GPFS Adminstration Guide, telling > how to permanently remove GPFS  from a node,   that might help  on how > to stop the mmaddnode command  to cope with a node already belonging to > a cluster ... > I removed the GPFS packages from the node, nuked /var/mmfs and /usr/lpp reinstalled GPFS and that clears the issue but it is stuck again There is a tsgskkm process running at 100%, been running 15 minutes now :-( Before I go any further I am going to presume GPFS is fine on Zen5 CPUs? Specifically dual EPYC 9555. The node is running 5.2.2-1 JAB. -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG From cblack at nygenome.org Tue Jun 24 18:22:40 2025 From: cblack at nygenome.org (Christopher Black) Date: Tue, 24 Jun 2025 13:22:40 -0400 Subject: [gpfsug-discuss] Node in cluster but not in cluster In-Reply-To: <79c7cc17-2afd-4f12-9499-dfcb08ca97b2@strath.ac.uk> References: <4d9ca2a7-080a-4872-9703-5cd4fa74d37e@strath.ac.uk> <79c7cc17-2afd-4f12-9499-dfcb08ca97b2@strath.ac.uk> Message-ID: I'd try running the mmdelnode process from a different node (we use cluster manager, mmlsmgr). Ensure the target node to be removed is unreachable from system where you run mmdelnode by either shutting host down or setting up a null route from cluster manager. After mmdelnode'ing, fix networking, and mmaddnode. Re CPU compatibility, unsure about Zen5, but we have gpfs 5.1.x.x running on AMD Zen4. Best, Chris On Tue, Jun 24, 2025 at 12:56 PM Jonathan Buzzard < jonathan.buzzard at strath.ac.uk> wrote: > On 24/06/2025 16:59, Achim Rehor wrote: > > > > I think the node was not added to the cluster, but some changes on the > > node itself took place already (like creating /var/mmfs/gen ... and > > possibly others) > > > > There is/was a chapter/appendix to the GPFS Adminstration Guide, telling > > how to permanently remove GPFS from a node, that might help on how > > to stop the mmaddnode command to cope with a node already belonging to > > a cluster ... > > > > I removed the GPFS packages from the node, nuked /var/mmfs and /usr/lpp > reinstalled GPFS and that clears the issue but it is stuck again > > There is a tsgskkm process running at 100%, been running 15 minutes now :-( > > Before I go any further I am going to presume GPFS is fine on Zen5 CPUs? > Specifically dual EPYC 9555. The node is running 5.2.2-1 > > > JAB. > > -- > Jonathan A. Buzzard Tel: +44141-5483420 > HPC System Administrator, ARCHIE-WeSt. > University of Strathclyde, John Anderson Building, Glasgow. G4 0NG > > > _______________________________________________ > gpfsug-discuss mailing list > gpfsug-discuss at gpfsug.org > http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org > -- This message is for the recipient’s use only, and may contain confidential, privileged or protected information. Any unauthorized use or dissemination of this communication is prohibited. If you received this message in error, please immediately notify the sender and destroy all copies of this message. The recipient should check this email and any attachments for the presence of viruses, as we accept no liability for any damage caused by any virus transmitted by this email. -------------- next part -------------- An HTML attachment was scrubbed... URL: From truongv at us.ibm.com Tue Jun 24 19:00:14 2025 From: truongv at us.ibm.com (Truong Vu) Date: Tue, 24 Jun 2025 18:00:14 +0000 Subject: [gpfsug-discuss] gpfsug-discuss Digest, Vol 155, Issue 2 In-Reply-To: References: Message-ID: <0A270B9E-6B84-4B8B-8CA3-336CCC3C4A61@us.ibm.com> There is an undocumented option for this purpose. You can issue mmdelnode -f on the node bad node. This cleans up leftover configuration and stop/start services if needed. >>> Specifically dual EPYC 9555 If tsgskkm is hung, you may hit a known gskit issue. Can you manually apply the workaround and see if it works? Insert the following lines to file /usr/lpp/mmfs/lib/gsk8/C/icc/icclib/ICCSIG.txt ICC_SHIFT=3 ICC_TRNG=TRNG_ALT4 Insert the following lines to file /usr/lpp/mmfs/lib/gsk8/N/icc/icclib/ICCSIG.txt ICC_TRNG=TRNG_ALT4 Can you post lscpu output? Thanks, Tru. On 6/24/25, 12:32 PM, "gpfsug-discuss on behalf of gpfsug-discuss-request at gpfsug.org " on behalf of gpfsug-discuss-request at gpfsug.org > wrote: Send gpfsug-discuss mailing list submissions to gpfsug-discuss at gpfsug.org To subscribe or unsubscribe via the World Wide Web, visit http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org or, via email, send a message with subject or body 'help' to gpfsug-discuss-request at gpfsug.org You can reach the person managing the list at gpfsug-discuss-owner at gpfsug.org When replying, please edit your Subject line so it is more specific than "Re: Contents of gpfsug-discuss digest..." Today's Topics: 1. Node in cluster but not in cluster (Laurence Horrocks-Barlow) ---------------------------------------------------------------------- Message: 1 Date: Tue, 24 Jun 2025 16:29:03 +0000 From: Laurence Horrocks-Barlow > To: gpfsug main discussion list > Subject: [gpfsug-discuss] Node in cluster but not in cluster Message-ID: > Content-Type: text/plain; charset="utf-8" Hi JAB, Inline with Achim, I?d be tempted to do the following (based on what?s in your email) >From the failed node ssh rm -rfv /var/mmf/gen confirm ssh and/or api auth etc ? fix if required confirm firewall ports ? fix if required halt -p --------------------------------- >From a node in the cluster mmdelnode -p --------------------------------- Boot new node up again Check pings etc, and try again. If that doesn?t work; let us know. Laurence Horrocks-Barlow | Technical Director [https://ocf.co.uk/media/0i1lzfjz/ocf-logo-strapline-2022-blk.png ] > Phone: 0114 257 2200 Address: OCF Limited, Unit 5 Rotunda Business Centre, Thorncliffe Park, Chapeltown, Sheffield S35 2PG Website: www.ocf.co.uk > [LinkedIn icon] > [Twitter icon] > [https://ocf.co.uk/media/imkoyi2j/line.jpg ] OCF Limited is a company registered in England and Wales. Registered number 4132533, VAT number GB 780 6803 14. Registered office address: OCF Limited, 5 Rotunda Business Centre, Thorncliffe Park, Chapeltown, Sheffield, S35 2PG. This message is private and confidential. If you have received this message in error, please notify us immediately and remove it from your system. ------------- Scheduled Annual Leave: Click to see my full annual leave calendar > "It is well known that a vital ingredient of success is not knowing that what you're attempting can't be done." -- Sir Terry Pratchett I think the node was not added to the cluster, but some changes on the node itself took place already (like creating /var/mmfs/gen ... and possibly others) There is/was a chapter/appendix to the GPFS Adminstration Guide, telling how to permanently remove GPFS from a node, that might help on how to stop the mmaddnode command to cope with a node already belonging to a cluster ... -- Mit freundlichen Gr??en / Kind regards Achim Rehor Technical Support Specialist S?pectrum Scale and ESS (SME) Advisory Product Services Professional IBM Systems Storage Support - EMEA Achim.Rehor at de.ibm.com > >> +49-170-4521194 IBM Deutschland GmbH Vorsitzender des Aufsichtsrats: Sebastian Krause Gesch?ftsf?hrung: Gregor Pillen (Vorsitzender), Nicole Reimer, Gabriele Schwarenthorer, Christine Rupp, Frank Theisen Sitz der Gesellschaft: Ehningen / Registergericht: Amtsgericht Stuttgart, HRB 14562 / WEEE-Reg.-Nr. DE 99369940 -----Original Message----- From: Jonathan Buzzard > >%3e>> Reply-To: gpfsug main discussion list > >%3e>> To: gpfsug-discuss at gpfsug.org > > >%22%20%3cgpfsug-discuss at gpfsug.org >%3e>> Subject: [EXTERNAL] [gpfsug-discuss] Node in cluster but not in cluster Date: Tue, 24 Jun 2025 16:13:48 +0100 I was trying to add a node to the cluster but after an hour I ctrl-c it as it was clearly not in a good place. It told me that it didn't make changes ^Cmmdsh: Caught SIG INT - terminating the child processes. mmaddnode: Interrupt received: No changes made. I checked to make sure all the networking is correct, GPFS packages are installed on the node etc. and it is all good so I go to add it again and mmaddnode: Node <##### censored> was not added to the cluster. The node appears to already belong to a GPFS cluster. mmaddnode: mmaddnode quitting. None of the specified nodes are valid. mmaddnode: Command failed. Examine previous error messages to determine cause. Ok, lets try and delete it mmdelnode: Incorrect node <##### censored> specified for command. mmdelnode: No nodes were found that matched the input specification. mmdelnode: Command failed. Examine previous error messages to determine cause. Trying to do mmsdrrestore on the node gives mmsdrrestore: There is no record for this node in file gpfs0:/var/mmfs/gen/mmsdrfs. Either the node is not part of the cluster, or the file is for a different cluster, or not all of the node's adapter interfaces have been activated yet. mmsdrrestore: Command failed. Examine previous error messages to determine cause. The node does not show up in mmlsnode or mmlscluster Anyone and idea what is going on? JAB. -------------- next part -------------- An HTML attachment was scrubbed... URL: > ------------------------------ Subject: Digest Footer _______________________________________________ gpfsug-discuss mailing list gpfsug-discuss at gpfsug.org http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org ------------------------------ End of gpfsug-discuss Digest, Vol 155, Issue 2 ********************************************** From jonathan.buzzard at strath.ac.uk Wed Jun 25 09:38:19 2025 From: jonathan.buzzard at strath.ac.uk (Jonathan Buzzard) Date: Wed, 25 Jun 2025 09:38:19 +0100 Subject: [gpfsug-discuss] gpfsug-discuss Digest, Vol 155, Issue 2 In-Reply-To: <0A270B9E-6B84-4B8B-8CA3-336CCC3C4A61@us.ibm.com> References: <0A270B9E-6B84-4B8B-8CA3-336CCC3C4A61@us.ibm.com> Message-ID: <575f9664-76af-43fd-8205-6ed2ea5ab750@strath.ac.uk> On 24/06/2025 19:00, Truong Vu wrote: > > There is an undocumented option for this purpose. You can issue > mmdelnode -f on the node bad node. This cleans up leftover > configuration and stop/start services if needed. Thanks, one to remember then. Though to be honest having used GPFS now for 18+ years that is the first time I have needed it. >>>> Specifically dual EPYC 9555 > If tsgskkm is hung, you may hit a known gskit issue. Can you > manually apply the workaround and see if it works? > > Insert the following lines to file > /usr/lpp/mmfs/lib/gsk8/C/icc/icclib/ICCSIG.txt > > ICC_SHIFT=3 > ICC_TRNG=TRNG_ALT4 > > Insert the following lines to file > /usr/lpp/mmfs/lib/gsk8/N/icc/icclib/ICCSIG.txt > > ICC_TRNG=TRNG_ALT4 > That did the trick. I quick Google shows this being an issue back in 2020 (on this list) with GPFS 4.2 on AMD Epyc. And also this APAR from 2023 https://www.ibm.com/support/pages/apar/IJ43790 The suggest fix is a little different too. However I already have some AMD EPYC 7513 based servers on the system running 5.1.9-6 (to be upgraded real soon now to 5.2.2-1) which are according to lscpu CPU family 25. I have no recollection of doing anything special and I don't notice the fix in the files. > Can you post lscpu output? > See below, my educated guess is that this is CPU family 26 and whatever fix IBM introduced for CPU a family 25 doesn't work on Zen 5 CPU's. JAB. Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 52 bits physical, 57 bits virtual Byte Order: Little Endian CPU(s): 256 On-line CPU(s) list: 0-255 Vendor ID: AuthenticAMD BIOS Vendor ID: AMD Model name: AMD EPYC 9555 64-Core Processor BIOS Model name: AMD EPYC 9555 64-Core Processor CPU family: 26 Model: 2 Thread(s) per core: 2 Core(s) per socket: 64 Socket(s): 2 Stepping: 1 Frequency boost: enabled CPU(s) scaling MHz: 72% CPU max MHz: 4409.3750 CPU min MHz: 1500.0000 BogoMIPS: 6390.74 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_go od amd_lbr_v2 nopl nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdran d lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba perfmon_v2 ibrs ibpb stibp ibrs_enhanced vmmcall fsgsbase tsc_adjust bmi1 avx2 smep bmi2 invpcid cqm rdt_a avx512f avx512dq rdseed a dx smap avx512ifma clflushopt clwb avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local avx_vnni avx512 _bf16 clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin cppc arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold a vic v_vmsave_vmload vgif x2avic v_spec_ctrl vnmi avx512vbmi umip pku ospke avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg avx512_vpopcntdq la57 rdpid bu s_lock_detect movdiri movdir64b overflow_recov succor smca avx512_vp2intersect flush_l1d debug_swap amd_lbr_pmc_freeze Virtualization features: Virtualization: AMD-V Caches (sum of all): L1d: 6 MiB (128 instances) L1i: 4 MiB (128 instances) L2: 128 MiB (128 instances) L3: 512 MiB (16 instances) NUMA: NUMA node(s): 2 NUMA node0 CPU(s): 0-63,128-191 NUMA node1 CPU(s): 64-127,192-255 Vulnerabilities: Gather data sampling: Not affected Itlb multihit: Not affected L1tf: Not affected Mds: Not affected Meltdown: Not affected Mmio stale data: Not affected Reg file data sampling: Not affected Retbleed: Not affected Spec rstack overflow: Not affected Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; STIBP always-on; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected Srbds: Not affected Tsx async abort: Not affected -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG From truongv at us.ibm.com Wed Jun 25 17:12:13 2025 From: truongv at us.ibm.com (Truong Vu) Date: Wed, 25 Jun 2025 16:12:13 +0000 Subject: [gpfsug-discuss] gpfsug-discuss Digest, Vol 155, Issue 4 In-Reply-To: References: Message-ID: <6CF7E5DF-FC13-40B6-8367-BE19B829188D@us.ibm.com> Zen5 (CPU family 26) is not include in the list that GPFS checks and applies the workaround if needed. If you are running 5.2.2.1 on CPU family 25 model 1 and 17, it should automatically checks and applies the workaround. Thanks for providing the lscpu output. We will add family 26 to the list. Thanks, Tru. On 6/25/25, 7:01 AM, "gpfsug-discuss on behalf of gpfsug-discuss-request at gpfsug.org " on behalf of gpfsug-discuss-request at gpfsug.org > wrote: Send gpfsug-discuss mailing list submissions to gpfsug-discuss at gpfsug.org To subscribe or unsubscribe via the World Wide Web, visit http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org or, via email, send a message with subject or body 'help' to gpfsug-discuss-request at gpfsug.org You can reach the person managing the list at gpfsug-discuss-owner at gpfsug.org When replying, please edit your Subject line so it is more specific than "Re: Contents of gpfsug-discuss digest..." Today's Topics: 1. Re: gpfsug-discuss Digest, Vol 155, Issue 2 (Jonathan Buzzard) ---------------------------------------------------------------------- Message: 1 Date: Wed, 25 Jun 2025 09:38:19 +0100 From: Jonathan Buzzard > To: gpfsug-discuss at gpfsug.org Subject: Re: [gpfsug-discuss] gpfsug-discuss Digest, Vol 155, Issue 2 Message-ID: <575f9664-76af-43fd-8205-6ed2ea5ab750 at strath.ac.uk > Content-Type: text/plain; charset=UTF-8; format=flowed On 24/06/2025 19:00, Truong Vu wrote: > > There is an undocumented option for this purpose. You can issue > mmdelnode -f on the node bad node. This cleans up leftover > configuration and stop/start services if needed. Thanks, one to remember then. Though to be honest having used GPFS now for 18+ years that is the first time I have needed it. >>>> Specifically dual EPYC 9555 > If tsgskkm is hung, you may hit a known gskit issue. Can you > manually apply the workaround and see if it works? > > Insert the following lines to file > /usr/lpp/mmfs/lib/gsk8/C/icc/icclib/ICCSIG.txt > > ICC_SHIFT=3 > ICC_TRNG=TRNG_ALT4 > > Insert the following lines to file > /usr/lpp/mmfs/lib/gsk8/N/icc/icclib/ICCSIG.txt > > ICC_TRNG=TRNG_ALT4 > That did the trick. I quick Google shows this being an issue back in 2020 (on this list) with GPFS 4.2 on AMD Epyc. And also this APAR from 2023 https://www.ibm.com/support/pages/apar/IJ43790 The suggest fix is a little different too. However I already have some AMD EPYC 7513 based servers on the system running 5.1.9-6 (to be upgraded real soon now to 5.2.2-1) which are according to lscpu CPU family 25. I have no recollection of doing anything special and I don't notice the fix in the files. > Can you post lscpu output? > See below, my educated guess is that this is CPU family 26 and whatever fix IBM introduced for CPU a family 25 doesn't work on Zen 5 CPU's. JAB. Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 52 bits physical, 57 bits virtual Byte Order: Little Endian CPU(s): 256 On-line CPU(s) list: 0-255 Vendor ID: AuthenticAMD BIOS Vendor ID: AMD Model name: AMD EPYC 9555 64-Core Processor BIOS Model name: AMD EPYC 9555 64-Core Processor CPU family: 26 Model: 2 Thread(s) per core: 2 Core(s) per socket: 64 Socket(s): 2 Stepping: 1 Frequency boost: enabled CPU(s) scaling MHz: 72% CPU max MHz: 4409.3750 CPU min MHz: 1500.0000 BogoMIPS: 6390.74 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_go od amd_lbr_v2 nopl nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdran d lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba perfmon_v2 ibrs ibpb stibp ibrs_enhanced vmmcall fsgsbase tsc_adjust bmi1 avx2 smep bmi2 invpcid cqm rdt_a avx512f avx512dq rdseed a dx smap avx512ifma clflushopt clwb avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local avx_vnni avx512 _bf16 clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin cppc arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold a vic v_vmsave_vmload vgif x2avic v_spec_ctrl vnmi avx512vbmi umip pku ospke avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg avx512_vpopcntdq la57 rdpid bu s_lock_detect movdiri movdir64b overflow_recov succor smca avx512_vp2intersect flush_l1d debug_swap amd_lbr_pmc_freeze Virtualization features: Virtualization: AMD-V Caches (sum of all): L1d: 6 MiB (128 instances) L1i: 4 MiB (128 instances) L2: 128 MiB (128 instances) L3: 512 MiB (16 instances) NUMA: NUMA node(s): 2 NUMA node0 CPU(s): 0-63,128-191 NUMA node1 CPU(s): 64-127,192-255 Vulnerabilities: Gather data sampling: Not affected Itlb multihit: Not affected L1tf: Not affected Mds: Not affected Meltdown: Not affected Mmio stale data: Not affected Reg file data sampling: Not affected Retbleed: Not affected Spec rstack overflow: Not affected Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; STIBP always-on; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected Srbds: Not affected Tsx async abort: Not affected -- Jonathan A. Buzzard Tel: +44141-5483420 HPC System Administrator, ARCHIE-WeSt. University of Strathclyde, John Anderson Building, Glasgow. G4 0NG ------------------------------ Subject: Digest Footer _______________________________________________ gpfsug-discuss mailing list gpfsug-discuss at gpfsug.org http://gpfsug.org/mailman/listinfo/gpfsug-discuss_gpfsug.org ------------------------------ End of gpfsug-discuss Digest, Vol 155, Issue 4 **********************************************