Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #265707 > unrolled thread
| Started by | The Wanderer <wanderer@fastmail.fm> |
|---|---|
| First post | 2024-01-09 14:20 +0100 |
| Last post | 2024-02-15 16:50 +0100 |
| Articles | 20 on this page of 40 — 12 participants |
Back to article view | Back to linux.debian.user
SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-01-09 14:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Dan Ritter <dsr@randomstring.org> - 2024-01-09 16:00 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-01-09 16:30 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-09 19:50 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Curt <curty@free.fr> - 2024-01-09 17:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-01-09 19:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-09 17:30 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-01-09 19:30 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-09 20:10 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-01-09 20:30 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-02-15 04:00 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? songbird <songbird@anthive.com> - 2024-02-15 07:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-02-15 16:50 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-02-15 09:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-02-15 16:50 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-02-15 18:40 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-02-16 11:00 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? "Roy J. Tellason, Sr." <roy@rtellason.com> - 2024-02-16 20:00 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-02-16 21:00 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Gremlin <scott-andrews@columbus.rr.com> - 2024-02-16 22:50 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? "Roy J. Tellason, Sr." <roy@rtellason.com> - 2024-02-17 19:50 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? gene heskett <gheskett@shentel.net> - 2024-02-18 02:50 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? debian-user@howorth.org.uk - 2024-02-15 13:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-02-15 16:10 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-01-09 23:40 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-01-09 23:50 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-10 10:40 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-01-10 16:40 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Curt <curty@free.fr> - 2024-01-10 18:10 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-10 18:40 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-01-11 03:30 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Dan Ritter <dsr@randomstring.org> - 2024-01-10 18:40 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-01-11 03:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Dan Ritter <dsr@randomstring.org> - 2024-01-11 15:10 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? David Christensen <dpchrist@holgerdanske.com> - 2024-01-11 16:10 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Dan Ritter <dsr@randomstring.org> - 2024-01-11 19:30 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Stefan Monnier <monnier@iro.umontreal.ca> - 2024-01-11 21:30 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Michael Stone <mstone@debian.org> - 2024-01-12 00:00 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? Dan Ritter <dsr@randomstring.org> - 2024-01-12 12:20 +0100
Re: SMART Uncorrectable_Error_Cnt rising - should I be worried? The Wanderer <wanderer@fastmail.fm> - 2024-02-15 16:50 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | "Roy J. Tellason, Sr." <roy@rtellason.com> |
|---|---|
| Date | 2024-02-17 19:50 +0100 |
| Message-ID | <I8xV7-aO4s-5@gated-at.bofh.it> |
| In reply to | #267509 |
On Friday 16 February 2024 04:42:12 pm Gremlin wrote: > On 2/16/24 13:56, Roy J. Tellason, Sr. wrote: > > On Friday 16 February 2024 04:52:22 am David Christensen wrote: > >> I think the Raspberry Pi, etc., users on this list live with USB storage > >> and have found it to be reliable enough for personal and SOHO network use. > > > > I have one, haven't done much with it. Are there any alternative ways to interface storage? Maybe add SATA ports or something? > > > On raspberry Pi 1 to 4 No, you have a choice of USB 2 or USB 3 Looks like I'll have to go with a USB - SATA adapter, then. It's a 4, I bought the "Canakit" package that included an enclosure, keyboard, mouse, and a small touch screen (4"?). > Raspberry Pi 5 Yes with and NVME hat interfaced to the pcie "port" > > I am using a Pi 5 (desktop) with USB 3 port hooked to an NVME external > drive and it works just fine. > > It is much faster than the Pi 4 I was using because of the new "south > bridge" I'm aware of the 5 having come out, but haven't explored the possibility of getting one of those yet. -- Member of the toughest, meanest, deadliest, most unrelenting -- and ablest -- form of life in this section of space, a critter that can be killed but can't be tamed. --Robert A. Heinlein, "The Puppet Masters" - Information is more dangerous than cannon to a society ruled by lies. --James M Dakin
[toc] | [prev] | [next] | [standalone]
| From | gene heskett <gheskett@shentel.net> |
|---|---|
| Date | 2024-02-18 02:50 +0100 |
| Message-ID | <I8Etz-aS84-1@gated-at.bofh.it> |
| In reply to | #267544 |
On 2/17/24 13:45, Roy J. Tellason, Sr. wrote: > On Friday 16 February 2024 04:42:12 pm Gremlin wrote: >> On 2/16/24 13:56, Roy J. Tellason, Sr. wrote: >>> On Friday 16 February 2024 04:52:22 am David Christensen wrote: >>>> I think the Raspberry Pi, etc., users on this list live with USB storage >>>> and have found it to be reliable enough for personal and SOHO network use. >>> >>> I have one, haven't done much with it. Are there any alternative ways to interface storage? Maybe add SATA ports or something? >>> >> On raspberry Pi 1 to 4 No, you have a choice of USB 2 or USB 3 > > Looks like I'll have to go with a USB - SATA adapter, then. It's a 4, I bought the "Canakit" package that included an enclosure, keyboard, mouse, and a small touch screen (4"?). > >> Raspberry Pi 5 Yes with and NVME hat interfaced to the pcie "port" >> >> I am using a Pi 5 (desktop) with USB 3 port hooked to an NVME external >> drive and it works just fine. >> >> It is much faster than the Pi 4 I was using because of the new "south >> bridge" > > I'm aware of the 5 having come out, but haven't explored the possibility of getting one of those yet. > StarTech makes an excellent sata-III to usb3 adapter for about a tenner a copy. So a 7 port hub takes up only 1 or the 4 usb3 ports on a bpi-m5, leaving 3 more ports available on the bpi-m5 itself. See at <startech.com> ssh into it from the Main system and run the pi headless. > Cheers, Gene Heskett, CET. -- "There are four boxes to be used in defense of liberty: soap, ballot, jury, and ammo. Please use in that order." -Ed Howdershelt (Author, 1940) If we desire respect for the law, we must first make the law respectable. - Louis D. Brandeis
[toc] | [prev] | [next] | [standalone]
| From | debian-user@howorth.org.uk |
|---|---|
| Date | 2024-02-15 13:20 +0100 |
| Message-ID | <I7ISB-ajue-3@gated-at.bofh.it> |
| In reply to | #267416 |
The Wanderer <wanderer@fastmail.fm> wrote: > It turns out that there is a hard limit of 65000 > hardlinks per on-disk file; That's a filesystem dependent value. That's the value for ext4. XFS has a much larger limit I believe. As well as some other helpful properties for large filesystems. btrfs has different limits, depending on where the hardlinks are, apparently. Some larger, some ridiculously smaller.
[toc] | [prev] | [next] | [standalone]
| From | The Wanderer <wanderer@fastmail.fm> |
|---|---|
| Date | 2024-02-15 16:10 +0100 |
| Message-ID | <I7Lx8-al8w-23@gated-at.bofh.it> |
| In reply to | #267425 |
[Multipart message — attachments visible in raw view] — view raw
On 2024-02-15 at 07:14, debian-user@howorth.org.uk wrote: > The Wanderer <wanderer@fastmail.fm> wrote: > >> It turns out that there is a hard limit of 65000 hardlinks per >> on-disk file; > > That's a filesystem dependent value. That's the value for ext4. I think I recall reading that while I was flailing over this, yes. ext4 is what I use for daily-driver purposes these days; from the little I've looked into the matter, everything else seems to be either too complicated, or too non-robust, to be worth risking my live data on. > XFS has a much larger limit I believe. As well as some other helpful > properties for large filesystems. > > btrfs has different limits, depending on where the hardlinks are, > apparently. Some larger, some ridiculously smaller. So it might make sense to use one of those as the underpinning for whatever external system I wind up setting up for tiered backup, then. Though experimentation to determine the limits would be warranted. That's not immediately actionable, but it's good to have in the background as planning etc. takes place. -- The Wanderer The reasonable man adapts himself to the world; the unreasonable one persists in trying to adapt the world to himself. Therefore all progress depends on the unreasonable man. -- George Bernard Shaw
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2024-01-09 23:40 +0100 |
| Message-ID | <HUsVj-1ZNF-11@gated-at.bofh.it> |
| In reply to | #265707 |
On 1/9/24 05:11, The Wanderer wrote: > I have an eight-drive RAID-6 array of 2TB SSDs, built > back in early-to-mid 2021. > Within the past few weeks, I got root-mail notifications from smartd > that the ATA error count on two of the drives had increased ... > On Sunday (two days ago), I got root-mail notifications from smartd > about *all* of the drives in the array. > One thing I don't know, which may or may not be important, is whether > these alert mails are being triggered when the error-count increase > happens, or when a scheduled check of some type is run. Please do a full backup to a portable HDD ASAP. Put that HDD off-site. Get another HDD and do a full backup. Then do incremental backups daily. After a week, two weeks, or a month, swap the drives. At some point, start destroying older backups to make room for new backups. Please burn your most critical data to a high-quality optical media. Enable checksums by some means (extended attributes, checksum file in volume root, etc.). Validate the burn using the checksums. Then burn and validate new critical data every week, two weeks, month, etc.. Validate checksums periodically. AIUI smartd runs periodically via systemd, Perhaps another reader can post the incantation required to display the settings and/or locate past SMART reports on disk. You can always run smartctl manual to get a SMART report whenever you want (I like the --xall/-x option): # smartctl -x DEV > I've looked at the SMART attributes for the drives, and am having a hard > time determining whether or not there's anything worth being actually > concerned about here. Some of the information I'm seeing seems to > suggest yes, but other information seems to suggest no. Reading SMART reports has a learning curve. STFW for the terms you do not understand. And, beware that different manufacturers with different engineers make different long-term predictions based upon different short-term test data. Looking at SMART reports over time for the same drive, looking for trends, and noticing problems is exactly the right thing to do. You and smartd did good. :-) > Most of the attributes are listed as of type "Old_age". Samsung EVO 870 are good drives, but they are "consumer" drives -- e.g. intended for laptop/ desktop computers that are powered off or hibernating most of the time. The SMART report you attached showed a "Power_On_Hours" attribute value of 22286. Assuming an operational specification of 40 hours/week, that SSD has usage equivalent to 10.7 years. So, it is old. > I don't know how to interpret the "Pre-fail" notation for the other > attributes. AIUI "Pre-fail" indicates the drive is going to fail soon and should be replaced. > My default plan is to identify an appropriate model and buy a pair of > replacement drives, but not install them yet; buy another two drives > every six months, until I have a full replacement set; and start failing > drives out of the RAID array and installing replacements as soon as one > either fails, or looks like it's imminently about to fail. If you want 24x7 storage at minimum total cost of ownership, I suggest 3.5" enterprise HDD's. I buy "new" or "open box" older model drives on eBay, the older the cheaper. They typically die within a month or run for years. SAS has more features, yet can be cheaper (assuming you have compatible hardware). I prefer RAID-10 over RAID-5 or RAID-6 because IOPS scales up linearly with the number of mirrors (spindles). So, if one mirror does 120 IOPS (7200 RPM), two mirrors do 240 IOPS, three do 360 IOPS, etc.. Also, resilvering is a direct disk-to-disk copy at sequential read and write speeds. To get protection against two-device failure, you need 3-day mirrors; or, a hot spare and a time delay longer than resilvering time between failures. Finally, depending upon your choice of RAID, volume management, filesystem, etc., you might be able to re-use those SSD's as accelerators -- read cache, write cache, metadata, etc.. (This is easy on ZFS. Perhaps other readers with madm, LVM, btrfs, etc., can comment on SSD acceleration for those.) David
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2024-01-09 23:50 +0100 |
| Message-ID | <HUt4Z-1ZQY-1@gated-at.bofh.it> |
| In reply to | #265732 |
On 1/9/24 14:34, David Christensen wrote: > You can always run smartctl manual ... Correction: manually > To get protection against two-device failure, you need 3-day mirrors ... Correction: 3-way > Perhaps other readers with madm, ... Correction: mdadm David
[toc] | [prev] | [next] | [standalone]
| From | Michael Kjörling <2695bd53d63c@ewoof.net> |
|---|---|
| Date | 2024-01-10 10:40 +0100 |
| Message-ID | <HUDe1-265T-7@gated-at.bofh.it> |
| In reply to | #265732 |
On 9 Jan 2024 14:34 -0800, from dpchrist@holgerdanske.com (David Christensen): >> I don't know how to interpret the "Pre-fail" notation for the other >> attributes. > > AIUI "Pre-fail" indicates the drive is going to fail soon and should be > replaced. Only if the attribute hits the "failure" threshold, whatever that happens to be or mean for that particular attribute. -- Michael Kjörling 🔗 https://michael.kjorling.se “Remember when, on the Internet, nobody cared that you were a dog?”
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2024-01-10 16:40 +0100 |
| Message-ID | <HUIQq-29vn-7@gated-at.bofh.it> |
| In reply to | #265740 |
On 1/10/24 01:35, Michael Kjörling wrote: > On 9 Jan 2024 14:34 -0800, from dpchrist@holgerdanske.com (David Christensen): >>> I don't know how to interpret the "Pre-fail" notation for the other >>> attributes. >> >> AIUI "Pre-fail" indicates the drive is going to fail soon and should be >> replaced. > > Only if the attribute hits the "failure" threshold, whatever that > happens to be or mean for that particular attribute. If you choose to run RAID member drives all the way to failure, then you need to have a reasonable expectation that the remaining drives have enough reliability and the RAID has enough redundancy to protect the data until the sysadmin notices the failed drive(s), the sysadmin replaces the failed drive(s), and the RAID resilvers. Given the OP's situation -- 8 consumer SSD's, same make and model, possibly from a defective manufacturing batch, all purchased at the same time, all deployed in the same RAID-6, all run 2.5 years 24x7, and all suddenly showing lots of SMART warnings -- I would not have confidence in that RAID. David
[toc] | [prev] | [next] | [standalone]
| From | Curt <curty@free.fr> |
|---|---|
| Date | 2024-01-10 18:10 +0100 |
| Message-ID | <HUKfv-2av3-1@gated-at.bofh.it> |
| In reply to | #265741 |
On 2024-01-10, David Christensen <dpchrist@holgerdanske.com> wrote: > > > Given the OP's situation -- 8 consumer SSD's, same make and model, > possibly from a defective manufacturing batch, all purchased at the same > time, all deployed in the same RAID-6, all run 2.5 years 24x7, and all > suddenly showing lots of SMART warnings -- I would not have confidence > in that RAID. It's curious, but I just heard something on French TV from a journalist that's relevant to this. She said she'd covered the aeronautics field in the past and mentioned the *principe de dissemblance* (dissimilarity principle). Critical redundant parts on aircraft, she claimed, would be sourced from different manufacturers in order to obviate the possibility of redundant failures you've raised here. > > David > >
[toc] | [prev] | [next] | [standalone]
| From | Michael Kjörling <2695bd53d63c@ewoof.net> |
|---|---|
| Date | 2024-01-10 18:40 +0100 |
| Message-ID | <HUKIy-2aF8-5@gated-at.bofh.it> |
| In reply to | #265743 |
On 10 Jan 2024 17:07 -0000, from curty@free.fr (Curt): > It's curious, but I just heard something on French TV from a journalist > that's relevant to this. She said she'd covered the aeronautics field in > the past and mentioned the *principe de dissemblance* (dissimilarity > principle). Critical redundant parts on aircraft, she claimed, would be > sourced from different manufacturers in order to obviate the possibility > of redundant failures you've raised here. Indeed. My understanding is that it's even relatively common, at least for flight-critical components, to use totally different implementations (of both hardware and software), not just sourced from different vendors, resellers or batches, such that the same software bug _cannot_ reasonably appear in both, reducing the scope of software errors to _specification_ bugs, which an inherently engineering field (physical engineering, fluid dynamics, ...) is better equipped to deal with early. Recent events notwithstanding. As for David's note on OP's RAID array, I think that point has been sufficiently made by now in this thread; and let's hope that the new backup drive arrives soon enough that a full copy can be made before there is any actual data loss. -- Michael Kjörling 🔗 https://michael.kjorling.se “Remember when, on the Internet, nobody cared that you were a dog?”
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2024-01-11 03:30 +0100 |
| Message-ID | <HUSZr-2ggV-3@gated-at.bofh.it> |
| In reply to | #265744 |
On 1/10/24 09:30, Michael Kjörling wrote: > My understanding is that it's even relatively common, at least > for flight-critical components, to use totally different > implementations (of both hardware and software), not just sourced from > different vendors, resellers or batches, such that the same software > bug _cannot_ reasonably appear in both, reducing the scope of software > errors to _specification_ bugs, which an inherently engineering field > (physical engineering, fluid dynamics, ...) is better equipped to deal > with early. Recent events notwithstanding. Erlang has a different and interesting philosophy to software systems: https://medium.com/pragmatic-programmers/error-handling-philosophy-d820bd68a469 David
[toc] | [prev] | [next] | [standalone]
| From | Dan Ritter <dsr@randomstring.org> |
|---|---|
| Date | 2024-01-10 18:40 +0100 |
| Message-ID | <HUKIy-2aF8-7@gated-at.bofh.it> |
| In reply to | #265743 |
Curt wrote: > On 2024-01-10, David Christensen <dpchrist@holgerdanske.com> wrote: > > > > > > Given the OP's situation -- 8 consumer SSD's, same make and model, > > possibly from a defective manufacturing batch, all purchased at the same > > time, all deployed in the same RAID-6, all run 2.5 years 24x7, and all > > suddenly showing lots of SMART warnings -- I would not have confidence > > in that RAID. > > It's curious, but I just heard something on French TV from a journalist > that's relevant to this. She said she'd covered the aeronautics field in > the past and mentioned the *principe de dissemblance* (dissimilarity > principle). Critical redundant parts on aircraft, she claimed, would be > sourced from different manufacturers in order to obviate the possibility > of redundant failures you've raised here. I don't know whether that's true in aeronautics, but at the home and small business scale, that's always something I've practiced. At the large scale, server assemblers don't want to mix parts very often (you can get some of them to do it), so you usually need your servers as a whole to be the unit of redundancy, not disks in an array. -dsr-
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2024-01-11 03:20 +0100 |
| Message-ID | <HUSPL-2gdQ-1@gated-at.bofh.it> |
| In reply to | #265743 |
On 1/10/24 09:07, Curt wrote: > On 2024-01-10, David Christensen <dpchrist@holgerdanske.com> wrote: >> Given the OP's situation -- 8 consumer SSD's, same make and model, >> possibly from a defective manufacturing batch, all purchased at the same >> time, all deployed in the same RAID-6, all run 2.5 years 24x7, and all >> suddenly showing lots of SMART warnings -- I would not have confidence >> in that RAID. > > It's curious, but I just heard something on French TV from a journalist > that's relevant to this. She said she'd covered the aeronautics field in > the past and mentioned the *principe de dissemblance* (dissimilarity > principle). Critical redundant parts on aircraft, she claimed, would be > sourced from different manufacturers in order to obviate the possibility > of redundant failures you've raised here. https://en.wikipedia.org/wiki/Failure_analysis Using components from different vendors is a known mitigation technique and can help in the right situations. Relevant to this this thread, some people use disk drives from different manufacturers in their RAID's. Doing so with RAID-10 (stripe of mirrors) is straight forward -- within each mirror, use brand A for the first disk, brand B for the second, brand C for the third (or hot spares), etc.. It then makes sense to do the same with HBA's -- use HBA brand X for the first disks in each mirror, HBA brand Y for the second disks, HBA brand Z for the third/ spare disks, etc.. For x86 workstations and servers, ECC memory, dual network interfaces, and dual power supplies come to mind. I am unclear about dual processors and/or dual memory banks. Moving beyond one computer, the process continues with KVM/ serial console fabric, networks, electric power, cooling, etc.. It's just a question of what failure modes you want to protect against and how much time and money you want to spend. David
[toc] | [prev] | [next] | [standalone]
| From | Dan Ritter <dsr@randomstring.org> |
|---|---|
| Date | 2024-01-11 15:10 +0100 |
| Message-ID | <HV3UR-2n7J-1@gated-at.bofh.it> |
| In reply to | #265764 |
David Christensen wrote: > On 1/10/24 09:07, Curt wrote: > > On 2024-01-10, David Christensen <dpchrist@holgerdanske.com> wrote: > > dual network interfaces, and dual power supplies come to mind. I am unclear > about dual processors and/or dual memory banks. Moving beyond one computer, There are no systems that I'm aware of which allow you to use 2 or more processors of different models; they always have to be exact duplicates. Sometimes different step revisions of the same model will not work -- if Intel makes a Xeon 5254 in March, and fixes things in June, August and November, sometimes the November release won't work perfectly with a chip produced in March. You can always use identically spec'd RAM from different manufacturers in different memory banks, but since it's always possible to power down, replace or just remove memory, and power up again, I don't know that there's any reason to bother distributing the manufacturers in a single machine. -dsr-
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2024-01-11 16:10 +0100 |
| Message-ID | <HV4QW-2nGk-3@gated-at.bofh.it> |
| In reply to | #265788 |
On 1/11/24 05:50, Dan Ritter wrote: > David Christensen wrote: >> dual network interfaces, and dual power supplies come to mind. I am >> unclear about dual processors and/or dual memory banks. > > There are no systems that I'm aware of which allow you to use 2 > or more processors of different models; they always have to be > exact duplicates. Sometimes different step revisions of the same > model will not work -- if Intel makes a Xeon 5254 in March, and > fixes things in June, August and November, sometimes the > November release won't work perfectly with a chip produced in > March. Okay. > You can always use identically spec'd RAM from different > manufacturers in different memory banks, but since it's always > possible to power down, replace or just remove memory, and power > up again, I don't know that there's any reason to bother > distributing the manufacturers in a single machine. Okay. STFW the Dell PowerEdge 6850 (circa 2004) featured "hot plug" disk drives, expansion slots, memory risers, power supplies, and system cooling fans: https://downloads.dell.com/manuals/all-products/esuprt_ser_stor_net/esuprt_poweredge/poweredge-6850_user%27s%20guide4_en-us.pdf STFW dell.com today, I see servers with: * hot plug hard drives * hot spare hard drives * dual hot plug redundant power supplies * dual hot plug fully redundant power supplies * dual hot plug fault tolerant power supplies * dual hot plug fault tolerant redundant power supplies. This Dell article explains some of the PSU options: https://infohub.delltechnologies.com/p/full-redundancy-vs-fault-tolerant-redundancy-for-poweredge-server-psus/ David
[toc] | [prev] | [next] | [standalone]
| From | Dan Ritter <dsr@randomstring.org> |
|---|---|
| Date | 2024-01-11 19:30 +0100 |
| Message-ID | <HV7Yt-2ptL-11@gated-at.bofh.it> |
| In reply to | #265792 |
David Christensen wrote: > On 1/11/24 05:50, Dan Ritter wrote: > > David Christensen wrote: > STFW the Dell PowerEdge 6850 (circa 2004) featured "hot plug" disk drives, > expansion slots, memory risers, power supplies, and system cooling fans: > > https://downloads.dell.com/manuals/all-products/esuprt_ser_stor_net/esuprt_poweredge/poweredge-6850_user%27s%20guide4_en-us.pdf > > > STFW dell.com today, I see servers with: > > * hot plug hard drives > * hot spare hard drives > * dual hot plug redundant power supplies > * dual hot plug fully redundant power supplies > * dual hot plug fault tolerant power supplies > * dual hot plug fault tolerant redundant power supplies. Hot plug disks are easy -- SATA, SAS and NVMe U.2 interfaces are all specified so that the chassis manufacturer can arrange for data to be disconnected before power. This is nearly but not quite ubiquitous in rackmountable servers; somewhat rare in desktops. Hot plugged fans are extremely easy -- there's no data being stored and no state of any consequence. Arranging an easy disconnect mechanism is the maximum difficulty. Hot spare is more a property of the RAID management software. mdadm, btrfs and zfs all support marking a disk as 'spare' and not using it until another disk is marked as failing. Interestingly: all PCIe cards are nominally hot-pluggable. -dsr-
[toc] | [prev] | [next] | [standalone]
| From | Stefan Monnier <monnier@iro.umontreal.ca> |
|---|---|
| Date | 2024-01-11 21:30 +0100 |
| Message-ID | <HV9QB-2qBs-3@gated-at.bofh.it> |
| In reply to | #265788 |
> manufacturers in different memory banks, but since it's always
> possible to power down, replace or just remove memory, and power
> up again,
Hmm... "always"? What about long running computations like that
simulation (or LLM training) launched a month ago and that's expected to
finish in another month or so?
Some mainframes have supported hot (un)plugging RAM modules as well and
I wouldn't be surprised if some x86 servers also support it nowadays.
Stefan
[toc] | [prev] | [next] | [standalone]
| From | Michael Stone <mstone@debian.org> |
|---|---|
| Date | 2024-01-12 00:00 +0100 |
| Message-ID | <HVcbL-2rR5-1@gated-at.bofh.it> |
| In reply to | #265805 |
On Thu, Jan 11, 2024 at 03:25:51PM -0500, Stefan Monnier wrote: >> manufacturers in different memory banks, but since it's always >> possible to power down, replace or just remove memory, and power >> up again, > >Hmm... "always"? What about long running computations like that >simulation (or LLM training) launched a month ago and that's expected to >finish in another month or so? I'd expect something like that to have a checkpoint/restart capability if not starting over actually matters. >Some mainframes have supported hot (un)plugging RAM modules as well Yes, mainframes have been engineered that way for a long time. It makes them very expensive, and their market share has been declining for decades because most problems can be solved more cheaply in software (even while maintaining high availability). Hot *spare* memory is relatively common, as it solves most problems without the complexity of hot *swapping*, at the (generally low) cost of having to schedule downtime at some point in the future to actually replace the failed module.
[toc] | [prev] | [next] | [standalone]
| From | Dan Ritter <dsr@randomstring.org> |
|---|---|
| Date | 2024-01-12 12:20 +0100 |
| Message-ID | <HVnJT-2zhz-1@gated-at.bofh.it> |
| In reply to | #265805 |
Stefan Monnier wrote: > > manufacturers in different memory banks, but since it's always > > possible to power down, replace or just remove memory, and power > > up again, > > Hmm... "always"? What about long running computations like that > simulation (or LLM training) launched a month ago and that's expected to > finish in another month or so? If the job is that big, it's being run on multiple machines. This machine's current chunk is corrupt, so you can't use it anyway. The orchestrator stops using this machine, someone comes in to replace the RAM. Later the machine is re-added to the pool. > Some mainframes have supported hot (un)plugging RAM modules as well and > I wouldn't be surprised if some x86 servers also support it nowadays. https://www.kernel.org/doc/html/latest/admin-guide/mm/memory-hotplug.html That said, you won't find this feature without specifying it when you buy it, and very few have a use case for it. -dsr-
[toc] | [prev] | [next] | [standalone]
| From | The Wanderer <wanderer@fastmail.fm> |
|---|---|
| Date | 2024-02-15 16:50 +0100 |
| Message-ID | <I7M9Q-all9-11@gated-at.bofh.it> |
| In reply to | #265805 |
[Multipart message — attachments visible in raw view] — view raw
On 2024-01-11 at 15:25, Stefan Monnier wrote: >> manufacturers in different memory banks, but since it's always >> possible to power down, replace or just remove memory, and power up >> again, > > Hmm... "always"? What about long running computations like that > simulation (or LLM training) launched a month ago and that's expected > to finish in another month or so? > > Some mainframes have supported hot (un)plugging RAM modules as well > and I wouldn't be surprised if some x86 servers also support it > nowadays. I remember, in my previous job (back in the oughts, now), one occasion on which I was going around adding RAM to various desktop computers in the area under my purview, by adding more DIMMs to the open slots - and discovering, when I put the case back together on one of those computers and went to power it back on, that *it was already powered on and the system was still booted*. Surprisingly, none of the hardware showed any sign of damage, and the system recognized the RAM just fine after a reboot. But it was a bit of a jolt at the time to realize that I'd just done parts surgery, however mild, on a powered and running system. -- The Wanderer The reasonable man adapts himself to the world; the unreasonable one persists in trying to adapt the world to himself. Therefore all progress depends on the unreasonable man. -- George Bernard Shaw
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.debian.user
csiph-web