Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #188783 > unrolled thread

Talking about RAID - disks with same id

Started bydeloptes <deloptes@gmail.com>
First post2017-11-08 22:50 +0100
Last post2017-11-12 20:20 +0100
Articles 19 — 5 participants

Back to article view | Back to linux.debian.user


Contents

  Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-08 22:50 +0100
    Re: Talking about RAID - disks with same id David Christensen <dpchrist@holgerdanske.com> - 2017-11-08 23:40 +0100
      Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-09 08:20 +0100
        Re: Talking about RAID - disks with same id David Christensen <dpchrist@holgerdanske.com> - 2017-11-09 19:20 +0100
          Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-09 22:10 +0100
            Re: Talking about RAID - disks with same id David Christensen <dpchrist@holgerdanske.com> - 2017-11-09 23:50 +0100
    Re: Talking about RAID - disks with same id Tobx <net@tobx.de> - 2017-11-09 10:40 +0100
      Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-09 22:20 +0100
    Re: Talking about RAID - disks with same id Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2017-11-10 00:00 +0100
      Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-10 09:20 +0100
        Re: Talking about RAID - disks with same id Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2017-11-10 18:10 +0100
          Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-10 20:50 +0100
          Re: Talking about RAID - disks with same id Pascal Hambourg <pascal@plouf.fr.eu.org> - 2017-11-12 17:20 +0100
        Re: Talking about RAID - disks with same id David Christensen <dpchrist@holgerdanske.com> - 2017-11-10 21:10 +0100
          Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-10 23:30 +0100
        Re: Talking about RAID - disks with same id Pascal Hambourg <pascal@plouf.fr.eu.org> - 2017-11-12 17:00 +0100
          Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-12 18:00 +0100
            Re: Talking about RAID - disks with same id Pascal Hambourg <pascal@plouf.fr.eu.org> - 2017-11-12 19:50 +0100
              Re: Talking about RAID - disks with same id deloptes <deloptes@gmail.com> - 2017-11-12 20:20 +0100

#188783 — Talking about RAID - disks with same id

Fromdeloptes <deloptes@gmail.com>
Date2017-11-08 22:50 +0100
SubjectTalking about RAID - disks with same id
Message-ID<uJGop-4j2-5@gated-at.bofh.it>
Hi,
I noticed recently by accident that when I read/write from the oldest raid
disks I have - only one of the tray leds blinks. Of course the led could be
damaged, but rather not, so looking into it I found that both disks in
question return same UUID. So I am concerned now that I don't have any true
RAID there and that there is very important data on those disks.

How is this possible and how to solve it - I would simply add 3rd 500MB disk
to the raid and remove one of the others, but still what is the impact of
this (stupid) coincidence ...

# blkid /dev/sdf
/dev/sdf: PTUUID="13e17ac7" PTTYPE="dos"
                  ^^^^^^^^
# blkid /dev/sdg
/dev/sdg: PTUUID="13e17ac7" PTTYPE="dos"
                  ^^^^^^^^
# blkid /dev/sda
/dev/sda: PTUUID="4184002e" PTTYPE="dos"
# blkid /dev/sdb
/dev/sdb: PTUUID="9570a766" PTTYPE="dos"
# blkid /dev/sdd
/dev/sdd: PTUUID="d898a14c" PTTYPE="dos"
# blkid /dev/sde
/dev/sde: PTUUID="b6346d5e" PTTYPE="dos"


# blkid /dev/sdf1
/dev/sdf1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
TYPE="linux_raid_member" PARTUUID="13e17ac7-01"
# blkid /dev/sdg1
/dev/sdg1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
TYPE="linux_raid_member" PARTUUID="13e17ac7-01"


# smartctl -i /dev/sdf
smartctl 6.6 2016-05-31 r4324 [x86_64-linux-4.12.10] (local build)
Copyright (C) 2002-16, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Western Digital Caviar Blue (SATA)
Device Model:     WDC WD5000AAKS-00A7B0
Serial Number:    WD-WMASY0076274
LU WWN Device Id: 5 0014ee 055d4abcc
Firmware Version: 01.03B01
User Capacity:    500,107,862,016 bytes [500 GB]
Sector Size:      512 bytes logical/physical
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ATA8-ACS (minor revision not indicated)
SATA Version is:  SATA 2.5, 3.0 Gb/s
Local Time is:    Wed Nov  8 22:30:10 2017 CET
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

# smartctl -i /dev/sdg
smartctl 6.6 2016-05-31 r4324 [x86_64-linux-4.12.10] (local build)
Copyright (C) 2002-16, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Seagate Barracuda 7200.10
Device Model:     ST3500630AS
Serial Number:    9QG87WTQ
Firmware Version: 3.AAK
User Capacity:    500,107,862,016 bytes [500 GB]
Sector Size:      512 bytes logical/physical
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ATA/ATAPI-7 (minor revision not indicated)
Local Time is:    Wed Nov  8 22:30:12 2017 CET
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

[toc] | [next] | [standalone]


#188787

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2017-11-08 23:40 +0100
Message-ID<uJHaN-4Te-5@gated-at.bofh.it>
In reply to#188783
On 11/08/17 13:40, deloptes wrote:
> Hi,
> I noticed recently by accident that when I read/write from the oldest raid
> disks I have - only one of the tray leds blinks. Of course the led could be
> damaged, but rather not, so looking into it I found that both disks in
> question return same UUID. So I am concerned now that I don't have any true
> RAID there and that there is very important data on those disks.
> 
> How is this possible and how to solve it - I would simply add 3rd 500MB disk
> to the raid and remove one of the others, but still what is the impact of
> this (stupid) coincidence ...
> 
> # blkid /dev/sdf
> /dev/sdf: PTUUID="13e17ac7" PTTYPE="dos"
>                    ^^^^^^^^
> # blkid /dev/sdg
> /dev/sdg: PTUUID="13e17ac7" PTTYPE="dos"

My file server had a 1.5 TB desktop drive with LUKS and btrfs, created 
with Debian 7.  When I rebuilt my SOHO network with Debian 8, all was 
well.  But, when I rebuilt my SOHO network with Debian 9, I noted 
weirdness.  I don't know if it was Debian, GNU, Linux, LUKS, btrfs, 
smbd, something on the client, PEBKAS, etc..


While trouble-shooting PEBKAS issues is important to me, I have found 
that my attempts at trouble-shooting GNU/Linux issues is usually an 
exercise in futility.  The best I can hope for is finding a way to 
reproduce the issue and filing a bug report.  But as for fixing an 
issue, my best bets is fresh software and known-good hardware.


So, for the file server issue, I built a 2 @ 1.5 TB mdadm RAID1 with 
LUKS and ext4 in another Debian 9 machine, tested it, backed up the file 
server, migrated the data, and then migrated the drives.  The weirdness 
is now gone.  :-)


We'll see what happens when I rebuild with Debian 10...


David

[toc] | [prev] | [next] | [standalone]


#188802

Fromdeloptes <deloptes@gmail.com>
Date2017-11-09 08:20 +0100
Message-ID<uJPi2-1Ha-5@gated-at.bofh.it>
In reply to#188787
David Christensen wrote:

> My file server had a 1.5 TB desktop drive with LUKS and btrfs, created
> with Debian 7.  When I rebuilt my SOHO network with Debian 8, all was
> well.  But, when I rebuilt my SOHO network with Debian 9, I noted
> weirdness.  I don't know if it was Debian, GNU, Linux, LUKS, btrfs,
> smbd, something on the client, PEBKAS, etc..
> 
> 
> While trouble-shooting PEBKAS issues is important to me, I have found
> that my attempts at trouble-shooting GNU/Linux issues is usually an
> exercise in futility.  The best I can hope for is finding a way to
> reproduce the issue and filing a bug report.  But as for fixing an
> issue, my best bets is fresh software and known-good hardware.
> 
> 
> So, for the file server issue, I built a 2 @ 1.5 TB mdadm RAID1 with
> LUKS and ext4 in another Debian 9 machine, tested it, backed up the file
> server, migrated the data, and then migrated the drives.  The weirdness
> is now gone.  :-)
> 
> 
> We'll see what happens when I rebuild with Debian 10...

Hi David and thanks for sharing your experience. However it does not bring
an answer to my question.
I personally never had problems with fixing issues and I must admit the
community is often more helpful than some commercial companies.
Since Etch I never had to reinstall my servers upgrade worked more or less
pretty well. Of course performing an upgrade on a test machine is a must.

What I want to know is if this

# blkid /dev/sdf1
/dev/sdf1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
TYPE="linux_raid_member" PARTUUID="13e17ac7-01"
# blkid /dev/sdg1
/dev/sdg1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
TYPE="linux_raid_member" PARTUUID="13e17ac7-01"

has some effect and I should replace one of the disks. I think the Seagate
is >10y old.

regads

[toc] | [prev] | [next] | [standalone]


#188818

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2017-11-09 19:20 +0100
Message-ID<uJZAK-lt-3@gated-at.bofh.it>
In reply to#188802
On 11/08/17 23:17, deloptes wrote:
> David Christensen wrote:
>> While trouble-shooting PEBKAS issues is important to me, I have found
>> that my attempts at trouble-shooting GNU/Linux issues is usually an
>> exercise in futility.  The best I can hope for is finding a way to
>> reproduce the issue and filing a bug report.  But as for fixing an
>> issue, my best bets is fresh software and known-good hardware.
> Hi David and thanks for sharing your experience. However it does not bring
> an answer to my question.
> I personally never had problems with fixing issues and I must admit the
> community is often more helpful than some commercial companies.
> Since Etch I never had to reinstall my servers upgrade worked more or less
> pretty well. Of course performing an upgrade on a test machine is a must.
> 
> What I want to know is if this
> 
> # blkid /dev/sdf1
> /dev/sdf1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
> TYPE="linux_raid_member" PARTUUID="13e17ac7-01"
> # blkid /dev/sdg1
> /dev/sdg1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
> TYPE="linux_raid_member" PARTUUID="13e17ac7-01"
> 
> has some effect 

Answering that question definitively would involve reviewing the source 
code of all the software and firmware on your computer for anything that 
is affected, directly or indirectly, by UUID's or PARTUUID's.


You should probably start with the source code for whatever RAID 
technology you are using. (What RAID technology are you using?)


Alternatively, look for a FAQ.


STFW.  Ideally, using any warnings or error messages you are seeing.


> and I should replace one of the disks. I think the Seagate
> is >10y old.

Take a look at:

# smartctl --xall /dev/sdg


If you learn smartctl well enough, capture reports on a schedule 
(weekly?), and look for trends, you might be able to predict failure. 
STFW for information on this approach.


Download the bootable CD image of Seagate Seatools and run it:

     https://www.seagate.com/support/downloads/seatools/


The group consensus seems to be:

1.  When you hear the "click of death", failure is imminent.

2.  When you put HDD's on the shelf for long periods, they often fail 
shortly after being returned to service (e.g. within a day).

3.  Hard disk drives last the longest if you leave them in a computer 
and powered up, even if not in use.

4.  All drives fail eventually.  Plan on it and be prepared.


I would estimate a dozen of my HDD's have failed by #1 over the years, 
and another dozen by #2.  I've got a dozen or more on the shelf that 
could end up #2.


David

[toc] | [prev] | [next] | [standalone]


#188829

Fromdeloptes <deloptes@gmail.com>
Date2017-11-09 22:10 +0100
Message-ID<uK2ff-25u-3@gated-at.bofh.it>
In reply to#188818

[Multipart message — attachments visible in raw view] — view raw

Thank you for the kind answer. Here are my toughts

David Christensen wrote:

> Answering that question definitively would involve reviewing the source
> code of all the software and firmware on your computer for anything that
> is affected, directly or indirectly, by UUID's or PARTUUID's.
> 
> 
> You should probably start with the source code for whatever RAID
> technology you are using. (What RAID technology are you using?)
> 

Linux software raid - kernel is 4.12.10

> 
> Alternatively, look for a FAQ.
> 
> 
> STFW.  Ideally, using any warnings or error messages you are seeing.
> 
> 
>> and I should replace one of the disks. I think the Seagate
>> is >10y old.
> 
> Take a look at:
> 
> # smartctl --xall /dev/sdg
> 
> 

This is nothing spectacular - see attachment. In fact I think it does not
write to this disk at all as the partition in the raid setup shows to a
disk with same id.

I think the problem is that blkid reports the same ID for both and that
somehow RAID is using this information, rather than using some of the other
mechanisms - UUID or UDEV Maker/Model/Serial .. which can be found
under /dev/disk

> If you learn smartctl well enough, capture reports on a schedule
> (weekly?), and look for trends, you might be able to predict failure.
> STFW for information on this approach.
> 
> 
> Download the bootable CD image of Seagate Seatools and run it:
> 
> https://www.seagate.com/support/downloads/seatools/
> 

might do that, but I think the problem is in raid itself as it does not
indicate activity on the second disk and blkid reports the same id for two
disks - I really might need to look into the raid code if blkid is used in
any way.

> 
> The group consensus seems to be:
> 
> 1.  When you hear the "click of death", failure is imminent.
> 
> 2.  When you put HDD's on the shelf for long periods, they often fail
> shortly after being returned to service (e.g. within a day).
> 
> 3.  Hard disk drives last the longest if you leave them in a computer
> and powered up, even if not in use.
> 
> 4.  All drives fail eventually.  Plan on it and be prepared.
> 
> 
> I would estimate a dozen of my HDD's have failed by #1 over the years,
> and another dozen by #2.  I've got a dozen or more on the shelf that
> could end up #2.

well those are in server that virtually runs 24/7 and indeed I have replaced
many over the years. In fact most of the old disks are gone. The Seagate is
the oldest there ... the only left, so I think I'll just replace it so that
I may sleep well ... the problem is I don't know which disk is really
writing, might be the Seagate and the WD is not operational ... I think it
is best to be on the safe side :)

regards

[toc] | [prev] | [next] | [standalone]


#188831

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2017-11-09 23:50 +0100
Message-ID<uK3O1-2Si-7@gated-at.bofh.it>
In reply to#188829
On 11/09/17 13:04, deloptes wrote:
> David Christensen wrote:
>> What RAID technology are you using?
> 
> Linux software raid - kernel is 4.12.10

Most people call it 'mdadm', after the command-line tool.  I am running 
the same, but on Debian "stable":

2017-11-09 14:00:32 root@dipsy ~
# dpkg-query --show mdadm
mdadm	3.4-4+b1

2017-11-09 14:00:40 root@dipsy ~
# cat /etc/debian_version
9.2

2017-11-09 14:01:00 root@dipsy ~
# uname -a
Linux dipsy 4.9.0-4-amd64 #1 SMP Debian 4.9.51-1 (2017-09-28) x86_64 
GNU/Linux

2017-11-09 14:01:06 root@dipsy ~
# dpkg-query --show mdadm
mdadm	3.4-4+b1


>> Take a look at:
>>
>> # smartctl --xall /dev/sdg
> 
> This is nothing spectacular - see attachment. 

I'll comment on the information I think I understand...


 > Device Model:     ST3500630AS

I deal with 8 @ ST31500341AS drives, which I believe are of the same 
vintage.  They all seem good.


 > SMART overall-health self-assessment test result: PASSED

That is good.


 > ID# ATTRIBUTE_NAME          FLAGS    VALUE WORST THRESH FAIL RAW_VALUE
 >   1 Raw_Read_Error_Rate     POSR--   105   095   006    -    0
 >  10 Spin_Retry_Count        PO--C-   100   100   097    -    0
 > 187 Reported_Uncorrect      -O--CK   100   100   000    -    0
 > 198 Offline_Uncorrectable   ----C-   100   100   000    -    0

A RAW_VALUE of 0 for these attributes is good.


 > 199 UDMA_CRC_Error_Count    -OSRCK   200   200   000    -    7

7 is low, but the two in my file server are both 0.


Check your cable connections -- they should be fully engaged and not 
loose.  Otherwise swap the cable.  (I wrote a serial number on all of my 
SATA cables with Sharpie and track which cable is where.)


 >   9 Power_On_Hours          -O--CK   034   034   000    -    58404

If 58404 means ~6.6 years (and I think it does), that is a lot of time. 
But, I would not worry based on just this value.


 >   7 Seek_Error_Rate         POSR--   088   060   030    -    747385748
 > 195 Hardware_ECC_Recovered  -O-RC-   064   056   000    -    179548239

I don't know how to interpret these raw values.  STFW I am not alone.


 > SMART Extended Comprehensive Error Log Version: 1 (5 sectors)
 > No Errors Logged

That is good.


> In fact I think it does not
> write to this disk at all as the partition in the raid setup shows to a
> disk with same id.
>
> I think the problem is that blkid reports the same ID for both and that
> somehow RAID is using this information, rather than using some of the other
> mechanisms - UUID or UDEV Maker/Model/Serial .. which can be found
> under /dev/disk

As I understand it, when mdadm creates an array, mdadm puts a metadata 
header into each device that includes identification of the array and 
identification of each member.


When the system boots, mdadm reads /etc/mdadm/mdadm.conf for array 
specifications, scans all devices for mdadm metadata, and then assembles 
the specified arrays using the devices it finds (as best it can).


It looks like you partitioned your drives with one large partition on 
each drive, and then created the array on the partitions.


The matching PTUUID values for both drives, and matching UUID and 
PARTUUID values for both partitions, indicates that one drive was cloned 
onto the other at some point after creating the array.  I agree that 
this is likely a mistake, and is likely to confuse mdadm.


>> If you learn smartctl well enough, capture reports on a schedule
>> (weekly?), and look for trends, you might be able to predict failure.
>> STFW for information on this approach.
>>
>>
>> Download the bootable CD image of Seagate Seatools and run it:
>>
>> https://www.seagate.com/support/downloads/seatools/
>>
> 
> might do that, 

You want that CD as part of your tool kit -- it makes running the SMART 
tests easy, lets you know if everything passed, and helps you understand 
anything that is questionable.


> but I think the problem is in raid itself as it does not
> indicate activity on the second disk and blkid reports the same id for two
> disks - I really might need to look into the raid code if blkid is used in
> any way.

Another alternative to crawling code would be to build another array on 
a pair of USB flash drives using the same process as you used for your 
500 GB drives, and then see what blkid(8) says about the USB drives.


Do you have the console session from when you built the array?


Be sure to keep a console session of any and all mdadm commands you 
issue from now on.


> [the drives] are in server that virtually runs 24/7 and indeed I have replaced
> many over the years. In fact most of the old disks are gone. The Seagate is
> the oldest there ... the only left, so I think I'll just replace it so that
> I may sleep well ... the problem is I don't know which disk is really
> writing, might be the Seagate and the WD is not operational ... I think it
> is best to be on the safe side :)

If the array is working, leave it alone.  Backup/ archive, build a 
replacement array, rsync the data over, validate, migrate services to 
the new array, validate services, and backup again (to validate your 
backup process).  Once the new array has been up and running for a 
while, tear it down and pull the drives.


David

[toc] | [prev] | [next] | [standalone]


#188806

FromTobx <net@tobx.de>
Date2017-11-09 10:40 +0100
Message-ID<uJRtv-3f6-3@gated-at.bofh.it>
In reply to#188783
On 8. Nov 2017, at 22:40, deloptes <deloptes@gmail.com> wrote:

> How is this possible and how to solve it - I would simply add 3rd 500MB disk
> to the raid and remove one of the others, but still what is the impact of
> this (stupid) coincidence …

I would not call it coincidence, what are the odds? There must be a reason for the identical UUIDs (1:1 copy of the disks?, restore of a partition table backup?), but it does not really matter.

I am quite new to this, but I guess you are not assembling your arrays by partition UUID, this should probably not work with identical UUIDs. As far as I understand most of the time the UUID of the array in the disks superblock is used for assembling and the partition UUID does not matter. Someone might be able to confirm this.

Of course there could be other parts of your system that use the partition UUID, but then again, if no issues occurred yet with two identical UUIDs, this is probably not the case, but this is hard to say.

You can change your partition UUID with fdisk (press x for extra functionality). An easy way to create a random UUID is:

$ cat /proc/sys/kernel/random/uuid

If you have the chance to test this, I would give it a try.

On 9. Nov 2017, at 08:17, deloptes <deloptes@gmail.com> wrote:

> has some effect and I should replace one of the disks. I think the Seagate
> is >10y old.

Then I guess you should replace the disk in anyway.

Cheers,
Tobi

[toc] | [prev] | [next] | [standalone]


#188830

Fromdeloptes <deloptes@gmail.com>
Date2017-11-09 22:20 +0100
Message-ID<uK2oV-28A-1@gated-at.bofh.it>
In reply to#188806
Tobx wrote:

> You can change your partition UUID with fdisk (press x for extra
> functionality). An easy way to create a random UUID is:
> 
> $ cat /proc/sys/kernel/random/uuid
> 
> If you have the chance to test this, I would give it a try.

Thanks - good idea - I guess I'll first do a backup :)
Then I'll give it a try

regards

[toc] | [prev] | [next] | [standalone]


#188833

FromJoe Pfeiffer <pfeiffer@cs.nmsu.edu>
Date2017-11-10 00:00 +0100
Message-ID<uK3XH-2Vo-1@gated-at.bofh.it>
In reply to#188783
deloptes <deloptes@gmail.com> writes:

> Hi,
> I noticed recently by accident that when I read/write from the oldest raid
> disks I have - only one of the tray leds blinks. Of course the led could be
> damaged, but rather not, so looking into it I found that both disks in
> question return same UUID. So I am concerned now that I don't have any true
> RAID there and that there is very important data on those disks.
>
> How is this possible and how to solve it - I would simply add 3rd 500MB disk
> to the raid and remove one of the others, but still what is the impact of
> this (stupid) coincidence ...
>
> # blkid /dev/sdf
> /dev/sdf: PTUUID="13e17ac7" PTTYPE="dos"
>                   ^^^^^^^^
> # blkid /dev/sdg
> /dev/sdg: PTUUID="13e17ac7" PTTYPE="dos"
>                   ^^^^^^^^
> # blkid /dev/sda
> /dev/sda: PTUUID="4184002e" PTTYPE="dos"
> # blkid /dev/sdb
> /dev/sdb: PTUUID="9570a766" PTTYPE="dos"
> # blkid /dev/sdd
> /dev/sdd: PTUUID="d898a14c" PTTYPE="dos"
> # blkid /dev/sde
> /dev/sde: PTUUID="b6346d5e" PTTYPE="dos"
>
>
> # blkid /dev/sdf1
> /dev/sdf1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
> TYPE="linux_raid_member" PARTUUID="13e17ac7-01"
> # blkid /dev/sdg1
> /dev/sdg1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
> TYPE="linux_raid_member" PARTUUID="13e17ac7-01"

This is normal.  It's the identical UUIDs that tell the system that the
partitions go into the same RAID array.

Here's what I see when I look at my RAID disks:

/dev/sda2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae" UUID_SUB="48b56869-6f19-21b9-283f-3eee3ac90cf8" LABEL="snowball:1" TYPE="linux_raid_member" PARTUUID="3bb3729a-528b-4384-b6a5-b6d9e148ed2a"
/dev/sdb2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae" UUID_SUB="1f48f805-4173-78cd-1f52-957920f66335" LABEL="snowball:1" TYPE="linux_raid_member" PARTUUID="1bdd3893-9346-49d2-8292-a61075ad0c5e"

and here's the relevant line in my /etc/mdadm/mdadm.conf
ARRAY /dev/md/1  metadata=1.2 UUID=67d3c233:96a0737c:5f88ed9b:936ea3ae name=snowball:1

But...  if this data is that important, you should be running backups.
RAID is to keep you running if a disk fails, it isn't to keep you from
losing data.

[toc] | [prev] | [next] | [standalone]


#188835

Fromdeloptes <deloptes@gmail.com>
Date2017-11-10 09:20 +0100
Message-ID<uKcHD-J9-1@gated-at.bofh.it>
In reply to#188833
Hi Joe,

thank you for the mesage

Joe Pfeiffer wrote:

> This is normal.  It's the identical UUIDs that tell the system that the
> partitions go into the same RAID array.
> 
> Here's what I see when I look at my RAID disks:
> 
> /dev/sda2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae"
> UUID_SUB="48b56869-6f19-21b9-283f-3eee3ac90cf8" LABEL="snowball:1"
> TYPE="linux_raid_member" PARTUUID="3bb3729a-528b-4384-b6a5-b6d9e148ed2a"
> /dev/sdb2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae"
> UUID_SUB="1f48f805-4173-78cd-1f52-957920f66335" LABEL="snowball:1"
> TYPE="linux_raid_member" PARTUUID="1bdd3893-9346-49d2-8292-a61075ad0c5e"
> 

you see in your case PARTUUID is different for both members. In my case it
is identical and this is what is bothering me

> and here's the relevant line in my /etc/mdadm/mdadm.conf
> ARRAY /dev/md/1  metadata=1.2 UUID=67d3c233:96a0737c:5f88ed9b:936ea3ae
> name=snowball:1
> 

It looks like the new style raid (I don't recall in which version it was
introduced). However this raid was created ~12y ago without metadata.

> But...  if this data is that important, you should be running backups.
> RAID is to keep you running if a disk fails, it isn't to keep you from
> losing data.

Indeed this is true - I make backups but not that often as data changes not
that often, however it might be good idea to run on regular bases.

I guess I'll have to sit over the weekend and make a plan.

thanks

regards

[toc] | [prev] | [next] | [standalone]


#188852

FromJoe Pfeiffer <pfeiffer@cs.nmsu.edu>
Date2017-11-10 18:10 +0100
Message-ID<uKkYy-6dO-5@gated-at.bofh.it>
In reply to#188835
deloptes <deloptes@gmail.com> writes:

> Hi Joe,
>
> thank you for the mesage
>
> Joe Pfeiffer wrote:
>
>> This is normal.  It's the identical UUIDs that tell the system that the
>> partitions go into the same RAID array.
>> 
>> Here's what I see when I look at my RAID disks:
>> 
>> /dev/sda2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae"
>> UUID_SUB="48b56869-6f19-21b9-283f-3eee3ac90cf8" LABEL="snowball:1"
>> TYPE="linux_raid_member" PARTUUID="3bb3729a-528b-4384-b6a5-b6d9e148ed2a"
>> /dev/sdb2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae"
>> UUID_SUB="1f48f805-4173-78cd-1f52-957920f66335" LABEL="snowball:1"
>> TYPE="linux_raid_member" PARTUUID="1bdd3893-9346-49d2-8292-a61075ad0c5e"
>> 
>
> you see in your case PARTUUID is different for both members. In my case it
> is identical and this is what is bothering me

It's my understanding that PTUUID on a disk using an MBR corresponds to
the UUID on a disk using a GPT, not to PARTUUID (I don't know what on an
MBR-based disk would correspond to PARTUUID, if anything).

[toc] | [prev] | [next] | [standalone]


#188859

Fromdeloptes <deloptes@gmail.com>
Date2017-11-10 20:50 +0100
Message-ID<uKntn-7Bp-5@gated-at.bofh.it>
In reply to#188852
Joe Pfeiffer wrote:

> It's my understanding that PTUUID on a disk using an MBR corresponds to
> the UUID on a disk using a GPT, not to PARTUUID (I don't know what on an
> MBR-based disk would correspond to PARTUUID, if anything).

# blkid /dev/sdf1
/dev/sdf1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
TYPE="linux_raid_member" PARTUUID="13e17ac7-01"
# blkid /dev/sdg1
/dev/sdg1: UUID="5427071b-25c8-fff8-476d-ff8c9852b714"
TYPE="linux_raid_member" PARTUUID="13e17ac7-01"

# blkid /dev/sdf
/dev/sdf: PTUUID="13e17ac7" PTTYPE="dos"
                  ^^^^^^^^
# blkid /dev/sdg
/dev/sdg: PTUUID="13e17ac7" PTTYPE="dos"
                  ^^^^^^^^

Correct and exactly this is the subject of my concern. What blkid is
reporting for the disk matches (correctly) to the PARTUUID - which as it
looks like tells us which partition it is.

I can't believe that this blkid library might be used in raid as if you read
the description it can easily report identical IDs.

Might be really worth examining the code - the last thing I wanted to do
now. Perhaps I ask on the kernel list.

thanks and regards

[toc] | [prev] | [next] | [standalone]


#188910

FromPascal Hambourg <pascal@plouf.fr.eu.org>
Date2017-11-12 17:20 +0100
Message-ID<uL39f-2tm-5@gated-at.bofh.it>
In reply to#188852
Le 10/11/2017 à 17:46, Joe Pfeiffer a écrit :
> deloptes <deloptes@gmail.com> writes:
> 
>> you see in your case PARTUUID is different for both members. In my case it
>> is identical and this is what is bothering me
> 
> It's my understanding that PTUUID on a disk using an MBR corresponds to
> the UUID on a disk using a GPT, not to PARTUUID (I don't know what on an
> MBR-based disk would correspond to PARTUUID, if anything).

In the blkid syntax :
PTUUID = partition table UUID (in the partition table).
PARTUUID = partition UUID (in the partition table).
UUID = filesystem or other contents UUID (in the partition data).

There are no PTUUID nor PARTUUID in the MSDOS partition table format. 
There is only a 32-bit "disk identifier" field in the MBR, which can be 
displayed by fdisk. blkid uses it as a poor-man's PTUUID. Also, since 
version 3.8, the kernel combines the MSDOS disk identifier (PTUUID) and 
the partition numbers to create fake partition UUIDs (PARTUUIDs).

The GPT partition table format has real 128-bit independent PTUUID and 
PARTUUIDs.

[toc] | [prev] | [next] | [standalone]


#188861

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2017-11-10 21:10 +0100
Message-ID<uKnMJ-7Yj-1@gated-at.bofh.it>
In reply to#188835
On 11/10/17 00:12, deloptes wrote:
> this raid was created ~12y ago without metadata.

https://www.psychologytoday.com/basics/self-harm


;-)

David

[toc] | [prev] | [next] | [standalone]


#188867

Fromdeloptes <deloptes@gmail.com>
Date2017-11-10 23:30 +0100
Message-ID<uKpYe-SE-9@gated-at.bofh.it>
In reply to#188861
David Christensen wrote:

> On 11/10/17 00:12, deloptes wrote:
>> this raid was created ~12y ago without metadata.
> 
> https://www.psychologytoday.com/basics/self-harm
> 
> 
> ;-)
> 
> David

appreciated :D - one that respects sarcasm :D

[toc] | [prev] | [next] | [standalone]


#188909

FromPascal Hambourg <pascal@plouf.fr.eu.org>
Date2017-11-12 17:00 +0100
Message-ID<uL2PT-27l-1@gated-at.bofh.it>
In reply to#188835
Le 10/11/2017 à 09:12, deloptes a écrit :
> 
> Joe Pfeiffer wrote:
>>
>> Here's what I see when I look at my RAID disks:
>>
>> /dev/sda2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae"
>> UUID_SUB="48b56869-6f19-21b9-283f-3eee3ac90cf8" LABEL="snowball:1"
>> TYPE="linux_raid_member" PARTUUID="3bb3729a-528b-4384-b6a5-b6d9e148ed2a"
>> /dev/sdb2: UUID="67d3c233-96a0-737c-5f88-ed9b936ea3ae"
>> UUID_SUB="1f48f805-4173-78cd-1f52-957920f66335" LABEL="snowball:1"
>> TYPE="linux_raid_member" PARTUUID="1bdd3893-9346-49d2-8292-a61075ad0c5e"
> 
> you see in your case PARTUUID is different for both members. In my case it
> is identical and this is what is bothering me
> 
>> and here's the relevant line in my /etc/mdadm/mdadm.conf
>> ARRAY /dev/md/1  metadata=1.2 UUID=67d3c233:96a0737c:5f88ed9b:936ea3ae
>> name=snowball:1
> 
> It looks like the new style raid

Indeed, superblock format 1.x. In addition to the "array UUID" which is 
common to all members of the array, it adds a specific "device UUID" for 
each member. blkid labels it "UUID_SUB".

> However this raid was created ~12y ago without metadata.

I don't think so. If the array was created without metadata, blkid would 
not report the members as TYPE="linux_raid_member". It was rather 
probably created with the old metadata format 0.90. You can check with

mdadm --examine /dev/sd[fg]1

[toc] | [prev] | [next] | [standalone]


#188912

Fromdeloptes <deloptes@gmail.com>
Date2017-11-12 18:00 +0100
Message-ID<uL3LX-2Gu-9@gated-at.bofh.it>
In reply to#188909
Thanks Pascal,

Pascal Hambourg wrote:

> Le 10/11/2017 à 09:12, deloptes a écrit :
>> 

>> 
>> It looks like the new style raid
> 
> Indeed, superblock format 1.x. In addition to the "array UUID" which is
> common to all members of the array, it adds a specific "device UUID" for
> each member. blkid labels it "UUID_SUB".
> 
>> However this raid was created ~12y ago without metadata.
> 
> I don't think so. If the array was created without metadata, blkid would
> not report the members as TYPE="linux_raid_member". It was rather
> probably created with the old metadata format 0.90. You can check with
> 
> mdadm --examine /dev/sd[fg]1

Indeed   Creation Time : Mon Jul 26 21:49:51 2010
means data was migrated to those disks, which were assembled 2010 - data
dates back to 2002. Seagate is definitely older than 2010.

Anyway, thanks for the snip - useful!

So if you know if in this context I have real RAID - I mean data is written
to both drives? It would be nice, though I decided already to replace
2x1G+2x0.5G for 2x2G - wife blessed the budget :)

regards

[toc] | [prev] | [next] | [standalone]


#188917

FromPascal Hambourg <pascal@plouf.fr.eu.org>
Date2017-11-12 19:50 +0100
Message-ID<uL5up-3NL-7@gated-at.bofh.it>
In reply to#188912
Le 12/11/2017 à 17:57, deloptes a écrit :
> 
> So if you know if in this context I have real RAID - I mean data is written
> to both drives?

Sorry, I cannot tell. Insufficient data. You can check in /proc/mdstat.

[toc] | [prev] | [next] | [standalone]


#188919

Fromdeloptes <deloptes@gmail.com>
Date2017-11-12 20:20 +0100
Message-ID<uL5Xs-4cE-7@gated-at.bofh.it>
In reply to#188917
Pascal Hambourg wrote:

> /proc/mdstat

this looks good, however I tried writing to the disk and only one disk led
indicates writes, so I think it is not writing on both disks.

Anyway I will replace those disks

thanks and regards

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web