Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.kernel > #83175 > unrolled thread
| Started by | Stefan <debian@simg.de> |
|---|---|
| First post | 2024-07-24 17:40 +0200 |
| Last post | 2024-09-25 21:10 +0200 |
| Articles | 4 — 2 participants |
Back to article view | Back to linux.debian.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Bug#1076372: Re.: linux-image-6.5.0-0.deb12.4-amd64: ext4 file corruption with newer kernels Stefan <debian@simg.de> - 2024-07-24 17:40 +0200
Bug#1076372: Re.: linux-image-6.5.0-0.deb12.4-amd64: ext4 file corruption with newer kernels Stefan <debian@simg.de> - 2024-07-26 11:00 +0200
Bug#1076372: Re.: linux-image-6.5.0-0.deb12.4-amd64: ext4 file corruption with newer kernels Stefan <debian@simg.de> - 2024-07-27 02:40 +0200
Bug#1076372: Re.: linux-image-6.5.0-0.deb12.4-amd64: ext4 file corruption with newer kernels Salvatore Bonaccorso <carnil@debian.org> - 2024-09-25 21:10 +0200
| From | Stefan <debian@simg.de> |
|---|---|
| Date | 2024-07-24 17:40 +0200 |
| Subject | Bug#1076372: Re.: linux-image-6.5.0-0.deb12.4-amd64: ext4 file corruption with newer kernels |
| Message-ID | <J3MfT-1Ar-1@gated-at.bofh.it> |
Hi, I ran a few other tests: 1. tried package "linux-image-6.1.0-22-amd64": works 2. tried package "linux-image-6.9.7+bpo-amd64": does not work 3. installed firmware from https://git.kernel.org/pub/scm/linux/kernel/git/firmware/linux-firmware.git/ (the Debian package is a little bit old): 6.1.0 works, 6.9.7 does not work 4. Updated the firmware/BIOS (It is one of the new chipset-less X600 mainboards AM5/Ryzen released by Asrock. Firmware seems to be a little bit unfinished because version changed from 1.43 to 4.03 within 5 weeks): 6.1.0 works, 6.9.7 does not work The size of the damaged portion in the files is usually (but not always) 128kB, not 256kB as reported in the initial message. I still did not detected any file system corruption, just the files are damaged. (maybe random/luck) Here is the `lspci` output of the NVME under 6.9.7: ---------- 02:00.0 Non-Volatile memory controller: Shenzhen Longsys Electronics Co., Ltd. Lexar NM790 NVME SSD (DRAM-less) (rev 01) (prog-if 02 [NVM Express]) Subsystem: Shenzhen Longsys Electronics Co., Ltd. Lexar NM790 NVME SSD (DRAM-less) Control: I/O- Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr- Stepping- SERR- FastB2B- DisINTx+ Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- <MAbort- >SERR- <PERR- INTx- Latency: 0, Cache Line Size: 64 bytes Interrupt: pin A routed to IRQ 35 IOMMU group: 14 Region 0: Memory at f6e00000 (64-bit, non-prefetchable) [size=16K] Capabilities: [40] Power Management version 3 Flags: PMEClk- DSI- D1- D2- AuxCurrent=0mA PME(D0-,D1-,D2-,D3hot-,D3cold-) Status: D0 NoSoftRst+ PME-Enable- DSel=0 DScale=0 PME- Capabilities: [50] MSI: Enable- Count=1/32 Maskable+ 64bit+ Address: 0000000000000000 Data: 0000 Masking: 00000000 Pending: 00000000 Capabilities: [70] Express (v2) Endpoint, MSI 1f DevCap: MaxPayload 512 bytes, PhantFunc 0, Latency L0s unlimited, L1 unlimited ExtTag- AttnBtn- AttnInd- PwrInd- RBE+ FLReset+ SlotPowerLimit 75W DevCtl: CorrErr+ NonFatalErr+ FatalErr+ UnsupReq+ RlxdOrd- ExtTag- PhantFunc- AuxPwr- NoSnoop+ FLReset- MaxPayload 256 bytes, MaxReadReq 512 bytes DevSta: CorrErr+ NonFatalErr- FatalErr- UnsupReq+ AuxPwr+ TransPend- LnkCap: Port #0, Speed 16GT/s, Width x4, ASPM L1, Exit Latency L1 <64us ClockPM+ Surprise- LLActRep- BwNot- ASPMOptComp+ LnkCtl: ASPM L1 Enabled; RCB 64 bytes, Disabled- CommClk+ ExtSynch- ClockPM- AutWidDis- BWInt- AutBWInt- LnkSta: Speed 16GT/s, Width x4 TrErr- Train- SlotClk+ DLActive- BWMgmt- ABWMgmt- DevCap2: Completion Timeout: Range ABCD, TimeoutDis+ NROPrPrP- LTR+ 10BitTagComp+ 10BitTagReq- OBFF Via message, ExtFmt- EETLPPrefix- EmergencyPowerReduction Not Supported, EmergencyPowerReductionInit- FRS- TPHComp- ExtTPHComp- AtomicOpsCap: 32bit- 64bit- 128bitCAS- DevCtl2: Completion Timeout: 50us to 50ms, TimeoutDis- LTR+ 10BitTagReq- OBFF Disabled, AtomicOpsCtl: ReqEn- LnkCap2: Supported Link Speeds: 2.5-16GT/s, Crosslink- Retimer+ 2Retimers+ DRS- LnkCtl2: Target Link Speed: 16GT/s, EnterCompliance- SpeedDis- Transmit Margin: Normal Operating Range, EnterModifiedCompliance- ComplianceSOS- Compliance Preset/De-emphasis: -6dB de-emphasis, 0dB preshoot LnkSta2: Current De-emphasis Level: -3.5dB, EqualizationComplete+ EqualizationPhase1+ EqualizationPhase2+ EqualizationPhase3+ LinkEqualizationRequest- Retimer- 2Retimers- CrosslinkRes: Upstream Port Capabilities: [b0] MSI-X: Enable+ Count=17 Masked- Vector table: BAR=0 offset=00003000 PBA: BAR=0 offset=00002000 Capabilities: [100 v2] Advanced Error Reporting UESta: DLP- SDES- TLP- FCP- CmpltTO- CmpltAbrt- UnxCmplt- RxOF- MalfTLP- ECRC- UnsupReq- ACSViol- UEMsk: DLP- SDES- TLP- FCP- CmpltTO- CmpltAbrt- UnxCmplt- RxOF- MalfTLP- ECRC- UnsupReq- ACSViol- UESvrt: DLP+ SDES+ TLP- FCP+ CmpltTO- CmpltAbrt- UnxCmplt- RxOF+ MalfTLP+ ECRC- UnsupReq- ACSViol- CESta: RxErr- BadTLP- BadDLLP- Rollover- Timeout- AdvNonFatalErr- CEMsk: RxErr- BadTLP- BadDLLP- Rollover- Timeout- AdvNonFatalErr+ AERCap: First Error Pointer: 00, ECRCGenCap+ ECRCGenEn- ECRCChkCap+ ECRCChkEn- MultHdrRecCap- MultHdrRecEn- TLPPfxPres- HdrLogCap- HeaderLog: 00000000 00000000 00000000 0000000002:00.0 Non-Volatile memory controller: Shenzhen Longsys Electronics Co., Ltd. Lexar NM790 NVME SSD (DRAM-less) (rev 01) (prog-if 02 [NVM Express]) Subsystem: Shenzhen Longsys Electronics Co., Ltd. Lexar NM790 NVME SSD (DRAM-less) Control: I/O- Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr- Stepping- SERR- FastB2B- DisINTx+ Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- <MAbort- >SERR- <PERR- INTx- Latency: 0, Cache Line Size: 64 bytes Interrupt: pin A routed to IRQ 35 IOMMU group: 14 Region 0: Memory at f6e00000 (64-bit, non-prefetchable) [size=16K] Capabilities: [40] Power Management version 3 Flags: PMEClk- DSI- D1- D2- AuxCurrent=0mA PME(D0-,D1-,D2-,D3hot-,D3cold-) Status: D0 NoSoftRst+ PME-Enable- DSel=0 DScale=0 PME- Capabilities: [50] MSI: Enable- Count=1/32 Maskable+ 64bit+ Address: 0000000000000000 Data: 0000 Masking: 00000000 Pending: 00000000 Capabilities: [70] Express (v2) Endpoint, MSI 1f DevCap: MaxPayload 512 bytes, PhantFunc 0, Latency L0s unlimited, L1 unlimited ExtTag- AttnBtn- AttnInd- PwrInd- RBE+ FLReset+ SlotPowerLimit 75W DevCtl: CorrErr+ NonFatalErr+ FatalErr+ UnsupReq+ RlxdOrd- ExtTag- PhantFunc- AuxPwr- NoSnoop+ FLReset- MaxPayload 256 bytes, MaxReadReq 512 bytes DevSta: CorrErr+ NonFatalErr- FatalErr- UnsupReq+ AuxPwr+ TransPend- LnkCap: Port #0, Speed 16GT/s, Width x4, ASPM L1, Exit Latency L1 <64us ClockPM+ Surprise- LLActRep- BwNot- ASPMOptComp+ LnkCtl: ASPM L1 Enabled; RCB 64 bytes, Disabled- CommClk+ ExtSynch- ClockPM- AutWidDis- BWInt- AutBWInt- LnkSta: Speed 16GT/s, Width x4 TrErr- Train- SlotClk+ DLActive- BWMgmt- ABWMgmt- DevCap2: Completion Timeout: Range ABCD, TimeoutDis+ NROPrPrP- LTR+ 10BitTagComp+ 10BitTagReq- OBFF Via message, ExtFmt- EETLPPrefix- EmergencyPowerReduction Not Supported, EmergencyPowerReductionInit- FRS- TPHComp- ExtTPHComp- AtomicOpsCap: 32bit- 64bit- 128bitCAS- DevCtl2: Completion Timeout: 50us to 50ms, TimeoutDis- LTR+ 10BitTagReq- OBFF Disabled, AtomicOpsCtl: ReqEn- LnkCap2: Supported Link Speeds: 2.5-16GT/s, Crosslink- Retimer+ 2Retimers+ DRS- LnkCtl2: Target Link Speed: 16GT/s, EnterCompliance- SpeedDis- Transmit Margin: Normal Operating Range, EnterModifiedCompliance- ComplianceSOS- Compliance Preset/De-emphasis: -6dB de-emphasis, 0dB preshoot LnkSta2: Current De-emphasis Level: -3.5dB, EqualizationComplete+ EqualizationPhase1+ EqualizationPhase2+ EqualizationPhase3+ LinkEqualizationRequest- Retimer- 2Retimers- CrosslinkRes: Upstream Port Capabilities: [b0] MSI-X: Enable+ Count=17 Masked- Vector table: BAR=0 offset=00003000 PBA: BAR=0 offset=00002000 Capabilities: [100 v2] Advanced Error Reporting UESta: DLP- SDES- TLP- FCP- CmpltTO- CmpltAbrt- UnxCmplt- RxOF- MalfTLP- ECRC- UnsupReq- ACSViol- UEMsk: DLP- SDES- TLP- FCP- CmpltTO- CmpltAbrt- UnxCmplt- RxOF- MalfTLP- ECRC- UnsupReq- ACSViol- UESvrt: DLP+ SDES+ TLP- FCP+ CmpltTO- CmpltAbrt- UnxCmplt- RxOF+ MalfTLP+ ECRC- UnsupReq- ACSViol- CESta: RxErr- BadTLP- BadDLLP- Rollover- Timeout- AdvNonFatalErr- CEMsk: RxErr- BadTLP- BadDLLP- Rollover- Timeout- AdvNonFatalErr+ AERCap: First Error Pointer: 00, ECRCGenCap+ ECRCGenEn- ECRCChkCap+ ECRCChkEn- MultHdrRecCap- MultHdrRecEn- TLPPfxPres- HdrLogCap- HeaderLog: 00000000 00000000 00000000 00000000 Capabilities: [148 v1] Device Serial Number 00-00-00-00-00-00-00-00 Capabilities: [158 v1] Power Budgeting <?> Capabilities: [168 v1] Alternative Routing-ID Interpretation (ARI) ARICap: MFVC- ACS+, Next Function: 0 ARICtl: MFVC- ACS-, Function Group: 0 Capabilities: [178 v1] Secondary PCI Express LnkCtl3: LnkEquIntrruptEn- PerformEqu- LaneErrStat: 0 Capabilities: [198 v1] Physical Layer 16.0 GT/s <?> Capabilities: [1bc v1] Lane Margining at the Receiver <?> Capabilities: [220 v1] Latency Tolerance Reporting Max snoop latency: 1048576ns Max no snoop latency: 1048576ns Capabilities: [228 v1] L1 PM Substates L1SubCap: PCI-PM_L1.2+ PCI-PM_L1.1+ ASPM_L1.2+ ASPM_L1.1+ L1_PM_Substates+ PortCommonModeRestoreTime=10us PortTPowerOnTime=1000us L1SubCtl1: PCI-PM_L1.2- PCI-PM_L1.1- ASPM_L1.2- ASPM_L1.1- T_CommonMode=0us LTR1.2_Threshold=32768ns L1SubCtl2: T_PwrOn=1000us Capabilities: [238 v1] Vendor Specific Information: ID=0002 Rev=4 Len=100 <?> Capabilities: [338 v1] Vendor Specific Information: ID=0001 Rev=1 Len=038 <?> Capabilities: [370 v1] Data Link Feature <?> Kernel driver in use: nvme Kernel modules: nvme Capabilities: [148 v1] Device Serial Number 00-00-00-00-00-00-00-00 Capabilities: [158 v1] Power Budgeting <?> Capabilities: [168 v1] Alternative Routing-ID Interpretation (ARI) ARICap: MFVC- ACS+, Next Function: 0 ARICtl: MFVC- ACS-, Function Group: 0 Capabilities: [178 v1] Secondary PCI Express LnkCtl3: LnkEquIntrruptEn- PerformEqu- LaneErrStat: 0 Capabilities: [198 v1] Physical Layer 16.0 GT/s <?> Capabilities: [1bc v1] Lane Margining at the Receiver <?> Capabilities: [220 v1] Latency Tolerance Reporting Max snoop latency: 1048576ns Max no snoop latency: 1048576ns Capabilities: [228 v1] L1 PM Substates L1SubCap: PCI-PM_L1.2+ PCI-PM_L1.1+ ASPM_L1.2+ ASPM_L1.1+ L1_PM_Substates+ PortCommonModeRestoreTime=10us PortTPowerOnTime=1000us L1SubCtl1: PCI-PM_L1.2- PCI-PM_L1.1- ASPM_L1.2- ASPM_L1.1- T_CommonMode=0us LTR1.2_Threshold=32768ns L1SubCtl2: T_PwrOn=1000us Capabilities: [238 v1] Vendor Specific Information: ID=0002 Rev=4 Len=100 <?> Capabilities: [338 v1] Vendor Specific Information: ID=0001 Rev=1 Len=038 <?> Capabilities: [370 v1] Data Link Feature <?> Kernel driver in use: nvme Kernel modules: nvme ---------- Parameters seem to be the same ones as under 6.1.0. Regards Stefan
[toc] | [next] | [standalone]
| From | Stefan <debian@simg.de> |
|---|---|
| Date | 2024-07-26 11:00 +0200 |
| Message-ID | <J4oXT-r4z-1@gated-at.bofh.it> |
| In reply to | #83175 |
The complete sentence is: oldest non-working kernel is 6.3.7 (package linux-image-6.3.0-1-amd64), 6.3.5 (latest version of package package linux-image-6.3.0-0-amd64) works. Am 26.07.24 um 10:45 schrieb Stefan: > Hi, > > oldest non-working kernel is 6.3.7 (package linux-image-6.3.0-1-amd64), > 6.3.5 (latest version of package package linux-image-6.3.0-0-amd64) > > The file corruption does not occur if the file is still in buffer. > Furthermore the corruption do not occur during reading, i.e. damaged > files stay damaged after reboot. > > In kernels 6.3.* to early 6.5.* another strange error occurs when nvme > module is loaded: "Device not ready; aborting initialisation. CSTS=0x0." > > Waiting a while and reloading nvme module helps. Sometimes several tries > are required. See the attached screenshot. > > Regards Stefan > > > Am 24.07.24 um 17:53 schrieb Diederik de Haas: >> Control: found -1 6.9.7-1~bpo12+1 >> >> On Wednesday, 24 July 2024 17:24:44 CEST Stefan wrote: >>> I ran a few other tests: >>> >>> 1. tried package "linux-image-6.1.0-22-amd64": works >>> 2. tried package "linux-image-6.9.7+bpo-amd64": does not work >> >> Via https://snapshot.debian.org/package/linux-signed-amd64/ you can find >> older (bpo) kernels then 6.5.10+1~bpo12+1 (where it was initially filed >> against) and it's useful to know what the last kernel version after 6.1 >> was where it worked properly and the first one where it broke.
[toc] | [prev] | [next] | [standalone]
| From | Stefan <debian@simg.de> |
|---|---|
| Date | 2024-07-27 02:40 +0200 |
| Message-ID | <J4DDz-BaX-7@gated-at.bofh.it> |
| In reply to | #83220 |
Hi Salvatore, Am 26.07.24 um 23:09 schrieb Salvatore Bonaccorso: > Now that you were able to pinpoint two versions which are very close > together, would you be able to bisect the changes between 6.3.5 and > 6.3.7 to find the breaking commit upstream in that 6.3.y series? not without spending more time I can effort ATM and without finding a better test procedure (currently 1-2h per run). I'm not convinced that it is an ext4 bug. The corruption may also occur when the data is written to the nvm. The hardware (chipset-less AM5 mainboard) is quite new, the only vendor (Asrock) probably does not care about Linux and maybe AMD did not tested is carefully enough. Thus, there a quite a lot of changes from 6.3.5 -> 6.3.7 that could cause the bug. I'll try to find out more next week (other ways to reproduce the error, other hardware, ...), especially whether it is a more general issue or hardware specific. Regards Stefan
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2024-09-25 21:10 +0200 |
| Message-ID | <JqFyF-eNkT-9@gated-at.bofh.it> |
| In reply to | #83229 |
Hi Stefan, On Wed, Aug 14, 2024 at 01:22:05AM +0200, Stefan wrote: > Hi Salvatore, > > sorry, I had not the time to run all tests I planned. > > Here is what I found out so far: > > * The errors can be reproduced with the program `f3`, see > https://fight-flash-fraud.readthedocs.io/en/latest/index.html. Errors > are reported as "overwritten sectors", which probably means that they > are address errors. The strange thing is, that the errors seem to > disappear after the program runs a while (in contrast to the older test; > log file enclosed). I tried to induce the/more errors (`stress-ng > --cpu`, deleting files while f3 is running), but had no success. (It may > be an issue with the pattern generator of `f3`.) > > * The errors cannot be reproduced on other systems (tested 3 ones), so > it is hardware specific. > > * The bug is not fixed in 6.10 kernels > > I had not sufficient time to test another nvm. This would reveal whether > it is a problem with the mainboard+CPU or with the nvm. Did you had a chance to do these further tests? In meanwhile we have retitled the bug to make it more specific to the Longsys/Lexar NM790 NVMe drive. Regards, Salvatore
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.kernel
csiph-web