Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #77890 > unrolled thread

Bug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down"

Started byDiederik de Haas <didi.debian@cknow.org>
First post2023-01-11 22:00 +0100
Last post2023-01-12 16:10 +0100
Articles 5 — 2 participants

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down" Diederik de Haas <didi.debian@cknow.org> - 2023-01-11 22:00 +0100
    Bug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down" Julian <julian.g@posteo.de> - 2023-01-12 12:20 +0100
      Bug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down" Diederik de Haas <didi.debian@cknow.org> - 2023-01-12 15:00 +0100
      Bug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down" Diederik de Haas <didi.debian@cknow.org> - 2023-01-12 15:50 +0100
        Bug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down" Julian <julian.g@posteo.de> - 2023-01-12 16:10 +0100

#77890 — Bug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down"

FromDiederik de Haas <didi.debian@cknow.org>
Date2023-01-11 22:00 +0100
SubjectBug#1028309: linux-image-6.0.0-6-amd64: Regression in Kernel 6.0: System partially freezes with "nvme controller is down"
Message-ID<FMQmu-hhwA-3@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

On Wednesday, 11 January 2023 19:28:25 CET Julian Groß wrote:
> On Mon, 09 Jan 2023 14:09:30 +0100 Diederik de Haas
> <didi.debian@cknow.org> wrote:
>  > https://wiki.debian.org/DebianKernel/GitBisect describes a procedure
>  > to find the exact commit which introduced the issue you reported, but it's
>  > often faster to first narrow down the range using snapshot.d.o.
> 
> The way I understand the `git bisect`, and with the issue taking
> sometimes days to happen, I will be sitting on this for months by the way.

Yep, that would/could be the consequence, which I can fully understand is not 
desirable or very useful.
Having the exact offending commit is ideal, but not a '100%' requirement.
As you've determined that it already happened with 6.0~rc7 and not with 
5.19.x, that's already a reasonably small range (likely introduced in the 6.0 
merge window).

So the next thing to do, is present the issue to the relevant upstream 
maintainers. Searching for "nvme controller is down" brought up another bug 
(but that happened pretty instantly) and in there the request was made to 
report the issue (via email) to the linux-pci@vger.kernel.org and 
linux-nvme@lists.infradead.org lists.
And that seems like the next best step for you too.

Use the information you provided in your initial bug report and add the extra 
findings (ie 6.1-rc7) to that too.
When you've done that, please inform this bug report where you did that so 
that we can track it's progress too.

Cheers,
  Diederik

[toc] | [next] | [standalone]


#77902

FromJulian <julian.g@posteo.de>
Date2023-01-12 12:20 +0100
Message-ID<FN3MK-ht2m-9@gated-at.bofh.it>
In reply to#77890

[Multipart message — attachments visible in raw view] — view raw

On Wed, 11 Jan 2023 21:49:45 +0100 Diederik de Haas <didi.debian@cknow.org> wrote:
> So the next thing to do, is present the issue to the relevant upstream
> maintainers. Searching for "nvme controller is down" brought up another bug
> (but that happened pretty instantly) and in there the request was made to
> report the issue (via email) to the linux-pci@vger.kernel.org and
> linux-nvme@lists.infradead.org lists.
> And that seems like the next best step for you too.

The message to linux-nvme finally came through and the thread is here: http://lists.infradead.org/pipermail/linux-nvme/2023-January/037384.html

For linux-pci, I am not sure if it worked.
I got a "Delivery status OK" and a "BOUNCE linux-pci@vger.kernel.org:     Message too long (>100000 chars)".
Obviously my message isn't >100000, so I assume they have a problem with one of the attachments.
But the message doesn't contain any information about if the mail got refused or not.
And more importantly, the "Delivery status OK" message, specifically says that the mail got delivered to the mailing lists address.

[toc] | [prev] | [next] | [standalone]


#77906

FromDiederik de Haas <didi.debian@cknow.org>
Date2023-01-12 15:00 +0100
Message-ID<FN6hz-husL-3@gated-at.bofh.it>
In reply to#77902

[Multipart message — attachments visible in raw view] — view raw

Control: forwarded -1 http://lists.infradead.org/pipermail/linux-nvme/2023-January/037384.html

On Thursday, 12 January 2023 12:11:57 CET Julian wrote:
> The message to linux-nvme finally came through and the thread is here:
> http://lists.infradead.org/pipermail/linux-nvme/2023-January/037384.html

Thanks!

[toc] | [prev] | [next] | [standalone]


#77908

FromDiederik de Haas <didi.debian@cknow.org>
Date2023-01-12 15:50 +0100
Message-ID<FN73X-huZY-1@gated-at.bofh.it>
In reply to#77902

[Multipart message — attachments visible in raw view] — view raw

Hi Julian,

On Thursday, 12 January 2023 12:11:57 CET Julian wrote:
> The message to linux-nvme finally came through and the thread is here:
> http://lists.infradead.org/pipermail/linux-nvme/2023-January/037384.html

The following paragraph may not be ideally formulated:

"Currently, I am using git bisect to narrow down the window of possible 
commits, but since the issue appears seemingly random, it will take many 
months to identify the offending commit this way."

Why? It *could* be that the maintainers will wait for the result of the
`git bisect` before responding/acting upon it.

Hopefully I'm wrong, but if they don't respond in a 'reasonable' time frame, 
you may want to clarify that you actually don't want to do the `git bisect` 
exactly because it could take many months.

Cheers,
  Diederik

[toc] | [prev] | [next] | [standalone]


#77909

FromJulian <julian.g@posteo.de>
Date2023-01-12 16:10 +0100
Message-ID<FN7nj-hvmG-9@gated-at.bofh.it>
In reply to#77908

[Multipart message — attachments visible in raw view] — view raw

On Thu, 12 Jan 2023 15:41:03 +0100 Diederik de Haas <didi.debian@cknow.org> wrote:
> "Currently, I am using git bisect to narrow down the window of possible
> commits, but since the issue appears seemingly random, it will take many
> months to identify the offending commit this way."
>
> Why? It *could* be that the maintainers will wait for the result of the
> `git bisect` before responding/acting upon it.

My intention has been to continue with the git bisect.
However, I have already encountered two revisions that do not build, so I doubt I can get very far there.

I will wait till next week and then report whatever additional information I was able to get through git bisect to the Kernel maintainers. Then I will most likely also tell them that I will not continue with git bisect unless they have a smaller window of revisions for me to try.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web