Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #85265 > unrolled thread

Bug#1093243: Upgrade to 6.1.123 kernel causes mariadb hangs

Started bySalvatore Bonaccorso <carnil@debian.org>
First post2025-01-23 22:00 +0100
Last post2025-01-27 17:50 +0100
Articles 3 — 2 participants

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1093243: Upgrade to 6.1.123 kernel causes mariadb hangs Salvatore Bonaccorso <carnil@debian.org> - 2025-01-23 22:00 +0100
    Bug#1093243: Upgrade to 6.1.123 kernel causes mariadb hangs Xan Charbonnet <xan@charbonnet.com> - 2025-01-24 03:20 +0100
    Bug#1093243: Upgrade to 6.1.123 kernel causes mariadb hangs Xan Charbonnet <xan@charbonnet.com> - 2025-01-27 17:50 +0100

#85265 — Bug#1093243: Upgrade to 6.1.123 kernel causes mariadb hangs

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-01-23 22:00 +0100
SubjectBug#1093243: Upgrade to 6.1.123 kernel causes mariadb hangs
Message-ID<K8csV-bjBX-1@gated-at.bofh.it>
Hi Xan,

On Thu, Jan 23, 2025 at 02:31:34PM -0600, Xan Charbonnet wrote:
> I rented a Linode and have been trying to load it down with sysbench
> activity while doing a mariabackup and a mysqldump, also while spinning up
> the CPU with zstd benchmarks.  So far I've had no luck triggering the fault.
> 
> I've also been doing some kernel compilation.  I followed this guide:
> https://www.dwarmstrong.org/kernel/
> (except that I used make -j24 to build in parallel and used make
> localmodconfig to compile only the modules I need)
> 
> I've built the following kernels:
> 6.1.123 (equivalent to linux-image-6.1.0-29-amd64)
> 6.1.122
> 6.1.121
> 6.1.120
> 
> So far they have all exhibited the behavior.  Next up is 6.1.119 which is
> equivalent to linux-image-6.1.0-28-amd64.  My expectation is that the fault
> will not appear for this kernel.
> 
> It looks like the issue is here somewhere:
> https://www.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.1.120
> 
> I have to work on some other things, and it'll take a while to prove the
> negative (that is, to know that the failure isn't happening).  I'll post
> back with the 6.1.119 results when I have them.

Additionally please try with 6.1.120 and revert this commit 

3ab9326f93ec ("io_uring: wake up optimisations")

(which landed in 6.1.120).

If that solves the problem maybe we miss some prequisites in the 6.1.y
series here?

Regards,
Salvatore

[toc] | [next] | [standalone]


#85268

FromXan Charbonnet <xan@charbonnet.com>
Date2025-01-24 03:20 +0100
Message-ID<K8hsB-bonh-7@gated-at.bofh.it>
In reply to#85265
On 1/23/25 20:49, Salvatore Bonaccorso wrote:
> Additionally please try with 6.1.120 and revert this commit
>
> 3ab9326f93ec ("io_uring: wake up optimisations")
>
> (which landed in 6.1.120).
>
> If that solves the problem maybe we miss some prequisites in the 6.1.y
> series here?


I hope I did all this right.  I found this:
https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id=3181e22fb79910c7071e84a43af93ac89e8a7106

and attempted to undo that change in the vanilla 6.1.124 source by 
making the following change to io_uring/io_uring.c:

585,594d584
< static inline void __io_cq_unlock_post_flush(struct io_ring_ctx *ctx)
<       __releases(ctx->completion_lock)
< {
<       io_commit_cqring(ctx);
<       spin_unlock(&ctx->completion_lock);
<       io_commit_cqring_flush(ctx);
<       if (!(ctx->flags & IORING_SETUP_DEFER_TASKRUN))
<               __io_cqring_wake(ctx);
< }
<
1352c1342
<       __io_cq_unlock_post_flush(ctx);
---
 >         __io_cq_unlock_post(ctx);


I rebooted into the resulting kernel and am happy to report that the 
problem did NOT occur!

[toc] | [prev] | [next] | [standalone]


#85301

FromXan Charbonnet <xan@charbonnet.com>
Date2025-01-27 17:50 +0100
Message-ID<K9Atb-cmgl-1@gated-at.bofh.it>
In reply to#85265
The MariaDB developers are wondering whether another corruption bug, 
MDEV-35334 ( https://jira.mariadb.org/browse/MDEV-35334 ) might be related.

The symptom was described as:
the first 1 byte of a .ibd file is changed from 0 to 1, or the first 4 
bytes are changed from 0 0 0 0 to 1 0 0 0.

Is it possible that an io_uring issue might be causing that as well? 
Thanks.

-Xan

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web