Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1370323 > unrolled thread

4.6-rc1 regression in SPI core -- deadlock

Started byRich Felker <dalias@libc.org>
First post2016-04-04 03:30 +0200
Last post2016-04-04 09:10 +0200
Articles 3 — 2 participants

Back to article view | Back to linux.kernel


Contents

  4.6-rc1 regression in SPI core -- deadlock Rich Felker <dalias@libc.org> - 2016-04-04 03:30 +0200
    Re: 4.6-rc1 regression in SPI core -- deadlock Vignesh R <vigneshr@ti.com> - 2016-04-04 06:00 +0200
      Re: 4.6-rc1 regression in SPI core -- deadlock Rich Felker <dalias@libc.org> - 2016-04-04 09:10 +0200

#1370323 — 4.6-rc1 regression in SPI core -- deadlock

FromRich Felker <dalias@libc.org>
Date2016-04-04 03:30 +0200
Subject4.6-rc1 regression in SPI core -- deadlock
Message-ID<rk1v4-3Ki-3@gated-at.bofh.it>
I've spent several days trying to debug a deadlock using our local
(not yet ready for upstream) driver for the J-Core SPI device and it
seems to be a new deadlock in the SPI core caused by commit
556351f14e74 and unrelated to the particular driver. Commit
49023d2e4ead tried to solve a related deadlock problem, but there
still seems to be a lock order issue and it's affecting SPI use even
without the spi_flash_read optimization. The deadlock I'm observing
has a kworker thread stuck in wait_for_completion called from
spi_sync_locked (ultimately from mmc_rescan) and the completion is
never finishing because this kworker thread has the bus locked while
the spi master task has already started processing the queue but can't
proceed because the bus lock is taken.

Anyone else seen this? Ideas for a proper fix? I've got it working for
me by disabling the bus locking in __spi_pump_messages entirely (like
before commit 556351f14e74) but I'm pretty sure this breaks the new
feature that was added.

Rich

[toc] | [next] | [standalone]


#1370337

FromVignesh R <vigneshr@ti.com>
Date2016-04-04 06:00 +0200
Message-ID<rk3Qd-5vv-1@gated-at.bofh.it>
In reply to#1370323

On 04/04/2016 06:50 AM, Rich Felker wrote:
> I've spent several days trying to debug a deadlock using our local
> (not yet ready for upstream) driver for the J-Core SPI device and it
> seems to be a new deadlock in the SPI core caused by commit
> 556351f14e74 and unrelated to the particular driver. Commit
> 49023d2e4ead tried to solve a related deadlock problem, but there
> still seems to be a lock order issue and it's affecting SPI use even
> without the spi_flash_read optimization. The deadlock I'm observing
> has a kworker thread stuck in wait_for_completion called from
> spi_sync_locked (ultimately from mmc_rescan) and the completion is
> never finishing because this kworker thread has the bus locked while
> the spi master task has already started processing the queue but can't
> proceed because the bus lock is taken.
> 
> Anyone else seen this? Ideas for a proper fix?

Could you try 24c8cd1b081286("spi: fix possible deadlock between
internal bus locks and bus_lock_flag") from linux-next?

-- 
Regards
Vignesh

[toc] | [prev] | [next] | [standalone]


#1370404

FromRich Felker <dalias@libc.org>
Date2016-04-04 09:10 +0200
Message-ID<rk6O6-7YC-5@gated-at.bofh.it>
In reply to#1370337
On Mon, Apr 04, 2016 at 09:23:31AM +0530, Vignesh R wrote:
> 
> 
> On 04/04/2016 06:50 AM, Rich Felker wrote:
> > I've spent several days trying to debug a deadlock using our local
> > (not yet ready for upstream) driver for the J-Core SPI device and it
> > seems to be a new deadlock in the SPI core caused by commit
> > 556351f14e74 and unrelated to the particular driver. Commit
> > 49023d2e4ead tried to solve a related deadlock problem, but there
> > still seems to be a lock order issue and it's affecting SPI use even
> > without the spi_flash_read optimization. The deadlock I'm observing
> > has a kworker thread stuck in wait_for_completion called from
> > spi_sync_locked (ultimately from mmc_rescan) and the completion is
> > never finishing because this kworker thread has the bus locked while
> > the spi master task has already started processing the queue but can't
> > proceed because the bus lock is taken.
> > 
> > Anyone else seen this? Ideas for a proper fix?
> 
> Could you try 24c8cd1b081286("spi: fix possible deadlock between
> internal bus locks and bus_lock_flag") from linux-next?

This seems to fix the problem. I'm still hitting an issue where
mmc_rescan deadlocks (at best) or crashes/corrupts the card state, but
I think that's a separate bug since it happened with my workaround
hack too. The patch in linux-next makes sense to me from what I've
read of the SPI core code so far. Thanks!

Rich

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web