Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1210262 > unrolled thread
| Started by | Christoph Hellwig <hch@infradead.org> |
|---|---|
| First post | 2015-08-20 10:10 +0200 |
| Last post | 2015-09-07 09:40 +0200 |
| Articles | 7 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: IDE Floppy support for IOMEGA Zip Drive broken in 3.16 -> 3.17 transition Christoph Hellwig <hch@infradead.org> - 2015-08-20 10:10 +0200
Re: IDE Floppy support for IOMEGA Zip Drive broken in 3.16 -> 3.17 transition Sergio Callegari <sergio.callegari@gmail.com> - 2015-08-20 17:10 +0200
"scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations Sergio Callegari <sergio.callegari@gmail.com> - 2015-08-24 19:10 +0200
Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations Christoph Hellwig <hch@infradead.org> - 2015-08-25 13:10 +0200
Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations Sergio Callegari <sergio.callegari@gmail.com> - 2015-08-26 09:20 +0200
Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations Sergio Callegari <sergio.callegari@gmail.com> - 2015-08-30 13:00 +0200
Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations Sergio Callegari <sergio.callegari@gmail.com> - 2015-09-07 09:40 +0200
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2015-08-20 10:10 +0200 |
| Subject | Re: IDE Floppy support for IOMEGA Zip Drive broken in 3.16 -> 3.17 transition |
| Message-ID | <pZt58-7r8-29@gated-at.bofh.it> |
Hi Sergio, On Tue, Aug 18, 2015 at 09:44:28AM +0200, Sergio Callegari wrote: > Hi, > > I have bisected the issue down to > > [045065d8a300a37218c548e9aa7becd581c6a0e8] [SCSI] fix qemu boot hang problem > > Bisecting has been a painful job due to the fact that the bug may show only > many hours after the system boot. > > The commit above in fact is not the culprit, but a fix to an issue that was > hiding the real bug on my system. See > > http://marc.info/?l=linux-kernel&m=143973820612978&w=2 > > The real issue is with sata host lock and seems to be biting a few other > people as well > > https://bbs.archlinux.org/viewtopic.php?id=189324 > > A patch fixing the issue was sent to the LKML back in Nov 2014 by Christoph > Hellwig (who is reading in CC) > > https://lkml.org/lkml/2014/11/20/581 > > I have tested the patch and it works for me. > > What is expected to happen now? As mentioned in that thread we need more input from the libata people on what kind of race this is papering over. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Sergio Callegari <sergio.callegari@gmail.com> |
|---|---|
| Date | 2015-08-20 17:10 +0200 |
| Message-ID | <pZzDz-8vO-3@gated-at.bofh.it> |
| In reply to | #1210262 |
Hi Christoph, thanks for your answer! Any clue on how to help/speed up this? Was there any exchange with them when you first provided your patch back in 2014? Can it be restarted? My feeling is that the issue fixed with 045065d8a300a37218c548e9aa7becd581c6a0e8 was in fact the real paper over the bug. As distros start picking kernels with 045065d8a300a37218c548e9aa7becd581c6a0e8 the issue is like to bite more and more. And it is a subtle bug. In some cases takes ages before revealing, making one feel confident on a system that is in fact broken. And takes ages to bisect (took me almost 7 days). Best, Sergio On 20/08/2015 10:08, Christoph Hellwig wrote: > Hi Sergio, > > On Tue, Aug 18, 2015 at 09:44:28AM +0200, Sergio Callegari wrote: >> Hi, >> >> I have bisected the issue down to >> >> [045065d8a300a37218c548e9aa7becd581c6a0e8] [SCSI] fix qemu boot hang problem >> >> Bisecting has been a painful job due to the fact that the bug may show only >> many hours after the system boot. >> >> The commit above in fact is not the culprit, but a fix to an issue that was >> hiding the real bug on my system. See >> >> http://marc.info/?l=linux-kernel&m=143973820612978&w=2 >> >> The real issue is with sata host lock and seems to be biting a few other >> people as well >> >> https://bbs.archlinux.org/viewtopic.php?id=189324 >> >> A patch fixing the issue was sent to the LKML back in Nov 2014 by Christoph >> Hellwig (who is reading in CC) >> >> https://lkml.org/lkml/2014/11/20/581 >> >> I have tested the patch and it works for me. >> >> What is expected to happen now? > As mentioned in that thread we need more input from the libata people > on what kind of race this is papering over. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Sergio Callegari <sergio.callegari@gmail.com> |
|---|---|
| Date | 2015-08-24 19:10 +0200 |
| Subject | "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations |
| Message-ID | <q13pT-5xE-1@gated-at.bofh.it> |
| In reply to | #1210262 |
Thanks Christoph for the answer! Apparently I missed a piece of the thread where the test patch was originally proposed . Now, I have gone through it and I see how the patch was not meant to be a final correction. My (possibly naive) understanding is that: - Even if this might be due to hardware that not fully conforms to the standard (but we do not know right now), commit 74665016086615bbaa3fa6f83af410a0a4e029ee ( scsi: convert host_busy to atomic_t ) certainly breaks the kernel for some hardware configurations causing a regression. - If the regression was immediately spotted, the patch would probably have been revised right after proposal. Unfortunately, another bug - that got fixed only much later with 045065d8a300a37218c - hid the original issue for a long time. - Now that a lot of time has passed with the "scsi: convert host_busy to atomic_t" series in the kernel, going back to look into it is much more difficult. Libata people might not be very interested as they moved to other topics and might need a lot of time to go through it (it has been known since November 2014 - 9 months ago), possibly due to the race like nature of the issue and the fact that the bug might not be reproducible on their hardware... Is this correct? Aren't commits that cause regressions confirmed by multiple users expected (at least in principle) to be reverted? If reverting is too costy, wouldn't your "papering over" or making the scsi delay configurable be an acceptable solution? Even better: can in some way the libata-people be helped find the real culprit, given that there are at least two hardware setups that are known to trigger the regression (mine and Barto's)? I have tried the linux-ide mailing list, but got silence. Best, Sergio On 20/08/2015 10:08, Christoph Hellwig wrote: > Hi Sergio, > > On Tue, Aug 18, 2015 at 09:44:28AM +0200, Sergio Callegari wrote: >> Hi, >> >> I have bisected the issue down to >> >> [045065d8a300a37218c548e9aa7becd581c6a0e8] [SCSI] fix qemu boot hang problem >> >> Bisecting has been a painful job due to the fact that the bug may show only >> many hours after the system boot. >> >> The commit above in fact is not the culprit, but a fix to an issue that was >> hiding the real bug on my system. See >> >> http://marc.info/?l=linux-kernel&m=143973820612978&w=2 >> >> The real issue is with sata host lock and seems to be biting a few other >> people as well >> >> https://bbs.archlinux.org/viewtopic.php?id=189324 >> >> A patch fixing the issue was sent to the LKML back in Nov 2014 by Christoph >> Hellwig (who is reading in CC) >> >> https://lkml.org/lkml/2014/11/20/581 >> >> I have tested the patch and it works for me. >> >> What is expected to happen now? > As mentioned in that thread we need more input from the libata people > on what kind of race this is papering over. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2015-08-25 13:10 +0200 |
| Subject | Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations |
| Message-ID | <q1kh4-4Fj-29@gated-at.bofh.it> |
| In reply to | #1212350 |
Hi Sergio,
can you give the patch below a try?
libata currently completes the SCSI command before freeing the internal
command structure, which could lead to various races that mess with
the ATA command state, which might cause issues like the one you see.
diff --git a/drivers/ata/libata-scsi.c b/drivers/ata/libata-scsi.c
index 641a61a..92cc156 100644
--- a/drivers/ata/libata-scsi.c
+++ b/drivers/ata/libata-scsi.c
@@ -1772,6 +1772,15 @@ nothing_to_do:
return 1;
}
+static void ata_qc_done(struct ata_queued_cmd *qc)
+{
+ struct scsi_cmnd *cmd = qc->scsicmd;
+ void (*done)(struct scsi_cmnd *) = qc->scsidone;
+
+ ata_qc_free(qc);
+ done(cmd);
+}
+
static void ata_scsi_qc_complete(struct ata_queued_cmd *qc)
{
struct ata_port *ap = qc->ap;
@@ -1810,9 +1819,7 @@ static void ata_scsi_qc_complete(struct ata_queued_cmd *qc)
if (need_sense && !ap->ops->error_handler)
ata_dump_status(ap->print_id, &qc->result_tf);
- qc->scsidone(cmd);
-
- ata_qc_free(qc);
+ ata_qc_done(qc);
}
/**
@@ -2611,8 +2618,7 @@ static void atapi_sense_complete(struct ata_queued_cmd *qc)
ata_gen_passthru_sense(qc);
}
- qc->scsidone(qc->scsicmd);
- ata_qc_free(qc);
+ ata_qc_done(qc);
}
/* is it pointless to prefer PIO for "safety reasons"? */
@@ -2707,8 +2713,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
qc->dev->sdev->locked = 0;
qc->scsicmd->result = SAM_STAT_CHECK_CONDITION;
- qc->scsidone(cmd);
- ata_qc_free(qc);
+ ata_qc_done(qc);
return;
}
@@ -2752,8 +2757,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
cmd->result = SAM_STAT_GOOD;
}
- qc->scsidone(cmd);
- ata_qc_free(qc);
+ ata_qc_done(qc);
}
/**
* atapi_xlat - Initialize PACKET taskfile
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Sergio Callegari <sergio.callegari@gmail.com> |
|---|---|
| Date | 2015-08-26 09:20 +0200 |
| Subject | Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations |
| Message-ID | <q1Da2-7cg-3@gated-at.bofh.it> |
| In reply to | #1212968 |
Sure, thanks!
I'll test this weekend.
Best,
Sergio
On 25/08/2015 13:00, Christoph Hellwig wrote:
> Hi Sergio,
>
> can you give the patch below a try?
>
> libata currently completes the SCSI command before freeing the internal
> command structure, which could lead to various races that mess with
> the ATA command state, which might cause issues like the one you see.
>
> diff --git a/drivers/ata/libata-scsi.c b/drivers/ata/libata-scsi.c
> index 641a61a..92cc156 100644
> --- a/drivers/ata/libata-scsi.c
> +++ b/drivers/ata/libata-scsi.c
> @@ -1772,6 +1772,15 @@ nothing_to_do:
> return 1;
> }
>
> +static void ata_qc_done(struct ata_queued_cmd *qc)
> +{
> + struct scsi_cmnd *cmd = qc->scsicmd;
> + void (*done)(struct scsi_cmnd *) = qc->scsidone;
> +
> + ata_qc_free(qc);
> + done(cmd);
> +}
> +
> static void ata_scsi_qc_complete(struct ata_queued_cmd *qc)
> {
> struct ata_port *ap = qc->ap;
> @@ -1810,9 +1819,7 @@ static void ata_scsi_qc_complete(struct ata_queued_cmd *qc)
> if (need_sense && !ap->ops->error_handler)
> ata_dump_status(ap->print_id, &qc->result_tf);
>
> - qc->scsidone(cmd);
> -
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> }
>
> /**
> @@ -2611,8 +2618,7 @@ static void atapi_sense_complete(struct ata_queued_cmd *qc)
> ata_gen_passthru_sense(qc);
> }
>
> - qc->scsidone(qc->scsicmd);
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> }
>
> /* is it pointless to prefer PIO for "safety reasons"? */
> @@ -2707,8 +2713,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
> qc->dev->sdev->locked = 0;
>
> qc->scsicmd->result = SAM_STAT_CHECK_CONDITION;
> - qc->scsidone(cmd);
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> return;
> }
>
> @@ -2752,8 +2757,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
> cmd->result = SAM_STAT_GOOD;
> }
>
> - qc->scsidone(cmd);
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> }
> /**
> * atapi_xlat - Initialize PACKET taskfile
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Sergio Callegari <sergio.callegari@gmail.com> |
|---|---|
| Date | 2015-08-30 13:00 +0200 |
| Subject | Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations |
| Message-ID | <q38v7-6MF-7@gated-at.bofh.it> |
| In reply to | #1212968 |
Hi Christoph,
just checked.
Unfortunately, the patch below, applied on top of Linus' v3.17 (which I
am using as a test kernel) *does not fix the issue*.
Best regards,
Sergio
On 25/08/2015 13:00, Christoph Hellwig wrote:
> Hi Sergio,
>
> can you give the patch below a try?
>
> libata currently completes the SCSI command before freeing the internal
> command structure, which could lead to various races that mess with
> the ATA command state, which might cause issues like the one you see.
>
> diff --git a/drivers/ata/libata-scsi.c b/drivers/ata/libata-scsi.c
> index 641a61a..92cc156 100644
> --- a/drivers/ata/libata-scsi.c
> +++ b/drivers/ata/libata-scsi.c
> @@ -1772,6 +1772,15 @@ nothing_to_do:
> return 1;
> }
>
> +static void ata_qc_done(struct ata_queued_cmd *qc)
> +{
> + struct scsi_cmnd *cmd = qc->scsicmd;
> + void (*done)(struct scsi_cmnd *) = qc->scsidone;
> +
> + ata_qc_free(qc);
> + done(cmd);
> +}
> +
> static void ata_scsi_qc_complete(struct ata_queued_cmd *qc)
> {
> struct ata_port *ap = qc->ap;
> @@ -1810,9 +1819,7 @@ static void ata_scsi_qc_complete(struct ata_queued_cmd *qc)
> if (need_sense && !ap->ops->error_handler)
> ata_dump_status(ap->print_id, &qc->result_tf);
>
> - qc->scsidone(cmd);
> -
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> }
>
> /**
> @@ -2611,8 +2618,7 @@ static void atapi_sense_complete(struct ata_queued_cmd *qc)
> ata_gen_passthru_sense(qc);
> }
>
> - qc->scsidone(qc->scsicmd);
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> }
>
> /* is it pointless to prefer PIO for "safety reasons"? */
> @@ -2707,8 +2713,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
> qc->dev->sdev->locked = 0;
>
> qc->scsicmd->result = SAM_STAT_CHECK_CONDITION;
> - qc->scsidone(cmd);
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> return;
> }
>
> @@ -2752,8 +2757,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
> cmd->result = SAM_STAT_GOOD;
> }
>
> - qc->scsidone(cmd);
> - ata_qc_free(qc);
> + ata_qc_done(qc);
> }
> /**
> * atapi_xlat - Initialize PACKET taskfile
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Sergio Callegari <sergio.callegari@gmail.com> |
|---|---|
| Date | 2015-09-07 09:40 +0200 |
| Subject | Re: "scsi: convert host_busy to atomic_t" series causes regressions for some hardware configurations |
| Message-ID | <q5ZbY-7oS-7@gated-at.bofh.it> |
| In reply to | #1215908 |
Hi Christoph (and those reading!),
I wonder if there might be any update or, most important, anything else for me
to test in order to provide info to address this issue...
I really would like to find out whether it is a bug in my hardware (which would
be OK, as I know already how to work around it) or something that may happen on
any hardware as soon as an unfortunate combination of storage equipment is adopted.
Thanks for the help so far,
Best regards,
Sergio
On 30/08/2015 12:54, Sergio Callegari wrote:
> Hi Christoph,
>
> just checked.
>
> Unfortunately, the patch below, applied on top of Linus' v3.17 (which I am
> using as a test kernel) *does not fix the issue*.
>
> Best regards,
>
> Sergio
>
> On 25/08/2015 13:00, Christoph Hellwig wrote:
>> Hi Sergio,
>>
>> can you give the patch below a try?
>>
>> libata currently completes the SCSI command before freeing the internal
>> command structure, which could lead to various races that mess with
>> the ATA command state, which might cause issues like the one you see.
>>
>> diff --git a/drivers/ata/libata-scsi.c b/drivers/ata/libata-scsi.c
>> index 641a61a..92cc156 100644
>> --- a/drivers/ata/libata-scsi.c
>> +++ b/drivers/ata/libata-scsi.c
>> @@ -1772,6 +1772,15 @@ nothing_to_do:
>> return 1;
>> }
>> +static void ata_qc_done(struct ata_queued_cmd *qc)
>> +{
>> + struct scsi_cmnd *cmd = qc->scsicmd;
>> + void (*done)(struct scsi_cmnd *) = qc->scsidone;
>> +
>> + ata_qc_free(qc);
>> + done(cmd);
>> +}
>> +
>> static void ata_scsi_qc_complete(struct ata_queued_cmd *qc)
>> {
>> struct ata_port *ap = qc->ap;
>> @@ -1810,9 +1819,7 @@ static void ata_scsi_qc_complete(struct ata_queued_cmd
>> *qc)
>> if (need_sense && !ap->ops->error_handler)
>> ata_dump_status(ap->print_id, &qc->result_tf);
>> - qc->scsidone(cmd);
>> -
>> - ata_qc_free(qc);
>> + ata_qc_done(qc);
>> }
>> /**
>> @@ -2611,8 +2618,7 @@ static void atapi_sense_complete(struct ata_queued_cmd
>> *qc)
>> ata_gen_passthru_sense(qc);
>> }
>> - qc->scsidone(qc->scsicmd);
>> - ata_qc_free(qc);
>> + ata_qc_done(qc);
>> }
>> /* is it pointless to prefer PIO for "safety reasons"? */
>> @@ -2707,8 +2713,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
>> qc->dev->sdev->locked = 0;
>> qc->scsicmd->result = SAM_STAT_CHECK_CONDITION;
>> - qc->scsidone(cmd);
>> - ata_qc_free(qc);
>> + ata_qc_done(qc);
>> return;
>> }
>> @@ -2752,8 +2757,7 @@ static void atapi_qc_complete(struct ata_queued_cmd *qc)
>> cmd->result = SAM_STAT_GOOD;
>> }
>> - qc->scsidone(cmd);
>> - ata_qc_free(qc);
>> + ata_qc_done(qc);
>> }
>> /**
>> * atapi_xlat - Initialize PACKET taskfile
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web