Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1295644 > unrolled thread
| Started by | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| First post | 2015-12-20 19:00 +0100 |
| Last post | 2015-12-22 00:00 +0100 |
| Articles | 20 on this page of 44 — 9 participants |
Back to article view | Back to linux.kernel
IO errors after "block: remove bio_get_nr_vecs()" Linus Torvalds <torvalds@linux-foundation.org> - 2015-12-20 19:00 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Christoph Hellwig <hch@lst.de> - 2015-12-20 19:20 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-20 19:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 00:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Linus Torvalds <torvalds@linux-foundation.org> - 2015-12-20 19:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 00:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Dan Aloni <dan@kernelim.com> - 2015-12-21 12:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 00:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-21 00:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 01:00 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 00:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Ming Lei <tom.leiming@gmail.com> - 2015-12-21 02:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 03:00 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Ming Lei <tom.leiming@gmail.com> - 2015-12-21 03:20 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 03:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-21 03:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Ming Lei <tom.leiming@gmail.com> - 2015-12-21 04:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 04:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Linus Torvalds <torvalds@linux-foundation.org> - 2015-12-21 05:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 05:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Linus Torvalds <torvalds@linux-foundation.org> - 2015-12-21 05:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Linus Torvalds <torvalds@linux-foundation.org> - 2015-12-21 06:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-21 08:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-22 05:10 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Tejun Heo <tj@kernel.org> - 2015-12-21 05:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Linus Torvalds <torvalds@linux-foundation.org> - 2015-12-21 06:20 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Tejun Heo <tj@kernel.org> - 2015-12-21 08:00 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Tejun Heo <tj@kernel.org> - 2015-12-21 20:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Tejun Heo <tj@kernel.org> - 2015-12-21 21:10 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Tejun Heo <tj@kernel.org> - 2015-12-21 22:10 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-22 04:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-22 05:10 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Junichi Nomura <j-nomura@ce.jp.nec.com> - 2015-12-22 06:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-22 06:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-22 07:00 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-22 07:00 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-22 07:00 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-22 07:10 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-22 06:40 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Jens Axboe <axboe@fb.com> - 2015-12-22 18:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Kent Overstreet <kent.overstreet@gmail.com> - 2015-12-22 05:50 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-22 06:20 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" "Artem S. Tashkinov" <t.artem@lycos.com> - 2015-12-22 06:30 +0100
Re: IO errors after "block: remove bio_get_nr_vecs()" Ming Lei <tom.leiming@gmail.com> - 2015-12-22 00:00 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2015-12-21 05:50 +0100 |
| Message-ID | <qI0A2-1P6-7@gated-at.bofh.it> |
| In reply to | #1295767 |
On Sun, Dec 20, 2015 at 8:43 PM, Artem S. Tashkinov <t.artem@lycos.com> wrote:
>
> In the past I happily ran an x86_64 bit kernel together with 32bit userland
> for quite some time but then I hit a wall: VirtualBox expects its kernel
> modules to have the same bitness as the application itself so I had to
> revert back to an i686 PAE setup.
Ugh, ok. That kind of forces your hand, yes.
Although:
> t's probably high time to try qemu however last time I looked at it a few
> years ago it lacked several crucial features I need from a VM.
kvm-qemu really ends up working pretty well.. Give it a try.
That said, we obviously need to figure out this current problem
regardless first..
Linus
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2015-12-21 06:30 +0100 |
| Message-ID | <qI1cJ-2jl-9@gated-at.bofh.it> |
| In reply to | #1295769 |
On Sun, Dec 20, 2015 at 8:47 PM, Linus Torvalds
<torvalds@linux-foundation.org> wrote:
>
> That said, we obviously need to figure out this current problem
> regardless first..
... although maybe it *would* be interesting to hear what happens if
you just compile a 64-bit kernel instead?
Do you still see the problem? Because if not, then we should look very
specifically for some 32-bit PAE issue.
For example, maybe we use "unsigned long" somewhere where we should
use "phys_addr_t". On x86-64, they obviously end up being the same. On
normal non-PAE x86-32, they are also the same. But ..
Linus
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Artem S. Tashkinov" <t.artem@lycos.com> |
|---|---|
| Date | 2015-12-21 08:40 +0100 |
| Message-ID | <qI3ex-3xA-9@gated-at.bofh.it> |
| In reply to | #1295774 |
On 2015-12-21 10:23, Linus Torvalds wrote: > On Sun, Dec 20, 2015 at 8:47 PM, Linus Torvalds > <torvalds@linux-foundation.org> wrote: >> >> That said, we obviously need to figure out this current problem >> regardless first.. > > ... although maybe it *would* be interesting to hear what happens if > you just compile a 64-bit kernel instead? > > Do you still see the problem? Because if not, then we should look very > specifically for some 32-bit PAE issue. > > For example, maybe we use "unsigned long" somewhere where we should > use "phys_addr_t". On x86-64, they obviously end up being the same. On > normal non-PAE x86-32, they are also the same. But .. > Let's wait for what Tejun Heo might say - I've applied his debugging patch and sent back the results. Building x86_64 kernel here involves installing a 64bit Linux VM, so I'd like it to be the last resort. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Artem S. Tashkinov" <t.artem@lycos.com> |
|---|---|
| Date | 2015-12-22 05:10 +0100 |
| Message-ID | <qImqR-7qS-1@gated-at.bofh.it> |
| In reply to | #1295774 |
On 2015-12-21 10:23, Linus Torvalds wrote: > On Sun, Dec 20, 2015 at 8:47 PM, Linus Torvalds > <torvalds@linux-foundation.org> wrote: >> >> That said, we obviously need to figure out this current problem >> regardless first.. > > ... although maybe it *would* be interesting to hear what happens if > you just compile a 64-bit kernel instead? Under x86-64 I cannot reproduce this problem. It seems like it's PAE specific (Kent Overstreet says he has reproduced it). > > Do you still see the problem? Because if not, then we should look very > specifically for some 32-bit PAE issue. > > For example, maybe we use "unsigned long" somewhere where we should > use "phys_addr_t". On x86-64, they obviously end up being the same. On > normal non-PAE x86-32, they are also the same. But .. > > Linus -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-12-21 05:30 +0100 |
| Message-ID | <qI0gF-1Hb-5@gated-at.bofh.it> |
| In reply to | #1295644 |
Hello, Linus. On Sun, Dec 20, 2015 at 09:51:14AM -0800, Linus Torvalds wrote: ... > (Also Tejun - maybe you can see what's up - maybe that error message > tells you something) Hmmm... all it says is that something went wrong on the PCI side. > I'm not sure what's up with his machine, the disk doesn't seem to be > anyuthing particularly unusual, it looks like a 1TB Seagate Barracuda: > > ata1.00: ATA-8: ST1000DM003-1CH162, CC44, max UDMA/133 > > which doesn't strike me as odd. > > Looking at the dmesg, it also looks like it's a pretty normal > Sandybridge setup with Intel chipset. Artem, can you confirm? The PCI > ID for the AHCI chip seems to be (INTEL, 0x1c02). > > Any ideas? Anybody? I wonder whether ahci is screwing up command / sg table setup in a way that e.g. if there are too many segments the sg table overflows into the neighboring one which is now being exposed by upper layer being fixed to send down larger commands. Looking into it. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2015-12-21 06:20 +0100 |
| Message-ID | <qI134-2e0-3@gated-at.bofh.it> |
| In reply to | #1295762 |
On Sun, Dec 20, 2015 at 8:26 PM, Tejun Heo <tj@kernel.org> wrote:
>
> I wonder whether ahci is screwing up command / sg table setup in a way
> that e.g. if there are too many segments the sg table overflows into
> the neighboring one which is now being exposed by upper layer being
> fixed to send down larger commands. Looking into it.
That would explain the
Corrupted low memory at c0001000 ...
that Artem also saw.
Anyway, it would be lovely to have some verification in the ATA
routines that the passed-on IO actually h9onors the limits it set.
Could you add a WARN_ON_ONCE(check_io_limits())" or similar, and maybe
we could catch whatever causes the overflow red-handed?
On a totally separate issue:
Just looking at some of the merging code, and I have to say that it
strikes me as insane. This in particular:
#define __BIO_SEG_BOUNDARY(addr1, addr2, mask) \
(((addr1) | (mask)) == (((addr2) - 1) | (mask)))
#define BIOVEC_SEG_BOUNDARY(q, b1, b2) \
__BIO_SEG_BOUNDARY(bvec_to_phys((b1)), bvec_to_phys((b2)) +
(b2)->bv_len, queue_segment_boundary((q)))
seems just *stupid*.
Why does it do that "bvec_to_phys((b2)) + (b2)->bv_len -1" on the
second bvec? That's the :"physical address of the last byte of the
second bvec".
I understand the "round both addresses up by the mask, and we want to
make sure that they are in the same segment" part.
But since an individual bvec had better be fully inside one segment
(since we split at bvec boundaries anyway, so if ). why do all that
crap anyway? The end address doesn't matter, you could just use the
beginning.
So remove the "-1" and remove the "+bv_len".
At which it would become just
#define __BIO_SEG_BOUNDARY(addr1, addr2, mask) \
((addr1) | (mask) == (addr2)|(mask))
#define BIOVEC_SEG_BOUNDARY(q, b1, b2) \
__BIO_SEG_BOUNDARY(bvec_to_phys((b1)), bvec_to_phys((b2)),
queue_segment_boundary((q)))
which seems simpler and more understandable. "Are the beginning
addresses in within the same segment"
Or are there ever bv_len == 0 things at the boundary that we want to
merge. Because then the "-1+bv_len" case migth make sense.
Anyway, that shouldn't change the end result in any way, so that
doesn't all *matter*, but it worries me when things look more
complicated than I think they should be.
Is there something I'm missing?
Linus
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-12-21 08:00 +0100 |
| Message-ID | <qI2BP-33y-1@gated-at.bofh.it> |
| In reply to | #1295644 |
Artem, can you please reproduce the issue with the following patch
applied and attach the kernel log?
Thanks.
---
drivers/ata/libahci.c | 40 ++++++++++++++++++++++++++++++++++++++--
drivers/ata/libata-eh.c | 4 ++++
drivers/ata/libata-scsi.c | 1 +
3 files changed, 43 insertions(+), 2 deletions(-)
--- a/drivers/ata/libahci.c
+++ b/drivers/ata/libahci.c
@@ -2278,7 +2278,7 @@ static int ahci_port_start(struct ata_po
struct ahci_host_priv *hpriv = ap->host->private_data;
struct device *dev = ap->host->dev;
struct ahci_port_priv *pp;
- void *mem;
+ void *mem, *base;
dma_addr_t mem_dma;
size_t dma_sz, rx_fis_sz;
@@ -2319,7 +2319,9 @@ static int ahci_port_start(struct ata_po
rx_fis_sz = AHCI_RX_FIS_SZ;
}
- mem = dmam_alloc_coherent(dev, dma_sz, &mem_dma, GFP_KERNEL);
+ base = mem = dmam_alloc_coherent(dev, dma_sz, &mem_dma, GFP_KERNEL);
+ printk("XXX port %d dma_sz=%zu mem=%p mem_dma=%p",
+ ap->port_no, dma_sz, mem, (void *)mem_dma);
if (!mem)
return -ENOMEM;
memset(mem, 0, dma_sz);
@@ -2331,6 +2333,8 @@ static int ahci_port_start(struct ata_po
pp->cmd_slot = mem;
pp->cmd_slot_dma = mem_dma;
+ pr_cont(" cmd_slot=%zu", mem - base);
+
mem += AHCI_CMD_SLOT_SZ;
mem_dma += AHCI_CMD_SLOT_SZ;
@@ -2340,6 +2344,8 @@ static int ahci_port_start(struct ata_po
pp->rx_fis = mem;
pp->rx_fis_dma = mem_dma;
+ pr_cont(" rx_fis=%zu", mem - base);
+
mem += rx_fis_sz;
mem_dma += rx_fis_sz;
@@ -2350,6 +2356,8 @@ static int ahci_port_start(struct ata_po
pp->cmd_tbl = mem;
pp->cmd_tbl_dma = mem_dma;
+ pr_cont(" cmd_tbl=%zu\n", mem - base);
+
/*
* Save off initial list of interrupts to be enabled.
* This could be changed later
@@ -2540,6 +2548,34 @@ int ahci_host_activate(struct ata_host *
}
EXPORT_SYMBOL_GPL(ahci_host_activate);
+void ahci_dump_dma(struct ata_queued_cmd *qc)
+{
+ struct ata_port *ap = qc->ap;
+ struct ahci_port_priv *pp = ap->private_data;
+ struct ahci_cmd_hdr *cmd = &pp->cmd_slot[qc->tag];
+ void *cmd_tbl = pp->cmd_tbl + qc->tag * AHCI_CMD_TBL_SZ;
+ u32 *fis = cmd_tbl;
+ struct ahci_sg *ahci_sg = cmd_tbl + AHCI_CMD_TBL_HDR_SZ;
+ int prdtl = (cmd->opts & 0xffff0000) >> 16;
+ int i;
+
+ printk("XXX cmd=%p cmd_tbl=%p ahci_sg=%p\n", cmd, cmd_tbl, ahci_sg);
+ printk("XXX opts=%x st=%x addr=%x addr_hi=%x rsvd=%x:%x:%x:%x\n",
+ cmd->opts, cmd->status, cmd->tbl_addr, cmd->tbl_addr_hi,
+ cmd->reserved[0], cmd->reserved[1], cmd->reserved[2], cmd->reserved[3]);
+ printk("XXX fis=%08x:%08x:%08x:%08x %08x:%08x:%08x:%08x\n",
+ fis[0], fis[1], fis[2], fis[3],
+ fis[4], fis[5], fis[6], fis[7]);
+
+ printk("XXX qc->n_elem=%d fis_len=%d prdtl=%d\n",
+ qc->n_elem, cmd->opts & 0xf, prdtl);
+
+ for (i = 0; i < prdtl; i++)
+ printk("XXX sg[%d] = %x %x %x (%d)\n",
+ i, ahci_sg[i].addr, ahci_sg[i].addr_hi, ahci_sg[i].flags_size,
+ (ahci_sg[i].flags_size & 0x7fffffff) + 1);
+}
+
MODULE_AUTHOR("Jeff Garzik");
MODULE_DESCRIPTION("Common AHCI SATA low-level routines");
MODULE_LICENSE("GPL");
--- a/drivers/ata/libata-eh.c
+++ b/drivers/ata/libata-eh.c
@@ -1059,6 +1059,7 @@ static int ata_do_link_abort(struct ata_
if (qc && (!link || qc->dev->link == link)) {
qc->flags |= ATA_QCFLAG_FAILED;
+ qc->err_mask = AC_ERR_DEV;
ata_qc_complete(qc);
nr_aborted++;
}
@@ -2416,6 +2417,8 @@ const char *ata_get_cmd_descript(u8 comm
}
EXPORT_SYMBOL_GPL(ata_get_cmd_descript);
+void ahci_dump_dma(struct ata_queued_cmd *qc);
+
/**
* ata_eh_link_report - report error handling to user
* @link: ATA link EH is going on
@@ -2590,6 +2593,7 @@ static void ata_eh_link_report(struct at
res->feature & ATA_IDNF ? "IDNF " : "",
res->feature & ATA_ABORTED ? "ABRT " : "");
#endif
+ ahci_dump_dma(qc);
}
}
--- a/drivers/ata/libata-scsi.c
+++ b/drivers/ata/libata-scsi.c
@@ -4035,6 +4035,7 @@ int ata_scsi_user_scan(struct Scsi_Host
}
if (rc == 0) {
+ ata_port_freeze(ap);
ata_port_schedule_eh(ap);
spin_unlock_irqrestore(ap->lock, flags);
ata_port_wait_eh(ap);
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-12-21 20:40 +0100 |
| Message-ID | <qIetk-2e0-11@gated-at.bofh.it> |
| In reply to | #1295787 |
Hello, Artem.
On Mon, Dec 21, 2015 at 12:25:06PM +0500, Artem S. Tashkinov wrote:
> I've applied this patch on top of vanilla 4.3.3 kernel (without Linus'es
> revert). Hopefully it's how you intended it to be.
>
> Here's the result (I skipped the beginning of dmesg - it's the same as
> always - see bugzilla).
I added some debug messages during init. It isn't critical but it'd
be great if you attach the full log from now on. Something seemingly
unrelated surprisingly often turns out to be an important clue.
> [ 60.387407] Corrupted low memory at c0001000 (1000 phys) = cba3d25f
...
> [ 60.389131] Corrupted low memory at c0001ffc (1ffc phys) = 2712322a
It looks like the controller shat on the entire second page, which is
really puzzling. Looks like the controller is being fed corrupt DMA
SG table.
...
> [ 74.367608] ata1.00: exception Emask 0x3 SAct 0x180000 SErr 0x3040400 action 0x6 frozen
> [ 74.367613] ata1.00: irq_stat 0x45000008
The intresting bit here is that the controller is indicating OVERFLOW
which means that it consumed all PRD entries (ahci's DMA SG table) for
the command but the disk is still sending data to the host.
> [ 74.367617] ata1: SError: { Proto CommWake TrStaTrns UnrecFIS }
> [ 74.367621] ata1.00: failed command: READ FPDMA QUEUED
> [ 74.367627] ata1.00: cmd 60/00:98:00:89:21/07:00:04:00:00/40 tag 19 ncq 917504 in
> [ 74.367627] res 40/00:98:00:89:21/00:00:04:00:00/40 Emask 0x3 (HSM violation)
> [ 74.367630] ata1.00: status: { DRDY }
The followings are the data fed to the controller as seen from the
CPU.
> [ 74.367632] XXX cmd=ee9e0260 cmd_tbl=ee9ed600 ahci_sg=ee9ed680
> [ 74.367634] XXX opts=140005 st=0 addr=2e9ed600 addr_hi=0 rsvd=0:0:0:0
> [ 74.367637] XXX fis=00608027:40218900:07000004:08000098 00000000:00000000:00000000:00001fff
> [ 74.367639] XXX qc->n_elem=20 fis_len=5 prdtl=20
> [ 74.367641] XXX sg[0] = 29007000 0 8fff (36864)
> [ 74.367643] XXX sg[1] = 29218000 0 7fff (32768)
> [ 74.367645] XXX sg[2] = 29298000 0 7fff (32768)
> [ 74.367646] XXX sg[3] = 29ad8000 0 7fff (32768)
> [ 74.367648] XXX sg[4] = 29338000 0 7fff (32768)
> [ 74.367650] XXX sg[5] = 29370000 0 2fff (12288)
> [ 74.367652] XXX sg[6] = 219000 0 2fff (12288)
> [ 74.367653] XXX sg[7] = 230000 0 3fff (16384)
> [ 74.367655] XXX sg[8] = 29373000 0 4fff (20480)
> [ 74.367657] XXX sg[9] = 29130000 0 ffff (65536)
> [ 74.367659] XXX sg[10] = 29170000 0 ffff (65536)
> [ 74.367660] XXX sg[11] = 29280000 0 ffff (65536)
> [ 74.367662] XXX sg[12] = 29200000 0 ffff (65536)
> [ 74.367664] XXX sg[13] = 29320000 0 ffff (65536)
> [ 74.367666] XXX sg[14] = 29360000 0 ffff (65536)
> [ 74.367667] XXX sg[15] = 29340000 0 ffff (65536)
> [ 74.367669] XXX sg[16] = 29350000 0 ffff (65536)
> [ 74.367671] XXX sg[17] = 29300000 0 ffff (65536)
> [ 74.367672] XXX sg[18] = 29310000 0 ffff (65536)
> [ 74.367674] XXX sg[19] = 29020000 0 7fff (32768)
And everything checks out. Data lenghts are consistent and all the
addresses look kosher - at least nothing should upset the data
transfer itself.
> [ 74.367677] ata1.00: failed command: READ FPDMA QUEUED
> [ 74.367682] ata1.00: cmd 60/90:a0:c0:fe:23/00:00:04:00:00/40 tag 20 ncq 73728 in
> [ 74.367682] res 40/00:00:00:00:00/00:00:00:00:00/40 Emask 0x7 (timeout)
This one looks like a collateral damage.
...
> [ 74.763895] ata1.00: exception Emask 0x3 SAct 0x800000 SErr 0x3040400 action 0x6
> [ 74.763900] ata1.00: irq_stat 0x45000008
> [ 74.763903] ata1: SError: { Proto CommWake TrStaTrns UnrecFIS }
> [ 74.763907] ata1.00: failed command: READ FPDMA QUEUED
> [ 74.763913] ata1.00: cmd 60/98:b8:00:c9:20/03:00:04:00:00/40 tag 23 ncq 471040 in
> [ 74.763913] res 40/00:b8:00:c9:20/00:00:04:00:00/40 Emask 0x3 (HSM violation)
> [ 74.763916] ata1.00: status: { DRDY }
> [ 74.763918] XXX cmd=ee9e02e0 cmd_tbl=ee9f0200 ahci_sg=ee9f0280
> [ 74.763921] XXX opts=160005 st=0 addr=2e9f0200 addr_hi=0 rsvd=0:0:0:0
> [ 74.763924] XXX fis=98608027:4020c900:03000004:080000b8 00000000:00000000:00000000:00000fff
> [ 74.763925] XXX qc->n_elem=22 fis_len=5 prdtl=22
> [ 74.763928] XXX sg[0] = 2ab1b000 0 fff (4096)
...
> [ 74.763964] XXX sg[21] = 29170000 0 dfff (57344)
This is a separate failure and shares the same pattern as before.
Everything looks proper.
The thing is ahci doesn't have much restrictions in terms of its DMA
capabilities. It can digest pretty much anything. The only
restriction is that each entry can't be larger than 4M - but our
segment maximum is 64k. There's no noticeable boundary crossing
happening both in target DMA regions and command tables. All
addresses are in linear mapped normal area.
If the controller is seeing what the host is seeing in the command
area, I can't see why it would be declaring overflow or dumping stuff
into the lowest pages.
Ming Lei reported a similar issue on 32bit ARM w/ PAE. I don't
understand why PAE is making any difference. Native 64bit machines
should be generating IOs just as large. Maybe iommu is hiding the
issue?
I'm afraid we'll have to go brute force with the problem - dump more
information on both 32 and 64bits and see where the differences lie.
At the moment, given that the DMA table looks completely fine and ahci
isn't picky about the shape of data areas, I think it's more likely
some obscure issue in address mapping under PAE or a silly bug in ahci
than block layer screwing up merging.
Thanks.
--
tejun
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-12-21 21:10 +0100 |
| Message-ID | <qIeWm-2Dm-21@gated-at.bofh.it> |
| In reply to | #1296171 |
Hello, Artem.
Can you please apply the following patch on top and see whether
anything changes? If it does make the issue go away, can you please
revert the ".can_queue" part and test again?
Thanks.
---
drivers/ata/ahci.h | 2 +-
drivers/ata/libahci.c | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
--- a/drivers/ata/ahci.h
+++ b/drivers/ata/ahci.h
@@ -365,7 +365,7 @@ extern struct device_attribute *ahci_sde
*/
#define AHCI_SHT(drv_name) \
ATA_NCQ_SHT(drv_name), \
- .can_queue = AHCI_MAX_CMDS - 1, \
+ .can_queue = 1/*AHCI_MAX_CMDS - 1*/, \
.sg_tablesize = AHCI_MAX_SG, \
.dma_boundary = AHCI_DMA_BOUNDARY, \
.shost_attrs = ahci_shost_attrs, \
--- a/drivers/ata/libahci.c
+++ b/drivers/ata/libahci.c
@@ -420,7 +420,7 @@ void ahci_save_initial_config(struct dev
hpriv->saved_cap2 = cap2 = 0;
/* some chips have errata preventing 64bit use */
- if ((cap & HOST_CAP_64) && (hpriv->flags & AHCI_HFLAG_32BIT_ONLY)) {
+ if ((cap & HOST_CAP_64)/* && (hpriv->flags & AHCI_HFLAG_32BIT_ONLY)*/) {
dev_info(dev, "controller can't do 64bit DMA, forcing 32bit\n");
cap &= ~HOST_CAP_64;
}
--
tejun
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-12-21 22:10 +0100 |
| Message-ID | <qIfSp-3eT-9@gated-at.bofh.it> |
| In reply to | #1296186 |
Hello, again. On Mon, Dec 21, 2015 at 03:07:21PM -0500, Tejun Heo wrote: > Hello, Artem. > > Can you please apply the following patch on top and see whether > anything changes? If it does make the issue go away, can you please > revert the ".can_queue" part and test again? If the patch doesn't change anything, can you please try the followings and see which one makes difference? 1. Exclude memory above 4G line with boot param "max_addr=4G". 2. Disable highmem with "highmem=0". 3. Try booting 64bit kernel. At the moment, the only thing I can think of which can explain the PAE + bio_get_nr_vecs() situation is that the bio split code which is activated by the bio_get_nr_vecs() somehow messes up 64bit or high addresses on 32bit kernels. I scanned for the obvious but at bio layer, memory is represented by struct page, so nothing obvious seems broken. Note that for now I'm ignoring the debug dumps from the ahci debug patch which indicates that the passed in addresses are all fine. It is possible that the controller gets confused with certain receiving addresses and reports failure on later commands or maybe there is a different sequence of events which can encompass both that we don't know of yet. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2015-12-22 04:50 +0100 |
| Message-ID | <qIm7v-75a-9@gated-at.bofh.it> |
| In reply to | #1296208 |
I just reproduced it - Artem, I'll let you know when we have a possible fix but hopefully there won't be any need for you to beat up your hardware any more :) On Mon, Dec 21, 2015 at 04:08:11PM -0500, Tejun Heo wrote: > Hello, again. > > On Mon, Dec 21, 2015 at 03:07:21PM -0500, Tejun Heo wrote: > > Hello, Artem. > > > > Can you please apply the following patch on top and see whether > > anything changes? If it does make the issue go away, can you please > > revert the ".can_queue" part and test again? > > If the patch doesn't change anything, can you please try the > followings and see which one makes difference? > > 1. Exclude memory above 4G line with boot param "max_addr=4G". > 2. Disable highmem with "highmem=0". > 3. Try booting 64bit kernel. > > At the moment, the only thing I can think of which can explain the PAE > + bio_get_nr_vecs() situation is that the bio split code which is > activated by the bio_get_nr_vecs() somehow messes up 64bit or high > addresses on 32bit kernels. I scanned for the obvious but at bio > layer, memory is represented by struct page, so nothing obvious seems > broken. > > Note that for now I'm ignoring the debug dumps from the ahci debug > patch which indicates that the passed in addresses are all fine. It > is possible that the controller gets confused with certain receiving > addresses and reports failure on later commands or maybe there is a > different sequence of events which can encompass both that we don't > know of yet. It repros pretty easily, I should be able to just write some test code that sends down single specific IOs so no matter what the controller is doing we know exactly what IO triggered the error. I'm checking all the 64 bit/pae/etc. stuff now. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2015-12-22 05:10 +0100 |
| Message-ID | <qImqR-7qS-7@gated-at.bofh.it> |
| In reply to | #1296208 |
On Mon, Dec 21, 2015 at 04:08:11PM -0500, Tejun Heo wrote: reproduced it with 32 bit pae: > 1. Exclude memory above 4G line with boot param "max_addr=4G". doesn't work - max_addr=1G doesn't work either > 2. Disable highmem with "highmem=0". works! > 3. Try booting 64bit kernel. works Ok, so maybe it actually is PAE specific... but like you noted the block layer works entirely in terms of pages so... The one idea I can think of is - maybe BIOVEC_PHYS_MERGEABLE() is broken in PAE mode? I am unfamiliar with anything PAE though. Where does the ahci code consume the sglist? -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Junichi Nomura <j-nomura@ce.jp.nec.com> |
|---|---|
| Date | 2015-12-22 06:30 +0100 |
| Message-ID | <qInGh-8ai-9@gated-at.bofh.it> |
| In reply to | #1296529 |
On 12/22/15 12:59, Kent Overstreet wrote:
> reproduced it with 32 bit pae:
>
>> 1. Exclude memory above 4G line with boot param "max_addr=4G".
>
> doesn't work - max_addr=1G doesn't work either
>
>> 2. Disable highmem with "highmem=0".
>
> works!
>
>> 3. Try booting 64bit kernel.
>
> works
blk_queue_bio() does split then bounce, which makes the segment
counting based on pages before bouncing and could go wrong.
What do you think of a patch like this?
--
Jun'ichi Nomura, NEC Corporation
diff --git a/block/blk-core.c b/block/blk-core.c
index 5131993b..1d1c3c7 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -1689,8 +1689,6 @@ static blk_qc_t blk_queue_bio(struct request_queue *q, struct bio *bio)
struct request *req;
unsigned int request_count = 0;
- blk_queue_split(q, &bio, q->bio_split);
-
/*
* low level driver can indicate that it wants pages above a
* certain limit bounced to low memory (ie for highmem, or even
@@ -1698,6 +1696,8 @@ static blk_qc_t blk_queue_bio(struct request_queue *q, struct bio *bio)
*/
blk_queue_bounce(q, &bio);
+ blk_queue_split(q, &bio, q->bio_split);
+
if (bio_integrity_enabled(bio) && bio_integrity_prep(bio)) {
bio->bi_error = -EIO;
bio_endio(bio);--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2015-12-22 06:40 +0100 |
| Message-ID | <qInPX-8dk-3@gated-at.bofh.it> |
| In reply to | #1296544 |
On Tue, Dec 22, 2015 at 05:26:12AM +0000, Junichi Nomura wrote:
> On 12/22/15 12:59, Kent Overstreet wrote:
> > reproduced it with 32 bit pae:
> >
> >> 1. Exclude memory above 4G line with boot param "max_addr=4G".
> >
> > doesn't work - max_addr=1G doesn't work either
> >
> >> 2. Disable highmem with "highmem=0".
> >
> > works!
> >
> >> 3. Try booting 64bit kernel.
> >
> > works
>
> blk_queue_bio() does split then bounce, which makes the segment
> counting based on pages before bouncing and could go wrong.
>
> What do you think of a patch like this?
Artem, can you give this patch a try?
>
> --
> Jun'ichi Nomura, NEC Corporation
>
> diff --git a/block/blk-core.c b/block/blk-core.c
> index 5131993b..1d1c3c7 100644
> --- a/block/blk-core.c
> +++ b/block/blk-core.c
> @@ -1689,8 +1689,6 @@ static blk_qc_t blk_queue_bio(struct request_queue *q, struct bio *bio)
> struct request *req;
> unsigned int request_count = 0;
>
> - blk_queue_split(q, &bio, q->bio_split);
> -
> /*
> * low level driver can indicate that it wants pages above a
> * certain limit bounced to low memory (ie for highmem, or even
> @@ -1698,6 +1696,8 @@ static blk_qc_t blk_queue_bio(struct request_queue *q, struct bio *bio)
> */
> blk_queue_bounce(q, &bio);
>
> + blk_queue_split(q, &bio, q->bio_split);
> +
> if (bio_integrity_enabled(bio) && bio_integrity_prep(bio)) {
> bio->bi_error = -EIO;
> bio_endio(bio);
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Artem S. Tashkinov" <t.artem@lycos.com> |
|---|---|
| Date | 2015-12-22 07:00 +0100 |
| Message-ID | <qIo9k-8jY-1@gated-at.bofh.it> |
| In reply to | #1296546 |
On 2015-12-22 10:38, Kent Overstreet wrote:
> On Tue, Dec 22, 2015 at 05:26:12AM +0000, Junichi Nomura wrote:
>> On 12/22/15 12:59, Kent Overstreet wrote:
>> > reproduced it with 32 bit pae:
>> >
>> >> 1. Exclude memory above 4G line with boot param "max_addr=4G".
>> >
>> > doesn't work - max_addr=1G doesn't work either
>> >
>> >> 2. Disable highmem with "highmem=0".
>> >
>> > works!
>> >
>> >> 3. Try booting 64bit kernel.
>> >
>> > works
>>
>> blk_queue_bio() does split then bounce, which makes the segment
>> counting based on pages before bouncing and could go wrong.
>>
>> What do you think of a patch like this?
>
> Artem, can you give this patch a try?
This patch ostensibly fixes the issue - at least I cannot immediately
reproduce it. You can count me in as "Tested-by: Artem S. Tashkinov"
>
>>
>> --
>> Jun'ichi Nomura, NEC Corporation
>>
>> diff --git a/block/blk-core.c b/block/blk-core.c
>> index 5131993b..1d1c3c7 100644
>> --- a/block/blk-core.c
>> +++ b/block/blk-core.c
>> @@ -1689,8 +1689,6 @@ static blk_qc_t blk_queue_bio(struct
>> request_queue *q, struct bio *bio)
>> struct request *req;
>> unsigned int request_count = 0;
>>
>> - blk_queue_split(q, &bio, q->bio_split);
>> -
>> /*
>> * low level driver can indicate that it wants pages above a
>> * certain limit bounced to low memory (ie for highmem, or even
>> @@ -1698,6 +1696,8 @@ static blk_qc_t blk_queue_bio(struct
>> request_queue *q, struct bio *bio)
>> */
>> blk_queue_bounce(q, &bio);
>>
>> + blk_queue_split(q, &bio, q->bio_split);
>> +
>> if (bio_integrity_enabled(bio) && bio_integrity_prep(bio)) {
>> bio->bi_error = -EIO;
>> bio_endio(bio);
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2015-12-22 07:00 +0100 |
| Message-ID | <qIo9k-8jY-5@gated-at.bofh.it> |
| In reply to | #1296557 |
On Tue, Dec 22, 2015 at 10:52:37AM +0500, Artem S. Tashkinov wrote: > On 2015-12-22 10:38, Kent Overstreet wrote: > >On Tue, Dec 22, 2015 at 05:26:12AM +0000, Junichi Nomura wrote: > >>On 12/22/15 12:59, Kent Overstreet wrote: > >>> reproduced it with 32 bit pae: > >>> > >>>> 1. Exclude memory above 4G line with boot param "max_addr=4G". > >>> > >>> doesn't work - max_addr=1G doesn't work either > >>> > >>>> 2. Disable highmem with "highmem=0". > >>> > >>> works! > >>> > >>>> 3. Try booting 64bit kernel. > >>> > >>> works > >> > >>blk_queue_bio() does split then bounce, which makes the segment > >>counting based on pages before bouncing and could go wrong. > >> > >>What do you think of a patch like this? > > > >Artem, can you give this patch a try? > > > This patch ostensibly fixes the issue - at least I cannot immediately > reproduce it. You can count me in as "Tested-by: Artem S. Tashkinov" Let's all contemplate the fact that blk_segment_map_sg() _overrunning the end of the provided sglist_ was this much of a clusterfuck to debug. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Artem S. Tashkinov" <t.artem@lycos.com> |
|---|---|
| Date | 2015-12-22 07:00 +0100 |
| Message-ID | <qIo9k-8jY-7@gated-at.bofh.it> |
| In reply to | #1296559 |
On 2015-12-22 10:55, Kent Overstreet wrote: > On Tue, Dec 22, 2015 at 10:52:37AM +0500, Artem S. Tashkinov wrote: >> On 2015-12-22 10:38, Kent Overstreet wrote: >> >On Tue, Dec 22, 2015 at 05:26:12AM +0000, Junichi Nomura wrote: >> >>On 12/22/15 12:59, Kent Overstreet wrote: >> >>> reproduced it with 32 bit pae: >> >>> >> >>>> 1. Exclude memory above 4G line with boot param "max_addr=4G". >> >>> >> >>> doesn't work - max_addr=1G doesn't work either >> >>> >> >>>> 2. Disable highmem with "highmem=0". >> >>> >> >>> works! >> >>> >> >>>> 3. Try booting 64bit kernel. >> >>> >> >>> works >> >> >> >>blk_queue_bio() does split then bounce, which makes the segment >> >>counting based on pages before bouncing and could go wrong. >> >> >> >>What do you think of a patch like this? >> > >> >Artem, can you give this patch a try? >> >> >> This patch ostensibly fixes the issue - at least I cannot immediately >> reproduce it. You can count me in as "Tested-by: Artem S. Tashkinov" > > Let's all contemplate the fact that blk_segment_map_sg() _overrunning > the end of > the provided sglist_ was this much of a clusterfuck to debug. From the look of it this fix has nothing to do with PAE, so then why only PAE users like me were affected by the original (b54ffb73cadcdcff9cc1ae0e11f502407e3e2e4c) patch? -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2015-12-22 07:10 +0100 |
| Message-ID | <qIoiZ-aT-13@gated-at.bofh.it> |
| In reply to | #1296560 |
On Tue, Dec 22, 2015 at 10:59:09AM +0500, Artem S. Tashkinov wrote: > On 2015-12-22 10:55, Kent Overstreet wrote: > >On Tue, Dec 22, 2015 at 10:52:37AM +0500, Artem S. Tashkinov wrote: > >>On 2015-12-22 10:38, Kent Overstreet wrote: > >>>On Tue, Dec 22, 2015 at 05:26:12AM +0000, Junichi Nomura wrote: > >>>>On 12/22/15 12:59, Kent Overstreet wrote: > >>>>> reproduced it with 32 bit pae: > >>>>> > >>>>>> 1. Exclude memory above 4G line with boot param "max_addr=4G". > >>>>> > >>>>> doesn't work - max_addr=1G doesn't work either > >>>>> > >>>>>> 2. Disable highmem with "highmem=0". > >>>>> > >>>>> works! > >>>>> > >>>>>> 3. Try booting 64bit kernel. > >>>>> > >>>>> works > >>>> > >>>>blk_queue_bio() does split then bounce, which makes the segment > >>>>counting based on pages before bouncing and could go wrong. > >>>> > >>>>What do you think of a patch like this? > >>> > >>>Artem, can you give this patch a try? > >> > >> > >>This patch ostensibly fixes the issue - at least I cannot immediately > >>reproduce it. You can count me in as "Tested-by: Artem S. Tashkinov" > > > >Let's all contemplate the fact that blk_segment_map_sg() _overrunning the > >end of > >the provided sglist_ was this much of a clusterfuck to debug. > > From the look of it this fix has nothing to do with PAE, so then why only > PAE users like me were affected by the original > (b54ffb73cadcdcff9cc1ae0e11f502407e3e2e4c) patch? The amusing thing is that I doubt PAE actually requires bouncing - addressing limits come from the device, not the cpu. But evidently in PAE mode, the block layer is in fact bouncing bios. Probably from some default setting in the queue limits that no one ever looks at. The whole queue limits design is an atrocity, it leads to exactly this kind of crap where no one can predict the actual behaviour of any given setup. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2015-12-22 06:40 +0100 |
| Message-ID | <qInPY-8dk-11@gated-at.bofh.it> |
| In reply to | #1296544 |
On Tue, Dec 22, 2015 at 05:26:12AM +0000, Junichi Nomura wrote:
> On 12/22/15 12:59, Kent Overstreet wrote:
> > reproduced it with 32 bit pae:
> >
> >> 1. Exclude memory above 4G line with boot param "max_addr=4G".
> >
> > doesn't work - max_addr=1G doesn't work either
> >
> >> 2. Disable highmem with "highmem=0".
> >
> > works!
> >
> >> 3. Try booting 64bit kernel.
> >
> > works
>
> blk_queue_bio() does split then bounce, which makes the segment
> counting based on pages before bouncing and could go wrong.
>
> What do you think of a patch like this?
Shit, you nailed it. Can't believe I didn't think to check that.
>
> --
> Jun'ichi Nomura, NEC Corporation
>
> diff --git a/block/blk-core.c b/block/blk-core.c
> index 5131993b..1d1c3c7 100644
> --- a/block/blk-core.c
> +++ b/block/blk-core.c
> @@ -1689,8 +1689,6 @@ static blk_qc_t blk_queue_bio(struct request_queue *q, struct bio *bio)
> struct request *req;
> unsigned int request_count = 0;
>
> - blk_queue_split(q, &bio, q->bio_split);
> -
> /*
> * low level driver can indicate that it wants pages above a
> * certain limit bounced to low memory (ie for highmem, or even
> @@ -1698,6 +1696,8 @@ static blk_qc_t blk_queue_bio(struct request_queue *q, struct bio *bio)
> */
> blk_queue_bounce(q, &bio);
>
> + blk_queue_split(q, &bio, q->bio_split);
> +
> if (bio_integrity_enabled(bio) && bio_integrity_prep(bio)) {
> bio->bi_error = -EIO;
> bio_endio(bio);
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2015-12-22 18:30 +0100 |
| Message-ID | <qIyV4-6R0-7@gated-at.bofh.it> |
| In reply to | #1296544 |
On 12/21/2015 10:26 PM, Junichi Nomura wrote:
> On 12/22/15 12:59, Kent Overstreet wrote:
>> reproduced it with 32 bit pae:
>>
>>> 1. Exclude memory above 4G line with boot param "max_addr=4G".
>>
>> doesn't work - max_addr=1G doesn't work either
>>
>>> 2. Disable highmem with "highmem=0".
>>
>> works!
>>
>>> 3. Try booting 64bit kernel.
>>
>> works
>
> blk_queue_bio() does split then bounce, which makes the segment
> counting based on pages before bouncing and could go wrong.
Good catch! The blk-mq parts aren't affected by this, the screw up only
happened in the old IO path. I've added this with the appropriate
tested-by from Artem, and CC stable and listed the commit that broke it:
commit 54efd50bfd873e2dbf784e0b21a8027ba4299a3e
Author: Kent Overstreet <kent.overstreet@gmail.com>
Date: Thu Apr 23 22:37:18 2015 -0700
block: make generic_make_request handle arbitrarily sized bios
Thanks to all involved in nailing this down, it'll go out shortly.
--
Jens Axboe
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web