Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1704238 > unrolled thread
| Started by | Haris Okanovic <haris.okanovic@ni.com> |
|---|---|
| First post | 2017-08-05 00:00 +0200 |
| Last post | 2017-08-09 00:00 +0200 |
| Articles | 3 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s Haris Okanovic <haris.okanovic@ni.com> - 2017-08-05 00:00 +0200
Re: [PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s Julia Cartwright <julia@ni.com> - 2017-08-07 17:00 +0200
Re: [PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s Jarkko Sakkinen <jarkko.sakkinen@linux.intel.com> - 2017-08-09 00:00 +0200
| From | Haris Okanovic <haris.okanovic@ni.com> |
|---|---|
| Date | 2017-08-05 00:00 +0200 |
| Subject | [PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s |
| Message-ID | <uaSNr-47E-9@gated-at.bofh.it> |
I have a latency issue using a SPI-based TPM chip with tpm_tis driver
from non-rt usermode application, which induces ~400 us latency spikes
in cyclictest (Intel Atom E3940 system, PREEMPT_RT_FULL kernel).
The spikes are caused by a stalling ioread8() operation, following a
sequence of 30+ iowrite8()s to the same address. I believe this happens
because the writes are cached (in cpu or somewhere along the bus), which
gets flushed on the first LOAD instruction (ioread*()) that follows.
The enclosed change appears to fix this issue: read the TPM chip's
access register (status code) after every iowrite*() operation.
I believe this works because it amortize the cost of flushing data to
chip across multiple instructions. However, I don't have any direct
evidence to support this theory.
Does this seem like a reasonable theory?
Any feedback on the change (a better way to do it, perhaps)?
Thanks,
Haris Okanovic
https://github.com/harisokanovic/linux/tree/dev/hokanovi/tpm-latency-spike-fix-rfc
---
drivers/char/tpm/tpm_tis.c | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
diff --git a/drivers/char/tpm/tpm_tis.c b/drivers/char/tpm/tpm_tis.c
index c7e1384f1b08..5cdbfec0ad67 100644
--- a/drivers/char/tpm/tpm_tis.c
+++ b/drivers/char/tpm/tpm_tis.c
@@ -89,6 +89,19 @@ static inline int is_itpm(struct acpi_device *dev)
}
#endif
+#ifdef CONFIG_PREEMPT_RT_FULL
+/*
+ * Flushes previous iowrite*() operations to chip so that a subsequent
+ * ioread*() won't stall a cpu.
+ */
+static void tpm_tcg_flush(struct tpm_tis_tcg_phy *phy)
+{
+ ioread8(phy->iobase + TPM_ACCESS(0));
+}
+#else
+#define tpm_tcg_flush do { } while(0)
+#endif
+
static int tpm_tcg_read_bytes(struct tpm_tis_data *data, u32 addr, u16 len,
u8 *result)
{
@@ -104,8 +117,10 @@ static int tpm_tcg_write_bytes(struct tpm_tis_data *data, u32 addr, u16 len,
{
struct tpm_tis_tcg_phy *phy = to_tpm_tis_tcg_phy(data);
- while (len--)
+ while (len--) {
iowrite8(*value++, phy->iobase + addr);
+ tpm_tcg_flush(phy);
+ }
return 0;
}
@@ -130,6 +145,7 @@ static int tpm_tcg_write32(struct tpm_tis_data *data, u32 addr, u32 value)
struct tpm_tis_tcg_phy *phy = to_tpm_tis_tcg_phy(data);
iowrite32(value, phy->iobase + addr);
+ tpm_tcg_flush(phy);
return 0;
}
--
2.13.2
[toc] | [next] | [standalone]
| From | Julia Cartwright <julia@ni.com> |
|---|---|
| Date | 2017-08-07 17:00 +0200 |
| Message-ID | <ubRFF-1cv-43@gated-at.bofh.it> |
| In reply to | #1704238 |
[Multipart message — attachments visible in raw view] — view raw
On Fri, Aug 04, 2017 at 04:56:51PM -0500, Haris Okanovic wrote: > I have a latency issue using a SPI-based TPM chip with tpm_tis driver > from non-rt usermode application, which induces ~400 us latency spikes > in cyclictest (Intel Atom E3940 system, PREEMPT_RT_FULL kernel). > > The spikes are caused by a stalling ioread8() operation, following a > sequence of 30+ iowrite8()s to the same address. I believe this happens > because the writes are cached (in cpu or somewhere along the bus), which > gets flushed on the first LOAD instruction (ioread*()) that follows. To use the ARM parlance, these accesses aren't "cached" (which would imply that a result could be returned to the load from any intermediate node in the interconnect), but instead are "bufferable". It is really unfortunate that we continue to run into this class of problem across various CPU vendors and various underlying bus technologies; it's the continuing curse of running an PREEMPT_RT on commodity hardware. RT is not easy :) > The enclosed change appears to fix this issue: read the TPM chip's > access register (status code) after every iowrite*() operation. Are we engaged in a game of wack-a-mole with all of the drivers which use this same access pattern (of which I imagine there are quite a few!)? I'm wondering if we should explore the idea of adding a load in the iowriteN()/writeX() macros (marking those accesses in which reads cause side effects explicitly, redirecting to a _raw() variant or something). Obviously that would be expensive for non-RT use cases, but for helping constrain latency, it may be worth it for RT. Julia
[toc] | [prev] | [next] | [standalone]
| From | Jarkko Sakkinen <jarkko.sakkinen@linux.intel.com> |
|---|---|
| Date | 2017-08-09 00:00 +0200 |
| Message-ID | <uckHF-64i-37@gated-at.bofh.it> |
| In reply to | #1705603 |
On Mon, Aug 07, 2017 at 09:59:35AM -0500, Julia Cartwright wrote: > On Fri, Aug 04, 2017 at 04:56:51PM -0500, Haris Okanovic wrote: > > I have a latency issue using a SPI-based TPM chip with tpm_tis driver > > from non-rt usermode application, which induces ~400 us latency spikes > > in cyclictest (Intel Atom E3940 system, PREEMPT_RT_FULL kernel). > > > > The spikes are caused by a stalling ioread8() operation, following a > > sequence of 30+ iowrite8()s to the same address. I believe this happens > > because the writes are cached (in cpu or somewhere along the bus), which > > gets flushed on the first LOAD instruction (ioread*()) that follows. > > To use the ARM parlance, these accesses aren't "cached" (which would > imply that a result could be returned to the load from any intermediate > node in the interconnect), but instead are "bufferable". > > It is really unfortunate that we continue to run into this class of > problem across various CPU vendors and various underlying bus > technologies; it's the continuing curse of running an PREEMPT_RT on > commodity hardware. RT is not easy :) > > > The enclosed change appears to fix this issue: read the TPM chip's > > access register (status code) after every iowrite*() operation. > > Are we engaged in a game of wack-a-mole with all of the drivers which > use this same access pattern (of which I imagine there are quite a > few!)? > > I'm wondering if we should explore the idea of adding a load in the > iowriteN()/writeX() macros (marking those accesses in which reads cause > side effects explicitly, redirecting to a _raw() variant or something). > > Obviously that would be expensive for non-RT use cases, but for helping > constrain latency, it may be worth it for RT. > > Julia What if we as quick resort we add tpm_tis_iowrite8() to the TPM driver. Would be easy to move to iowrite8() if the problem is sorted out there later on. /Jarkko
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web