Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1704238 > unrolled thread

[PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s

Started byHaris Okanovic <haris.okanovic@ni.com>
First post2017-08-05 00:00 +0200
Last post2017-08-09 00:00 +0200
Articles 3 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s Haris Okanovic <haris.okanovic@ni.com> - 2017-08-05 00:00 +0200
    Re: [PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s Julia Cartwright <julia@ni.com> - 2017-08-07 17:00 +0200
      Re: [PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s Jarkko Sakkinen <jarkko.sakkinen@linux.intel.com> - 2017-08-09 00:00 +0200

#1704238 — [PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s

FromHaris Okanovic <haris.okanovic@ni.com>
Date2017-08-05 00:00 +0200
Subject[PATCH] [RFC] tpm_tis: tpm_tcg_flush() after iowrite*()s
Message-ID<uaSNr-47E-9@gated-at.bofh.it>
I have a latency issue using a SPI-based TPM chip with tpm_tis driver
from non-rt usermode application, which induces ~400 us latency spikes
in cyclictest (Intel Atom E3940 system, PREEMPT_RT_FULL kernel).

The spikes are caused by a stalling ioread8() operation, following a
sequence of 30+ iowrite8()s to the same address. I believe this happens
because the writes are cached (in cpu or somewhere along the bus), which
gets flushed on the first LOAD instruction (ioread*()) that follows.

The enclosed change appears to fix this issue: read the TPM chip's
access register (status code) after every iowrite*() operation.

I believe this works because it amortize the cost of flushing data to
chip across multiple instructions. However, I don't have any direct
evidence to support this theory.

Does this seem like a reasonable theory?

Any feedback on the change (a better way to do it, perhaps)?

Thanks,
Haris Okanovic

https://github.com/harisokanovic/linux/tree/dev/hokanovi/tpm-latency-spike-fix-rfc
---
 drivers/char/tpm/tpm_tis.c | 18 +++++++++++++++++-
 1 file changed, 17 insertions(+), 1 deletion(-)

diff --git a/drivers/char/tpm/tpm_tis.c b/drivers/char/tpm/tpm_tis.c
index c7e1384f1b08..5cdbfec0ad67 100644
--- a/drivers/char/tpm/tpm_tis.c
+++ b/drivers/char/tpm/tpm_tis.c
@@ -89,6 +89,19 @@ static inline int is_itpm(struct acpi_device *dev)
 }
 #endif
 
+#ifdef CONFIG_PREEMPT_RT_FULL
+/*
+ * Flushes previous iowrite*() operations to chip so that a subsequent
+ * ioread*() won't stall a cpu.
+ */
+static void tpm_tcg_flush(struct tpm_tis_tcg_phy *phy)
+{
+	ioread8(phy->iobase + TPM_ACCESS(0));
+}
+#else
+#define tpm_tcg_flush do { } while(0)
+#endif
+
 static int tpm_tcg_read_bytes(struct tpm_tis_data *data, u32 addr, u16 len,
 			      u8 *result)
 {
@@ -104,8 +117,10 @@ static int tpm_tcg_write_bytes(struct tpm_tis_data *data, u32 addr, u16 len,
 {
 	struct tpm_tis_tcg_phy *phy = to_tpm_tis_tcg_phy(data);
 
-	while (len--)
+	while (len--) {
 		iowrite8(*value++, phy->iobase + addr);
+		tpm_tcg_flush(phy);
+	}
 	return 0;
 }
 
@@ -130,6 +145,7 @@ static int tpm_tcg_write32(struct tpm_tis_data *data, u32 addr, u32 value)
 	struct tpm_tis_tcg_phy *phy = to_tpm_tis_tcg_phy(data);
 
 	iowrite32(value, phy->iobase + addr);
+	tpm_tcg_flush(phy);
 	return 0;
 }
 
-- 
2.13.2

[toc] | [next] | [standalone]


#1705603

FromJulia Cartwright <julia@ni.com>
Date2017-08-07 17:00 +0200
Message-ID<ubRFF-1cv-43@gated-at.bofh.it>
In reply to#1704238

[Multipart message — attachments visible in raw view] — view raw

On Fri, Aug 04, 2017 at 04:56:51PM -0500, Haris Okanovic wrote:
> I have a latency issue using a SPI-based TPM chip with tpm_tis driver
> from non-rt usermode application, which induces ~400 us latency spikes
> in cyclictest (Intel Atom E3940 system, PREEMPT_RT_FULL kernel).
>
> The spikes are caused by a stalling ioread8() operation, following a
> sequence of 30+ iowrite8()s to the same address. I believe this happens
> because the writes are cached (in cpu or somewhere along the bus), which
> gets flushed on the first LOAD instruction (ioread*()) that follows.

To use the ARM parlance, these accesses aren't "cached" (which would
imply that a result could be returned to the load from any intermediate
node in the interconnect), but instead are "bufferable".

It is really unfortunate that we continue to run into this class of
problem across various CPU vendors and various underlying bus
technologies; it's the continuing curse of running an PREEMPT_RT on
commodity hardware.  RT is not easy :)

> The enclosed change appears to fix this issue: read the TPM chip's
> access register (status code) after every iowrite*() operation.

Are we engaged in a game of wack-a-mole with all of the drivers which
use this same access pattern (of which I imagine there are quite a
few!)?

I'm wondering if we should explore the idea of adding a load in the
iowriteN()/writeX() macros (marking those accesses in which reads cause
side effects explicitly, redirecting to a _raw() variant or something).

Obviously that would be expensive for non-RT use cases, but for helping
constrain latency, it may be worth it for RT.

   Julia

[toc] | [prev] | [next] | [standalone]


#1706902

FromJarkko Sakkinen <jarkko.sakkinen@linux.intel.com>
Date2017-08-09 00:00 +0200
Message-ID<uckHF-64i-37@gated-at.bofh.it>
In reply to#1705603
On Mon, Aug 07, 2017 at 09:59:35AM -0500, Julia Cartwright wrote:
> On Fri, Aug 04, 2017 at 04:56:51PM -0500, Haris Okanovic wrote:
> > I have a latency issue using a SPI-based TPM chip with tpm_tis driver
> > from non-rt usermode application, which induces ~400 us latency spikes
> > in cyclictest (Intel Atom E3940 system, PREEMPT_RT_FULL kernel).
> >
> > The spikes are caused by a stalling ioread8() operation, following a
> > sequence of 30+ iowrite8()s to the same address. I believe this happens
> > because the writes are cached (in cpu or somewhere along the bus), which
> > gets flushed on the first LOAD instruction (ioread*()) that follows.
> 
> To use the ARM parlance, these accesses aren't "cached" (which would
> imply that a result could be returned to the load from any intermediate
> node in the interconnect), but instead are "bufferable".
> 
> It is really unfortunate that we continue to run into this class of
> problem across various CPU vendors and various underlying bus
> technologies; it's the continuing curse of running an PREEMPT_RT on
> commodity hardware.  RT is not easy :)
> 
> > The enclosed change appears to fix this issue: read the TPM chip's
> > access register (status code) after every iowrite*() operation.
> 
> Are we engaged in a game of wack-a-mole with all of the drivers which
> use this same access pattern (of which I imagine there are quite a
> few!)?
> 
> I'm wondering if we should explore the idea of adding a load in the
> iowriteN()/writeX() macros (marking those accesses in which reads cause
> side effects explicitly, redirecting to a _raw() variant or something).
> 
> Obviously that would be expensive for non-RT use cases, but for helping
> constrain latency, it may be worth it for RT.
> 
>    Julia

What if we as quick resort we add tpm_tis_iowrite8() to the TPM driver.
Would be easy to move to iowrite8() if the problem is sorted out there
later on.

/Jarkko

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web