Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1585645 > unrolled thread

[PATCH v3 00/25] OpenRISC patches for 4.11 final call

Started byStafford Horne <shorne@gmail.com>
First post2017-02-21 20:20 +0100
Last post2017-02-21 20:30 +0100
Articles 20 on this page of 36 — 4 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v3 00/25] OpenRISC patches for 4.11 final call Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 16/25] openrisc: Add optimized memcpy routine Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 13/25] openrisc: Fix the bitmask for the unit present register Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 05/25] openrisc: head: refactor out tlb flush into it's own function Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 23/25] arch/openrisc/lib/memcpy.c: use correct OR1200 option Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 19/25] scripts/checkstack.pl: Add openrisc support Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 24/25] openrisc: Export ioremap symbols used by modules Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 21/25] openrisc: entry: Fix delay slot detection Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 04/25] openrisc: head: use THREAD_SIZE instead of magic constant Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 02/25] openrisc: add cache way information to cpuinfo Stafford Horne <shorne@gmail.com> - 2017-02-21 20:20 +0100
    [PATCH v3 14/25] openrisc: Initial support for the idle state Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
      Re: [PATCH v3 14/25] openrisc: Initial support for the idle state Joe Perches <joe@perches.com> - 2017-02-21 21:30 +0100
        Re: [PATCH v3 14/25] openrisc: Initial support for the idle state Stafford Horne <shorne@gmail.com> - 2017-02-22 15:20 +0100
    [PATCH v3 18/25] MAINTAINERS: Add the openrisc official repository Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
    [PATCH v3 07/25] openrisc: add atomic bitops Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
    [PATCH v3 09/25] openrisc: add optimized atomic operations Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
      Re: [PATCH v3 09/25] openrisc: add optimized atomic operations Peter Zijlstra <peterz@infradead.org> - 2017-02-22 12:30 +0100
        Re: [PATCH v3 09/25] openrisc: add optimized atomic operations Stafford Horne <shorne@gmail.com> - 2017-02-22 15:30 +0100
          Re: [OpenRISC] [PATCH v3 09/25] openrisc: add optimized atomic  operations Richard Henderson <rth@twiddle.net> - 2017-02-22 18:40 +0100
            Re: [OpenRISC] [PATCH v3 09/25] openrisc: add optimized atomic  operations Stafford Horne <shorne@gmail.com> - 2017-02-22 23:50 +0100
    [PATCH v3 11/25] openrisc: add futex_atomic_* implementations Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
    [PATCH v3 17/25] openrisc: Add .gitignore Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
    [PATCH v3 08/25] openrisc: add cmpxchg and xchg implementations Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
      Re: [PATCH v3 08/25] openrisc: add cmpxchg and xchg implementations Peter Zijlstra <peterz@infradead.org> - 2017-02-22 12:30 +0100
        Re: [PATCH v3 08/25] openrisc: add cmpxchg and xchg implementations Stafford Horne <shorne@gmail.com> - 2017-02-22 15:30 +0100
          Re: [OpenRISC] [PATCH v3 08/25] openrisc: add cmpxchg and xchg  implementations Richard Henderson <rth@twiddle.net> - 2017-02-22 18:40 +0100
            Re: [OpenRISC] [PATCH v3 08/25] openrisc: add cmpxchg and xchg  implementations Stafford Horne <shorne@gmail.com> - 2017-02-22 23:50 +0100
    [PATCH v3 10/25] openrisc: add spinlock implementation Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100
      Re: [PATCH v3 10/25] openrisc: add spinlock implementation Peter Zijlstra <peterz@infradead.org> - 2017-02-22 12:30 +0100
      Re: [PATCH v3 10/25] openrisc: add spinlock implementation Peter Zijlstra <peterz@infradead.org> - 2017-02-22 12:40 +0100
      Re: [PATCH v3 10/25] openrisc: add spinlock implementation Peter Zijlstra <peterz@infradead.org> - 2017-02-22 12:40 +0100
        Re: [PATCH v3 10/25] openrisc: add spinlock implementation Peter Zijlstra <peterz@infradead.org> - 2017-02-22 13:10 +0100
      Re: [PATCH v3 10/25] openrisc: add spinlock implementation Peter Zijlstra <peterz@infradead.org> - 2017-02-22 12:40 +0100
      Re: [PATCH v3 10/25] openrisc: add spinlock implementation Peter Zijlstra <peterz@infradead.org> - 2017-02-22 12:50 +0100
        Re: [PATCH v3 10/25] openrisc: add spinlock implementation Peter Zijlstra <peterz@infradead.org> - 2017-02-22 13:10 +0100
    [PATCH v3 20/25] openrisc: entry: Whitespace and comment cleanups Stafford Horne <shorne@gmail.com> - 2017-02-21 20:30 +0100

Page 1 of 2  [1] 2  Next page →


#1585645 — [PATCH v3 00/25] OpenRISC patches for 4.11 final call

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 00/25] OpenRISC patches for 4.11 final call
Message-ID<tdo8F-6Sm-3@gated-at.bofh.it>
Hi All,

Changes from v2
 o implemented all atomic ops pointed out by Peter Z
 o export ioremap symbols pointed out by allyesconfig
 o init r0 to 0 as per openrisc spec, suggested by Jakob Viketoft

Changes from v1
 o added change set from Valentin catching CONFIG issues
 o added missing test_and_change_bit atomic bitops patch

This is a attempt to get some final comments before I send the pull request
for this merge windot to Linus.  This comes up due to last minute findings
by Peter Zijlstra with the atomic operations patch.  If I get a bunch of
Nak's I will likely just skip this merge window, but lets hope that can be
avoided.

Any feedback is appreciated.

The interesting things here are:
 - optimized memset and memcpy routines, ~20% boot time saving
 - support for cpu idling
 - adding support for l.swa and l.lwa atomic operations (in spec from 2014)
 - use atomics to implement: bitops, cmpxchg, futex, spinlocks
 - the atomics are in preparation for SMP support

Testing:
I have used the kselftests to validate the changes especially the futex
operations with the futex test. Other atomic operations are common so no
explicit testing.  I have mainly done the tests on qemu.

Note for testers:
The l.swa and l.lwa emulation + native instruction support is NOW FIXED in
qemu upstream git.

I have send patches for get a recent openrisc toolchain added to the
lkp-tests make.cross script.  It should be up to date now I believe 
this is what most build sytems use. Let me know if different.

-Stafford


Jonas Bonn (1):
  openrisc: use SPARSE_IRQ

Olof Kindgren (1):
  openrisc: Add optimized memset

Sebastian Macke (2):
  openrisc: Fix the bitmask for the unit present register
  openrisc: Initial support for the idle state

Stafford Horne (9):
  openrisc: Add optimized memcpy routine
  openrisc: Add .gitignore
  MAINTAINERS: Add the openrisc official repository
  scripts/checkstack.pl: Add openrisc support
  openrisc: entry: Whitespace and comment cleanups
  openrisc: entry: Fix delay slot detection
  openrisc: head: Move init strings to rodata section
  openrisc: Export ioremap symbols used by modules
  openrisc: head: Init r0 to 0 on start

Stefan Kristiansson (11):
  openrisc: add cache way information to cpuinfo
  openrisc: tlb miss handler optimizations
  openrisc: head: use THREAD_SIZE instead of magic constant
  openrisc: head: refactor out tlb flush into it's own function
  openrisc: add l.lwa/l.swa emulation
  openrisc: add atomic bitops
  openrisc: add cmpxchg and xchg implementations
  openrisc: add optimized atomic operations
  openrisc: add spinlock implementation
  openrisc: add futex_atomic_* implementations
  openrisc: remove unnecessary stddef.h include

Valentin Rothberg (1):
  arch/openrisc/lib/memcpy.c: use correct OR1200 option

 MAINTAINERS                                |   1 +
 arch/openrisc/Kconfig                      |   1 +
 arch/openrisc/TODO.openrisc                |   1 -
 arch/openrisc/include/asm/Kbuild           |   5 +-
 arch/openrisc/include/asm/atomic.h         | 100 +++++++++++++
 arch/openrisc/include/asm/bitops.h         |   2 +-
 arch/openrisc/include/asm/bitops/atomic.h  | 123 +++++++++++++++
 arch/openrisc/include/asm/cmpxchg.h        |  82 ++++++++++
 arch/openrisc/include/asm/cpuinfo.h        |   2 +
 arch/openrisc/include/asm/futex.h          | 135 +++++++++++++++++
 arch/openrisc/include/asm/spinlock.h       | 232 ++++++++++++++++++++++++++++-
 arch/openrisc/include/asm/spinlock_types.h |  28 ++++
 arch/openrisc/include/asm/spr_defs.h       |   4 +-
 arch/openrisc/include/asm/string.h         |  10 ++
 arch/openrisc/kernel/.gitignore            |   1 +
 arch/openrisc/kernel/entry.S               |  60 +++++---
 arch/openrisc/kernel/head.S                | 200 ++++++++++---------------
 arch/openrisc/kernel/or32_ksyms.c          |   1 +
 arch/openrisc/kernel/process.c             |  17 +++
 arch/openrisc/kernel/ptrace.c              |   1 -
 arch/openrisc/kernel/setup.c               |  67 +++++----
 arch/openrisc/kernel/traps.c               | 183 +++++++++++++++++++++++
 arch/openrisc/lib/Makefile                 |   2 +-
 arch/openrisc/lib/memcpy.c                 | 124 +++++++++++++++
 arch/openrisc/lib/memset.S                 |  98 ++++++++++++
 arch/openrisc/mm/ioremap.c                 |   2 +
 scripts/checkstack.pl                      |   3 +
 27 files changed, 1297 insertions(+), 188 deletions(-)
 create mode 100644 arch/openrisc/include/asm/atomic.h
 create mode 100644 arch/openrisc/include/asm/bitops/atomic.h
 create mode 100644 arch/openrisc/include/asm/cmpxchg.h
 create mode 100644 arch/openrisc/include/asm/futex.h
 create mode 100644 arch/openrisc/include/asm/spinlock_types.h
 create mode 100644 arch/openrisc/include/asm/string.h
 create mode 100644 arch/openrisc/kernel/.gitignore
 create mode 100644 arch/openrisc/lib/memcpy.c
 create mode 100644 arch/openrisc/lib/memset.S

-- 
2.9.3

[toc] | [next] | [standalone]


#1585647 — [PATCH v3 16/25] openrisc: Add optimized memcpy routine

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 16/25] openrisc: Add optimized memcpy routine
Message-ID<tdo8H-6Sm-43@gated-at.bofh.it>
In reply to#1585645
The generic memcpy routine provided in kernel does only byte copies.
Using word copies we can lower boot time and cycles spend in memcpy
quite significantly.

Booting on my de0 nano I see boot times go from 7.2 to 5.6 seconds.
The avg cycles in memcpy during boot go from 6467 to 1887.

I tested several algorithms (see code in previous patch mails)

The implementations I tested and avg cycles:
  - Word Copies + Loop Unrolls + Non Aligned    1882
  - Word Copies + Loop Unrolls                  1887
  - Word Copies                                 2441
  - Byte Copies + Loop Unrolls                  6467
  - Byte Copies                                 7600

In the end I ended up going with Word Copies + Loop Unrolls as it
provides best tradeoff between simplicity and boot speedups.

Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/TODO.openrisc        |   1 -
 arch/openrisc/include/asm/string.h |   3 +
 arch/openrisc/lib/Makefile         |   2 +-
 arch/openrisc/lib/memcpy.c         | 124 +++++++++++++++++++++++++++++++++++++
 4 files changed, 128 insertions(+), 2 deletions(-)
 create mode 100644 arch/openrisc/lib/memcpy.c

diff --git a/arch/openrisc/TODO.openrisc b/arch/openrisc/TODO.openrisc
index 0eb04c8..c43d4e1 100644
--- a/arch/openrisc/TODO.openrisc
+++ b/arch/openrisc/TODO.openrisc
@@ -10,4 +10,3 @@ that are due for investigation shortly, i.e. our TODO list:
    or1k and this change is slowly trickling through the stack.  For the time
    being, or32 is equivalent to or1k.
 
--- Implement optimized version of memcpy and memset
diff --git a/arch/openrisc/include/asm/string.h b/arch/openrisc/include/asm/string.h
index 33470d4..64939cc 100644
--- a/arch/openrisc/include/asm/string.h
+++ b/arch/openrisc/include/asm/string.h
@@ -4,4 +4,7 @@
 #define __HAVE_ARCH_MEMSET
 extern void *memset(void *s, int c, __kernel_size_t n);
 
+#define __HAVE_ARCH_MEMCPY
+extern void *memcpy(void *dest, __const void *src, __kernel_size_t n);
+
 #endif /* __ASM_OPENRISC_STRING_H */
diff --git a/arch/openrisc/lib/Makefile b/arch/openrisc/lib/Makefile
index 67c583e..17d9d37 100644
--- a/arch/openrisc/lib/Makefile
+++ b/arch/openrisc/lib/Makefile
@@ -2,4 +2,4 @@
 # Makefile for or32 specific library files..
 #
 
-obj-y  = memset.o string.o delay.o
+obj-y	:= delay.o string.o memset.o memcpy.o
diff --git a/arch/openrisc/lib/memcpy.c b/arch/openrisc/lib/memcpy.c
new file mode 100644
index 0000000..4706f01
--- /dev/null
+++ b/arch/openrisc/lib/memcpy.c
@@ -0,0 +1,124 @@
+/*
+ * arch/openrisc/lib/memcpy.c
+ *
+ * Optimized memory copy routines for openrisc.  These are mostly copied
+ * from ohter sources but slightly entended based on ideas discuassed in
+ * #openrisc.
+ *
+ * The word unroll implementation is an extension to the arm byte
+ * unrolled implementation, but using word copies (if things are
+ * properly aligned)
+ *
+ * The great arm loop unroll algorithm can be found at:
+ *  arch/arm/boot/compressed/string.c
+ */
+
+#include <linux/export.h>
+
+#include <linux/string.h>
+
+#ifdef CONFIG_OR1200
+/*
+ * Do memcpy with word copies and loop unrolling. This gives the
+ * best performance on the OR1200 and MOR1KX archirectures
+ */
+void *memcpy(void *dest, __const void *src, __kernel_size_t n)
+{
+	int i = 0;
+	unsigned char *d, *s;
+	uint32_t *dest_w = (uint32_t *)dest, *src_w = (uint32_t *)src;
+
+	/* If both source and dest are word aligned copy words */
+	if (!((unsigned int)dest_w & 3) && !((unsigned int)src_w & 3)) {
+		/* Copy 32 bytes per loop */
+		for (i = n >> 5; i > 0; i--) {
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+		}
+
+		if (n & 1 << 4) {
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+		}
+
+		if (n & 1 << 3) {
+			*dest_w++ = *src_w++;
+			*dest_w++ = *src_w++;
+		}
+
+		if (n & 1 << 2)
+			*dest_w++ = *src_w++;
+
+		d = (unsigned char *)dest_w;
+		s = (unsigned char *)src_w;
+
+	} else {
+		d = (unsigned char *)dest_w;
+		s = (unsigned char *)src_w;
+
+		for (i = n >> 3; i > 0; i--) {
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+		}
+
+		if (n & 1 << 2) {
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+			*d++ = *s++;
+		}
+	}
+
+	if (n & 1 << 1) {
+		*d++ = *s++;
+		*d++ = *s++;
+	}
+
+	if (n & 1)
+		*d++ = *s++;
+
+	return dest;
+}
+#else
+/*
+ * Use word copies but no loop unrolling as we cannot assume there
+ * will be benefits on the archirecture
+ */
+void *memcpy(void *dest, __const void *src, __kernel_size_t n)
+{
+	unsigned char *d = (unsigned char *)dest, *s = (unsigned char *)src;
+	uint32_t *dest_w = (uint32_t *)dest, *src_w = (uint32_t *)src;
+
+	/* If both source and dest are word aligned copy words */
+	if (!((unsigned int)dest_w & 3) && !((unsigned int)src_w & 3)) {
+		for (; n >= 4; n -= 4)
+			*dest_w++ = *src_w++;
+	}
+
+	d = (unsigned char *)dest_w;
+	s = (unsigned char *)src_w;
+
+	/* For remaining or if not aligned, copy bytes */
+	for (; n >= 1; n -= 1)
+		*d++ = *s++;
+
+	return dest;
+
+}
+#endif
+
+EXPORT_SYMBOL(memcpy);
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585649 — [PATCH v3 13/25] openrisc: Fix the bitmask for the unit present register

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 13/25] openrisc: Fix the bitmask for the unit present register
Message-ID<tdo8H-6Sm-49@gated-at.bofh.it>
In reply to#1585645
From: Sebastian Macke <sebastian@macke.de>

The bits were swapped, as per spec and processor implementation the
power management present bit is 9 and PIC bit is 8. This patch brings
the definitions into spec.

Signed-off-by: Sebastian Macke <sebastian@macke.de>
[shorne@gmail.com: Added commit body]
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/include/asm/spr_defs.h | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/arch/openrisc/include/asm/spr_defs.h b/arch/openrisc/include/asm/spr_defs.h
index 5dbc668..367dac7 100644
--- a/arch/openrisc/include/asm/spr_defs.h
+++ b/arch/openrisc/include/asm/spr_defs.h
@@ -152,8 +152,8 @@
 #define SPR_UPR_MP	   0x00000020  /* MAC present */
 #define SPR_UPR_DUP	   0x00000040  /* Debug unit present */
 #define SPR_UPR_PCUP	   0x00000080  /* Performance counters unit present */
-#define SPR_UPR_PMP	   0x00000100  /* Power management present */
-#define SPR_UPR_PICP	   0x00000200  /* PIC present */
+#define SPR_UPR_PICP	   0x00000100  /* PIC present */
+#define SPR_UPR_PMP	   0x00000200  /* Power management present */
 #define SPR_UPR_TTP	   0x00000400  /* Tick timer present */
 #define SPR_UPR_RES	   0x00fe0000  /* Reserved */
 #define SPR_UPR_CUP	   0xff000000  /* Context units present */
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585650 — [PATCH v3 05/25] openrisc: head: refactor out tlb flush into it's own function

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 05/25] openrisc: head: refactor out tlb flush into it's own function
Message-ID<tdo8I-6Sm-65@gated-at.bofh.it>
In reply to#1585645
From: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>

This brings it inline with the other setup oprations done like the cache
enables _ic_enable and _dc_enable.  Also, this is going to make it
easier to initialize additional cpu's when smp is introduced.

Signed-off-by: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
[shorne@gmail.com: Added commit body]
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/kernel/head.S | 38 ++++++++++++++++++++++----------------
 1 file changed, 22 insertions(+), 16 deletions(-)

diff --git a/arch/openrisc/kernel/head.S b/arch/openrisc/kernel/head.S
index 63ba2d9..a22f1fc 100644
--- a/arch/openrisc/kernel/head.S
+++ b/arch/openrisc/kernel/head.S
@@ -522,22 +522,8 @@ enable_dc:
 	 l.nop
 
 flush_tlb:
-	/*
-	 *  I N V A L I D A T E   T L B   e n t r i e s
-	 */
-	LOAD_SYMBOL_2_GPR(r5,SPR_DTLBMR_BASE(0))
-	LOAD_SYMBOL_2_GPR(r6,SPR_ITLBMR_BASE(0))
-	l.addi	r7,r0,128 /* Maximum number of sets */
-1:
-	l.mtspr	r5,r0,0x0
-	l.mtspr	r6,r0,0x0
-
-	l.addi	r5,r5,1
-	l.addi	r6,r6,1
-	l.sfeq	r7,r0
-	l.bnf	1b
-	 l.addi	r7,r7,-1
-
+	l.jal	_flush_tlb
+	 l.nop
 
 /* The MMU needs to be enabled before or32_early_setup is called */
 
@@ -629,6 +615,26 @@ jump_start_kernel:
 	l.jr    r30
 	 l.nop
 
+_flush_tlb:
+	/*
+	 *  I N V A L I D A T E   T L B   e n t r i e s
+	 */
+	LOAD_SYMBOL_2_GPR(r5,SPR_DTLBMR_BASE(0))
+	LOAD_SYMBOL_2_GPR(r6,SPR_ITLBMR_BASE(0))
+	l.addi	r7,r0,128 /* Maximum number of sets */
+1:
+	l.mtspr	r5,r0,0x0
+	l.mtspr	r6,r0,0x0
+
+	l.addi	r5,r5,1
+	l.addi	r6,r6,1
+	l.sfeq	r7,r0
+	l.bnf	1b
+	 l.addi	r7,r7,-1
+
+	l.jr	r9
+	 l.nop
+
 /* ========================================[ cache ]=== */
 
 	/* aligment here so we don't change memory offsets with
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585655 — [PATCH v3 23/25] arch/openrisc/lib/memcpy.c: use correct OR1200 option

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 23/25] arch/openrisc/lib/memcpy.c: use correct OR1200 option
Message-ID<tdo8H-6Sm-59@gated-at.bofh.it>
In reply to#1585645
From: Valentin Rothberg <valentinrothberg@gmail.com>

The Kconfig option for OR12000 is OR1K_1200.

Signed-off-by: Valentin Rothberg <valentinrothberg@gmail.com>
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/lib/memcpy.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/arch/openrisc/lib/memcpy.c b/arch/openrisc/lib/memcpy.c
index 4706f01..669887a 100644
--- a/arch/openrisc/lib/memcpy.c
+++ b/arch/openrisc/lib/memcpy.c
@@ -17,7 +17,7 @@
 
 #include <linux/string.h>
 
-#ifdef CONFIG_OR1200
+#ifdef CONFIG_OR1K_1200
 /*
  * Do memcpy with word copies and loop unrolling. This gives the
  * best performance on the OR1200 and MOR1KX archirectures
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585656 — [PATCH v3 19/25] scripts/checkstack.pl: Add openrisc support

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 19/25] scripts/checkstack.pl: Add openrisc support
Message-ID<tdo8I-6Sm-75@gated-at.bofh.it>
In reply to#1585645
Openrisc stack pointer is managed by decrementing r1. Add regexes to
recognize this.

Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 scripts/checkstack.pl | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/scripts/checkstack.pl b/scripts/checkstack.pl
index dd83978..eea5b78 100755
--- a/scripts/checkstack.pl
+++ b/scripts/checkstack.pl
@@ -106,6 +106,9 @@ my (@stack, $re, $dre, $x, $xs, $funcre);
 	} elsif ($arch eq 'sparc' || $arch eq 'sparc64') {
 		# f0019d10:       9d e3 bf 90     save  %sp, -112, %sp
 		$re = qr/.*save.*%sp, -(([0-9]{2}|[3-9])[0-9]{2}), %sp/o;
+	} elsif ($arch eq 'openrisc') {
+		# c000043c:       9c 21 fe f0     l.addi r1,r1,-272
+		$re = qr/.*l\.addi.*r1,r1,-(([0-9]{2}|[3-9])[0-9]{2})/o;
 	} else {
 		print("wrong or unknown architecture \"$arch\"\n");
 		exit
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585657 — [PATCH v3 24/25] openrisc: Export ioremap symbols used by modules

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 24/25] openrisc: Export ioremap symbols used by modules
Message-ID<tdo8H-6Sm-63@gated-at.bofh.it>
In reply to#1585645
Noticed this when building with allyesconfig.  Got build failures due
to iounmap and __ioremap symbols missing.  This patch exports them so
modules can use them.  This is inline with other architectures.

Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/mm/ioremap.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/arch/openrisc/mm/ioremap.c b/arch/openrisc/mm/ioremap.c
index 8705a46..2175e4b 100644
--- a/arch/openrisc/mm/ioremap.c
+++ b/arch/openrisc/mm/ioremap.c
@@ -80,6 +80,7 @@ __ioremap(phys_addr_t addr, unsigned long size, pgprot_t prot)
 
 	return (void __iomem *)(offset + (char *)v);
 }
+EXPORT_SYMBOL(__ioremap);
 
 void iounmap(void *addr)
 {
@@ -106,6 +107,7 @@ void iounmap(void *addr)
 
 	return vfree((void *)(PAGE_MASK & (unsigned long)addr));
 }
+EXPORT_SYMBOL(iounmap);
 
 /**
  * OK, this one's a bit tricky... ioremap can get called before memory is
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585658 — [PATCH v3 21/25] openrisc: entry: Fix delay slot detection

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 21/25] openrisc: entry: Fix delay slot detection
Message-ID<tdo8I-6Sm-69@gated-at.bofh.it>
In reply to#1585645
Use execption SR stored in pt_regs for detection, the current SR is not
correct as the handler is running after return from exception.

Also, The code that checks for a delay slot uses a flag bitmask and then
wants to check if the result is not zero.  The test it implemented was
wrong.

Correct it by changing the test to check result against non zero.

Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/kernel/entry.S | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/arch/openrisc/kernel/entry.S b/arch/openrisc/kernel/entry.S
index daae2a4..bc65008 100644
--- a/arch/openrisc/kernel/entry.S
+++ b/arch/openrisc/kernel/entry.S
@@ -258,9 +258,9 @@ EXCEPTION_ENTRY(_data_page_fault_handler)
 
 #else
 
-	l.mfspr r6,r0,SPR_SR               // SR
+	l.lwz   r6,PT_SR(r3)               // SR
 	l.andi  r6,r6,SPR_SR_DSX           // check for delay slot exception
-	l.sfeqi r6,0x1                     // exception happened in delay slot
+	l.sfne  r6,r0                      // exception happened in delay slot
 	l.bnf   7f
 	 l.lwz  r6,PT_PC(r3)               // address of an offending insn
 
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585662 — [PATCH v3 04/25] openrisc: head: use THREAD_SIZE instead of magic constant

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 04/25] openrisc: head: use THREAD_SIZE instead of magic constant
Message-ID<tdo8I-6Sm-79@gated-at.bofh.it>
In reply to#1585645
From: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>

The stack size was hard coded to 0x2000, use the standard THREAD_SIZE
definition loaded from thread_info.h.

Signed-off-by: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
[shorne@gmail.com: Added body to the commit message]
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/kernel/head.S | 4 +++-
 1 file changed, 3 insertions(+), 1 deletion(-)

diff --git a/arch/openrisc/kernel/head.S b/arch/openrisc/kernel/head.S
index 2346c5b..63ba2d9 100644
--- a/arch/openrisc/kernel/head.S
+++ b/arch/openrisc/kernel/head.S
@@ -24,6 +24,7 @@
 #include <asm/page.h>
 #include <asm/mmu.h>
 #include <asm/pgtable.h>
+#include <asm/thread_info.h>
 #include <asm/cache.h>
 #include <asm/spr_defs.h>
 #include <asm/asm-offsets.h>
@@ -486,7 +487,8 @@ _start:
 	/*
 	 * set up initial ksp and current
 	 */
-	LOAD_SYMBOL_2_GPR(r1,init_thread_union+0x2000)	// setup kernel stack
+	/* setup kernel stack */
+	LOAD_SYMBOL_2_GPR(r1,init_thread_union + THREAD_SIZE)
 	LOAD_SYMBOL_2_GPR(r10,init_thread_union)	// setup current
 	tophys	(r31,r10)
 	l.sw	TI_KSP(r31), r1
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585663 — [PATCH v3 02/25] openrisc: add cache way information to cpuinfo

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:20 +0100
Subject[PATCH v3 02/25] openrisc: add cache way information to cpuinfo
Message-ID<tdo8I-6Sm-83@gated-at.bofh.it>
In reply to#1585645
From: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>

Motivation for this is to be able to print the way information
properly in print_cpuinfo(), instead of hardcoding it to one.

Signed-off-by: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
Signed-off-by: Jonas Bonn <jonas@southpole.se>
[shorne@gmail.com fixed conflict with show_cpuinfo change]
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/include/asm/cpuinfo.h |  2 ++
 arch/openrisc/kernel/setup.c        | 67 ++++++++++++++++++++-----------------
 2 files changed, 38 insertions(+), 31 deletions(-)

diff --git a/arch/openrisc/include/asm/cpuinfo.h b/arch/openrisc/include/asm/cpuinfo.h
index 917318b..ec10679 100644
--- a/arch/openrisc/include/asm/cpuinfo.h
+++ b/arch/openrisc/include/asm/cpuinfo.h
@@ -24,9 +24,11 @@ struct cpuinfo {
 
 	u32 icache_size;
 	u32 icache_block_size;
+	u32 icache_ways;
 
 	u32 dcache_size;
 	u32 dcache_block_size;
+	u32 dcache_ways;
 };
 
 extern struct cpuinfo cpuinfo;
diff --git a/arch/openrisc/kernel/setup.c b/arch/openrisc/kernel/setup.c
index cb797a3..dbf5ee9 100644
--- a/arch/openrisc/kernel/setup.c
+++ b/arch/openrisc/kernel/setup.c
@@ -117,13 +117,15 @@ static void print_cpuinfo(void)
 	if (upr & SPR_UPR_DCP)
 		printk(KERN_INFO
 		       "-- dcache: %4d bytes total, %2d bytes/line, %d way(s)\n",
-		       cpuinfo.dcache_size, cpuinfo.dcache_block_size, 1);
+		       cpuinfo.dcache_size, cpuinfo.dcache_block_size,
+		       cpuinfo.dcache_ways);
 	else
 		printk(KERN_INFO "-- dcache disabled\n");
 	if (upr & SPR_UPR_ICP)
 		printk(KERN_INFO
 		       "-- icache: %4d bytes total, %2d bytes/line, %d way(s)\n",
-		       cpuinfo.icache_size, cpuinfo.icache_block_size, 1);
+		       cpuinfo.icache_size, cpuinfo.icache_block_size,
+		       cpuinfo.icache_ways);
 	else
 		printk(KERN_INFO "-- icache disabled\n");
 
@@ -155,25 +157,25 @@ void __init setup_cpuinfo(void)
 {
 	struct device_node *cpu;
 	unsigned long iccfgr, dccfgr;
-	unsigned long cache_set_size, cache_ways;
+	unsigned long cache_set_size;
 
 	cpu = of_find_compatible_node(NULL, NULL, "opencores,or1200-rtlsvn481");
 	if (!cpu)
 		panic("No compatible CPU found in device tree...\n");
 
 	iccfgr = mfspr(SPR_ICCFGR);
-	cache_ways = 1 << (iccfgr & SPR_ICCFGR_NCW);
+	cpuinfo.icache_ways = 1 << (iccfgr & SPR_ICCFGR_NCW);
 	cache_set_size = 1 << ((iccfgr & SPR_ICCFGR_NCS) >> 3);
 	cpuinfo.icache_block_size = 16 << ((iccfgr & SPR_ICCFGR_CBS) >> 7);
 	cpuinfo.icache_size =
-	    cache_set_size * cache_ways * cpuinfo.icache_block_size;
+	    cache_set_size * cpuinfo.icache_ways * cpuinfo.icache_block_size;
 
 	dccfgr = mfspr(SPR_DCCFGR);
-	cache_ways = 1 << (dccfgr & SPR_DCCFGR_NCW);
+	cpuinfo.dcache_ways = 1 << (dccfgr & SPR_DCCFGR_NCW);
 	cache_set_size = 1 << ((dccfgr & SPR_DCCFGR_NCS) >> 3);
 	cpuinfo.dcache_block_size = 16 << ((dccfgr & SPR_DCCFGR_CBS) >> 7);
 	cpuinfo.dcache_size =
-	    cache_set_size * cache_ways * cpuinfo.dcache_block_size;
+	    cache_set_size * cpuinfo.dcache_ways * cpuinfo.dcache_block_size;
 
 	if (of_property_read_u32(cpu, "clock-frequency",
 				 &cpuinfo.clock_frequency)) {
@@ -308,30 +310,33 @@ static int show_cpuinfo(struct seq_file *m, void *v)
 	revision = vr & SPR_VR_REV;
 
 	seq_printf(m,
-		   "cpu\t\t: OpenRISC-%x\n"
-		   "revision\t: %d\n"
-		   "frequency\t: %ld\n"
-		   "dcache size\t: %d bytes\n"
-		   "dcache block size\t: %d bytes\n"
-		   "icache size\t: %d bytes\n"
-		   "icache block size\t: %d bytes\n"
-		   "immu\t\t: %d entries, %lu ways\n"
-		   "dmmu\t\t: %d entries, %lu ways\n"
-		   "bogomips\t: %lu.%02lu\n",
-		   version,
-		   revision,
-		   loops_per_jiffy * HZ,
-		   cpuinfo.dcache_size,
-		   cpuinfo.dcache_block_size,
-		   cpuinfo.icache_size,
-		   cpuinfo.icache_block_size,
-		   1 << ((mfspr(SPR_DMMUCFGR) & SPR_DMMUCFGR_NTS) >> 2),
-		   1 + (mfspr(SPR_DMMUCFGR) & SPR_DMMUCFGR_NTW),
-		   1 << ((mfspr(SPR_IMMUCFGR) & SPR_IMMUCFGR_NTS) >> 2),
-		   1 + (mfspr(SPR_IMMUCFGR) & SPR_IMMUCFGR_NTW),
-		   (loops_per_jiffy * HZ) / 500000,
-		   ((loops_per_jiffy * HZ) / 5000) % 100);
-
+		  "cpu\t\t: OpenRISC-%x\n"
+		  "revision\t: %d\n"
+		  "frequency\t: %ld\n"
+		  "dcache size\t: %d bytes\n"
+		  "dcache block size\t: %d bytes\n"
+		  "dcache ways\t: %d\n"
+		  "icache size\t: %d bytes\n"
+		  "icache block size\t: %d bytes\n"
+		  "icache ways\t: %d\n"
+		  "immu\t\t: %d entries, %lu ways\n"
+		  "dmmu\t\t: %d entries, %lu ways\n"
+		  "bogomips\t: %lu.%02lu\n",
+		  version,
+		  revision,
+		  loops_per_jiffy * HZ,
+		  cpuinfo.dcache_size,
+		  cpuinfo.dcache_block_size,
+		  cpuinfo.dcache_ways,
+		  cpuinfo.icache_size,
+		  cpuinfo.icache_block_size,
+		  cpuinfo.icache_ways,
+		  1 << ((mfspr(SPR_DMMUCFGR) & SPR_DMMUCFGR_NTS) >> 2),
+		  1 + (mfspr(SPR_DMMUCFGR) & SPR_DMMUCFGR_NTW),
+		  1 << ((mfspr(SPR_IMMUCFGR) & SPR_IMMUCFGR_NTS) >> 2),
+		  1 + (mfspr(SPR_IMMUCFGR) & SPR_IMMUCFGR_NTW),
+		  (loops_per_jiffy * HZ) / 500000,
+		  ((loops_per_jiffy * HZ) / 5000) % 100);
 	return 0;
 }
 
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585665 — [PATCH v3 14/25] openrisc: Initial support for the idle state

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:30 +0100
Subject[PATCH v3 14/25] openrisc: Initial support for the idle state
Message-ID<tdoil-6Y4-1@gated-at.bofh.it>
In reply to#1585645
From: Sebastian Macke <sebastian@macke.de>

This patch adds basic support for the idle state of the cpu.
The patch overrides the regular idle function, enables the interupts,
checks for the power management unit and enables the cpu doze mode
if available.

Signed-off-by: Sebastian Macke <sebastian@macke.de>
[shorne@gmail.com: Fixed checkpatch, blankline after declarations]
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/kernel/process.c | 14 ++++++++++++++
 1 file changed, 14 insertions(+)

diff --git a/arch/openrisc/kernel/process.c b/arch/openrisc/kernel/process.c
index c49350b..ffde77f 100644
--- a/arch/openrisc/kernel/process.c
+++ b/arch/openrisc/kernel/process.c
@@ -75,6 +75,20 @@ void machine_power_off(void)
 	__asm__("l.nop 1");
 }
 
+/*
+ * Send the doze signal to the cpu if available.
+ * Make sure, that all interrupts are enabled
+ */
+void arch_cpu_idle(void)
+{
+	unsigned long upr;
+
+	local_irq_enable();
+	upr = mfspr(SPR_UPR);
+	if (upr & SPR_UPR_PMP)
+		mtspr(SPR_PMR, mfspr(SPR_PMR) | SPR_PMR_DME);
+}
+
 void (*pm_power_off) (void) = machine_power_off;
 
 /*
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585716 — Re: [PATCH v3 14/25] openrisc: Initial support for the idle state

FromJoe Perches <joe@perches.com>
Date2017-02-21 21:30 +0100
SubjectRe: [PATCH v3 14/25] openrisc: Initial support for the idle state
Message-ID<tdpeq-7y1-25@gated-at.bofh.it>
In reply to#1585665
On Wed, 2017-02-22 at 04:11 +0900, Stafford Horne wrote:
> From: Sebastian Macke <sebastian@macke.de>
> 
> This patch adds basic support for the idle state of the cpu.
> The patch overrides the regular idle function, enables the interupts,
> checks for the power management unit and enables the cpu doze mode
> if available.

trivia:

> diff --git a/arch/openrisc/kernel/process.c b/arch/openrisc/kernel/process.c
[]
> @@ -75,6 +75,20 @@ void machine_power_off(void)
>  	__asm__("l.nop 1");
>  }
>  
> +/*
> + * Send the doze signal to the cpu if available.
> + * Make sure, that all interrupts are enabled
> + */
> +void arch_cpu_idle(void)
> +{
> +	unsigned long upr;
> +
> +	local_irq_enable();
> +	upr = mfspr(SPR_UPR);
> +	if (upr & SPR_UPR_PMP)
> +		mtspr(SPR_PMR, mfspr(SPR_PMR) | SPR_PMR_DME);
> +}

Perhaps this would be easier to read without the automatic

void arch_cpu_idle(void)
{
	local_irq_enable();
	if (mfspr(SPR_UPR) & SPR_UPR_PMP)
		mtspr(SPR_PMR, mfspr(SPR_PMR) | SPR_PMR_DME);
}

[toc] | [prev] | [next] | [standalone]


#1586183 — Re: [PATCH v3 14/25] openrisc: Initial support for the idle state

FromStafford Horne <shorne@gmail.com>
Date2017-02-22 15:20 +0100
SubjectRe: [PATCH v3 14/25] openrisc: Initial support for the idle state
Message-ID<tdFVT-2TB-3@gated-at.bofh.it>
In reply to#1585716
On Tue, Feb 21, 2017 at 12:24:41PM -0800, Joe Perches wrote:
> On Wed, 2017-02-22 at 04:11 +0900, Stafford Horne wrote:
> > From: Sebastian Macke <sebastian@macke.de>
> > 
> > This patch adds basic support for the idle state of the cpu.
> > The patch overrides the regular idle function, enables the interupts,
> > checks for the power management unit and enables the cpu doze mode
> > if available.
> 
> trivia:
> 
> > diff --git a/arch/openrisc/kernel/process.c b/arch/openrisc/kernel/process.c
> []
> > @@ -75,6 +75,20 @@ void machine_power_off(void)
> >  	__asm__("l.nop 1");
> >  }
> >  
> > +/*
> > + * Send the doze signal to the cpu if available.
> > + * Make sure, that all interrupts are enabled
> > + */
> > +void arch_cpu_idle(void)
> > +{
> > +	unsigned long upr;
> > +
> > +	local_irq_enable();
> > +	upr = mfspr(SPR_UPR);
> > +	if (upr & SPR_UPR_PMP)
> > +		mtspr(SPR_PMR, mfspr(SPR_PMR) | SPR_PMR_DME);
> > +}
> 
> Perhaps this would be easier to read without the automatic
> 
> void arch_cpu_idle(void)
> {
> 	local_irq_enable();
> 	if (mfspr(SPR_UPR) & SPR_UPR_PMP)
> 		mtspr(SPR_PMR, mfspr(SPR_PMR) | SPR_PMR_DME);
> }

Yeah, that looks better.  I made the change.  Will post another series. 

[toc] | [prev] | [next] | [standalone]


#1585668 — [PATCH v3 18/25] MAINTAINERS: Add the openrisc official repository

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:30 +0100
Subject[PATCH v3 18/25] MAINTAINERS: Add the openrisc official repository
Message-ID<tdoil-6Y4-11@gated-at.bofh.it>
In reply to#1585645
The openrisc official repository and patch work happens currently on
github. Add the repo for reference.

Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 MAINTAINERS | 1 +
 1 file changed, 1 insertion(+)

diff --git a/MAINTAINERS b/MAINTAINERS
index 187b961..57809d6 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -9187,6 +9187,7 @@ OPENRISC ARCHITECTURE
 M:	Jonas Bonn <jonas@southpole.se>
 M:	Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
 M:	Stafford Horne <shorne@gmail.com>
+T:	git git://github.com/openrisc/linux.git
 L:	openrisc@lists.librecores.org
 W:	http://openrisc.io
 S:	Maintained
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585672 — [PATCH v3 07/25] openrisc: add atomic bitops

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:30 +0100
Subject[PATCH v3 07/25] openrisc: add atomic bitops
Message-ID<tdoim-6Y4-27@gated-at.bofh.it>
In reply to#1585645
From: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>

This utilize the load-link/store-conditional l.lwa and l.swa
instructions to implement the atomic bitops.
When those instructions are not available emulation is provided.

Cc: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
[shorne@gmail.com: remove OPENRISC_HAVE_INST_LWA_SWA config suggesed by
Alan Cox https://lkml.org/lkml/2014/7/23/666, implement
test_and_change_bit]
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/include/asm/bitops.h        |   2 +-
 arch/openrisc/include/asm/bitops/atomic.h | 123 ++++++++++++++++++++++++++++++
 2 files changed, 124 insertions(+), 1 deletion(-)
 create mode 100644 arch/openrisc/include/asm/bitops/atomic.h

diff --git a/arch/openrisc/include/asm/bitops.h b/arch/openrisc/include/asm/bitops.h
index 3003cda..689f568 100644
--- a/arch/openrisc/include/asm/bitops.h
+++ b/arch/openrisc/include/asm/bitops.h
@@ -45,7 +45,7 @@
 #include <asm-generic/bitops/hweight.h>
 #include <asm-generic/bitops/lock.h>
 
-#include <asm-generic/bitops/atomic.h>
+#include <asm/bitops/atomic.h>
 #include <asm-generic/bitops/non-atomic.h>
 #include <asm-generic/bitops/le.h>
 #include <asm-generic/bitops/ext2-atomic.h>
diff --git a/arch/openrisc/include/asm/bitops/atomic.h b/arch/openrisc/include/asm/bitops/atomic.h
new file mode 100644
index 0000000..35fb85f
--- /dev/null
+++ b/arch/openrisc/include/asm/bitops/atomic.h
@@ -0,0 +1,123 @@
+/*
+ * Copyright (C) 2014 Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
+ *
+ * This file is licensed under the terms of the GNU General Public License
+ * version 2.  This program is licensed "as is" without any warranty of any
+ * kind, whether express or implied.
+ */
+
+#ifndef __ASM_OPENRISC_BITOPS_ATOMIC_H
+#define __ASM_OPENRISC_BITOPS_ATOMIC_H
+
+static inline void set_bit(int nr, volatile unsigned long *addr)
+{
+	unsigned long mask = BIT_MASK(nr);
+	unsigned long *p = ((unsigned long *)addr) + BIT_WORD(nr);
+	unsigned long tmp;
+
+	__asm__ __volatile__(
+		"1:	l.lwa	%0,0(%1)	\n"
+		"	l.or	%0,%0,%2	\n"
+		"	l.swa	0(%1),%0	\n"
+		"	l.bnf	1b		\n"
+		"	 l.nop			\n"
+		: "=&r"(tmp)
+		: "r"(p), "r"(mask)
+		: "cc", "memory");
+}
+
+static inline void clear_bit(int nr, volatile unsigned long *addr)
+{
+	unsigned long mask = BIT_MASK(nr);
+	unsigned long *p = ((unsigned long *)addr) + BIT_WORD(nr);
+	unsigned long tmp;
+
+	__asm__ __volatile__(
+		"1:	l.lwa	%0,0(%1)	\n"
+		"	l.and	%0,%0,%2	\n"
+		"	l.swa	0(%1),%0	\n"
+		"	l.bnf	1b		\n"
+		"	 l.nop			\n"
+		: "=&r"(tmp)
+		: "r"(p), "r"(~mask)
+		: "cc", "memory");
+}
+
+static inline void change_bit(int nr, volatile unsigned long *addr)
+{
+	unsigned long mask = BIT_MASK(nr);
+	unsigned long *p = ((unsigned long *)addr) + BIT_WORD(nr);
+	unsigned long tmp;
+
+	__asm__ __volatile__(
+		"1:	l.lwa	%0,0(%1)	\n"
+		"	l.xor	%0,%0,%2	\n"
+		"	l.swa	0(%1),%0	\n"
+		"	l.bnf	1b		\n"
+		"	 l.nop			\n"
+		: "=&r"(tmp)
+		: "r"(p), "r"(mask)
+		: "cc", "memory");
+}
+
+static inline int test_and_set_bit(int nr, volatile unsigned long *addr)
+{
+	unsigned long mask = BIT_MASK(nr);
+	unsigned long *p = ((unsigned long *)addr) + BIT_WORD(nr);
+	unsigned long old;
+	unsigned long tmp;
+
+	__asm__ __volatile__(
+		"1:	l.lwa	%0,0(%2)	\n"
+		"	l.or	%1,%0,%3	\n"
+		"	l.swa	0(%2),%1	\n"
+		"	l.bnf	1b		\n"
+		"	 l.nop			\n"
+		: "=&r"(old), "=&r"(tmp)
+		: "r"(p), "r"(mask)
+		: "cc", "memory");
+
+	return (old & mask) != 0;
+}
+
+static inline int test_and_clear_bit(int nr, volatile unsigned long *addr)
+{
+	unsigned long mask = BIT_MASK(nr);
+	unsigned long *p = ((unsigned long *)addr) + BIT_WORD(nr);
+	unsigned long old;
+	unsigned long tmp;
+
+	__asm__ __volatile__(
+		"1:	l.lwa	%0,0(%2)	\n"
+		"	l.and	%1,%0,%3	\n"
+		"	l.swa	0(%2),%1	\n"
+		"	l.bnf	1b		\n"
+		"	 l.nop			\n"
+		: "=&r"(old), "=&r"(tmp)
+		: "r"(p), "r"(~mask)
+		: "cc", "memory");
+
+	return (old & mask) != 0;
+}
+
+static inline int test_and_change_bit(int nr, volatile unsigned long *addr)
+{
+	unsigned long mask = BIT_MASK(nr);
+	unsigned long *p = ((unsigned long *)addr) + BIT_WORD(nr);
+	unsigned long old;
+	unsigned long tmp;
+
+	__asm__ __volatile__(
+		"1:	l.lwa	%0,0(%2)	\n"
+		"	l.xor	%1,%0,%3	\n"
+		"	l.swa	0(%2),%1	\n"
+		"	l.bnf	1b		\n"
+		"	 l.nop			\n"
+		: "=&r"(old), "=&r"(tmp)
+		: "r"(p), "r"(mask)
+		: "cc", "memory");
+
+	return (old & mask) != 0;
+}
+
+#endif /* __ASM_OPENRISC_BITOPS_ATOMIC_H */
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1585674 — [PATCH v3 09/25] openrisc: add optimized atomic operations

FromStafford Horne <shorne@gmail.com>
Date2017-02-21 20:30 +0100
Subject[PATCH v3 09/25] openrisc: add optimized atomic operations
Message-ID<tdoim-6Y4-23@gated-at.bofh.it>
In reply to#1585645
From: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>

Using the l.lwa and l.swa atomic instruction pair.
Most openrisc processor cores provide these instructions now. If the
instructions are not available emulation is provided.

Cc: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
[shorne@gmail.com: remove OPENRISC_HAVE_INST_LWA_SWA config suggesed by
Alan Cox https://lkml.org/lkml/2014/7/23/666]
[shorne@gmail.com: expand to implement all ops suggested by Peter
Zijlstra https://lkml.org/lkml/2017/2/20/317]
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
 arch/openrisc/include/asm/Kbuild   |   1 -
 arch/openrisc/include/asm/atomic.h | 100 +++++++++++++++++++++++++++++++++++++
 2 files changed, 100 insertions(+), 1 deletion(-)
 create mode 100644 arch/openrisc/include/asm/atomic.h

diff --git a/arch/openrisc/include/asm/Kbuild b/arch/openrisc/include/asm/Kbuild
index 15e6ed5..1cedd63 100644
--- a/arch/openrisc/include/asm/Kbuild
+++ b/arch/openrisc/include/asm/Kbuild
@@ -1,7 +1,6 @@
 
 header-y += ucontext.h
 
-generic-y += atomic.h
 generic-y += auxvec.h
 generic-y += barrier.h
 generic-y += bitsperlong.h
diff --git a/arch/openrisc/include/asm/atomic.h b/arch/openrisc/include/asm/atomic.h
new file mode 100644
index 0000000..66f47ae
--- /dev/null
+++ b/arch/openrisc/include/asm/atomic.h
@@ -0,0 +1,100 @@
+/*
+ * Copyright (C) 2014 Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>
+ *
+ * This file is licensed under the terms of the GNU General Public License
+ * version 2.  This program is licensed "as is" without any warranty of any
+ * kind, whether express or implied.
+ */
+
+#ifndef __ASM_OPENRISC_ATOMIC_H
+#define __ASM_OPENRISC_ATOMIC_H
+
+#include <linux/types.h>
+
+/* Atomically perform op with v->counter and i */
+#define ATOMIC_OP(op)							\
+static inline void atomic_##op(int i, atomic_t *v)			\
+{									\
+	int tmp;							\
+									\
+	__asm__ __volatile__(						\
+		"1:	l.lwa	%0,0(%1)	\n"			\
+		"	l." #op " %0,%0,%2	\n"			\
+		"	l.swa	0(%1),%0	\n"			\
+		"	l.bnf	1b		\n"			\
+		"	 l.nop			\n"			\
+		: "=&r"(tmp)						\
+		: "r"(&v->counter), "r"(i)				\
+		: "cc", "memory");					\
+}
+
+/* Atomically perform op with v->counter and i, return the result */
+#define ATOMIC_OP_RETURN(op)						\
+static inline int atomic_##op##_return(int i, atomic_t *v)		\
+{									\
+	int tmp;							\
+									\
+	__asm__ __volatile__(						\
+		"1:	l.lwa	%0,0(%1)	\n"			\
+		"	l." #op " %0,%0,%2	\n"			\
+		"	l.swa	0(%1),%0	\n"			\
+		"	l.bnf	1b		\n"			\
+		"	 l.nop			\n"			\
+		: "=&r"(tmp)						\
+		: "r"(&v->counter), "r"(i)				\
+		: "cc", "memory");					\
+									\
+	return tmp;							\
+}
+
+/* Atomically perform op with v->counter and i, return orig v->counter */
+#define ATOMIC_FETCH_OP(op)						\
+static inline int atomic_fetch_##op(int i, atomic_t *v)			\
+{									\
+	int tmp, old;							\
+									\
+	__asm__ __volatile__(						\
+		"1:	l.lwa	%0,0(%2)	\n"			\
+		"	l." #op " %1,%0,%3	\n"			\
+		"	l.swa	0(%2),%1	\n"			\
+		"	l.bnf	1b		\n"			\
+		"	 l.nop			\n"			\
+		: "=&r"(old), "=&r"(tmp)				\
+		: "r"(&v->counter), "r"(i)				\
+		: "cc", "memory");					\
+									\
+	return old;							\
+}
+
+ATOMIC_OP_RETURN(add)
+ATOMIC_OP_RETURN(sub)
+
+ATOMIC_FETCH_OP(add)
+ATOMIC_FETCH_OP(sub)
+ATOMIC_FETCH_OP(and)
+ATOMIC_FETCH_OP(or)
+ATOMIC_FETCH_OP(xor)
+
+ATOMIC_OP(and)
+ATOMIC_OP(or)
+ATOMIC_OP(xor)
+
+#undef ATOMIC_FETCH_OP
+#undef ATOMIC_OP_RETURN
+#undef ATOMIC_OP
+
+#define atomic_add_return	atomic_add_return
+#define atomic_sub_return	atomic_sub_return
+#define atomic_fetch_add	atomic_fetch_add
+#define atomic_fetch_sub	atomic_fetch_sub
+#define atomic_fetch_and	atomic_fetch_and
+#define atomic_fetch_or		atomic_fetch_or
+#define atomic_fetch_xor	atomic_fetch_xor
+#define atomic_and	atomic_and
+#define atomic_or	atomic_or
+#define atomic_xor	atomic_xor
+
+
+#include <asm-generic/atomic.h>
+
+#endif /* __ASM_OPENRISC_ATOMIC_H */
-- 
2.9.3

[toc] | [prev] | [next] | [standalone]


#1586078 — Re: [PATCH v3 09/25] openrisc: add optimized atomic operations

FromPeter Zijlstra <peterz@infradead.org>
Date2017-02-22 12:30 +0100
SubjectRe: [PATCH v3 09/25] openrisc: add optimized atomic operations
Message-ID<tdDho-V3-1@gated-at.bofh.it>
In reply to#1585674
On Wed, Feb 22, 2017 at 04:11:38AM +0900, Stafford Horne wrote:
> +#define atomic_add_return	atomic_add_return
> +#define atomic_sub_return	atomic_sub_return
> +#define atomic_fetch_add	atomic_fetch_add
> +#define atomic_fetch_sub	atomic_fetch_sub
> +#define atomic_fetch_and	atomic_fetch_and
> +#define atomic_fetch_or		atomic_fetch_or
> +#define atomic_fetch_xor	atomic_fetch_xor
> +#define atomic_and	atomic_and
> +#define atomic_or	atomic_or
> +#define atomic_xor	atomic_xor
> +

It would be good to also implement __atomic_add_unless().

Something like so, if I got your asm right..

static inline int __atomic_add_unless(atomic_t *v, int a, int u)
{
	int old, tmp;

	__asm__ __volatile__(
		"1:     l.lwa %0, 0(%2)         \n"
		"       l.sfeq %0, %4           \n"
		"       l.bf 2f                 \n"
		"        l.nop                  \n"
		"	l.add %1, %0, %3	\n"
		"       l.swa 0(%2), %1         \n"
		"       l.bnf 1b                \n"
		"2:      l.nop                  \n"
		: "=&r"(old), "=&r" (tmp)
		: "r"(&v->counter), "r"(a), "r"(u)
		: "cc", "memory");

	return old;
}

[toc] | [prev] | [next] | [standalone]


#1586198 — Re: [PATCH v3 09/25] openrisc: add optimized atomic operations

FromStafford Horne <shorne@gmail.com>
Date2017-02-22 15:30 +0100
SubjectRe: [PATCH v3 09/25] openrisc: add optimized atomic operations
Message-ID<tdG5A-2Z4-21@gated-at.bofh.it>
In reply to#1586078
On Wed, Feb 22, 2017 at 12:27:37PM +0100, Peter Zijlstra wrote:
> On Wed, Feb 22, 2017 at 04:11:38AM +0900, Stafford Horne wrote:
> > +#define atomic_add_return	atomic_add_return
> > +#define atomic_sub_return	atomic_sub_return
> > +#define atomic_fetch_add	atomic_fetch_add
> > +#define atomic_fetch_sub	atomic_fetch_sub
> > +#define atomic_fetch_and	atomic_fetch_and
> > +#define atomic_fetch_or		atomic_fetch_or
> > +#define atomic_fetch_xor	atomic_fetch_xor
> > +#define atomic_and	atomic_and
> > +#define atomic_or	atomic_or
> > +#define atomic_xor	atomic_xor
> > +
> 
> It would be good to also implement __atomic_add_unless().
> 
> Something like so, if I got your asm right..
> 
> static inline int __atomic_add_unless(atomic_t *v, int a, int u)
> {
> 	int old, tmp;
> 
> 	__asm__ __volatile__(
> 		"1:     l.lwa %0, 0(%2)         \n"
> 		"       l.sfeq %0, %4           \n"
> 		"       l.bf 2f                 \n"
> 		"        l.nop                  \n"
> 		"	l.add %1, %0, %3	\n"
> 		"       l.swa 0(%2), %1         \n"
> 		"       l.bnf 1b                \n"
> 		"2:      l.nop                  \n"
> 		: "=&r"(old), "=&r" (tmp)
> 		: "r"(&v->counter), "r"(a), "r"(u)
> 		: "cc", "memory");
> 
> 	return old;
> }

Ok, thanks this looks right. I tested it too and it looks to work ok.

Note, I still include <asm-generic/atomic.h> to avoid copy-n-pastes.  So
I also wrapped __atomic_add_unless with #ifndef __atomic_add_unless in
the generic code.

-Stafford

[toc] | [prev] | [next] | [standalone]


#1586340 — Re: [OpenRISC] [PATCH v3 09/25] openrisc: add optimized atomic operations

FromRichard Henderson <rth@twiddle.net>
Date2017-02-22 18:40 +0100
SubjectRe: [OpenRISC] [PATCH v3 09/25] openrisc: add optimized atomic operations
Message-ID<tdJ3r-5dB-1@gated-at.bofh.it>
In reply to#1586198
On 02/23/2017 01:22 AM, Stafford Horne wrote:
>> static inline int __atomic_add_unless(atomic_t *v, int a, int u)
>> {
>> 	int old, tmp;
>>
>> 	__asm__ __volatile__(
>> 		"1:     l.lwa %0, 0(%2)         \n"
>> 		"       l.sfeq %0, %4           \n"
>> 		"       l.bf 2f                 \n"
>> 		"        l.nop                  \n"
>> 		"	l.add %1, %0, %3	\n"

You can move this add into the delay slot and drop the preceding nop.


r~

[toc] | [prev] | [next] | [standalone]


#1586532 — Re: [OpenRISC] [PATCH v3 09/25] openrisc: add optimized atomic operations

FromStafford Horne <shorne@gmail.com>
Date2017-02-22 23:50 +0100
SubjectRe: [OpenRISC] [PATCH v3 09/25] openrisc: add optimized atomic operations
Message-ID<tdNTs-eQ-17@gated-at.bofh.it>
In reply to#1586340
On Thu, Feb 23, 2017 at 04:31:34AM +1100, Richard Henderson wrote:
> On 02/23/2017 01:22 AM, Stafford Horne wrote:
> > > static inline int __atomic_add_unless(atomic_t *v, int a, int u)
> > > {
> > > 	int old, tmp;
> > > 
> > > 	__asm__ __volatile__(
> > > 		"1:     l.lwa %0, 0(%2)         \n"
> > > 		"       l.sfeq %0, %4           \n"
> > > 		"       l.bf 2f                 \n"
> > > 		"        l.nop                  \n"
> > > 		"	l.add %1, %0, %3	\n"
> 
> You can move this add into the delay slot and drop the preceding nop.

Thanks, Thats right, also here the 2: label being after the l.nop can be
applied.   I should have thought about it.  I made the change, Ill also
fix/look again in the other places.

-Stafford

> 
> r~

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web