Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1587277 > unrolled thread
| Started by | Stafford Horne <shorne@gmail.com> |
|---|---|
| First post | 2017-02-24 05:40 +0100 |
| Last post | 2017-02-24 05:50 +0100 |
| Articles | 5 on this page of 45 — 5 participants |
Back to article view | Back to linux.kernel
[PATCH v4 00/24] OpenRISC patches for 4.11 Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 23/24] openrisc: Export ioremap symbols used by modules Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 13/24] openrisc: Initial support for the idle state Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 18/24] scripts/checkstack.pl: Add openrisc support Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 18/24] scripts/checkstack.pl: Add openrisc support Tobias Klauser <tklauser@distanz.ch> - 2017-02-24 16:00 +0100
Re: [PATCH v4 18/24] scripts/checkstack.pl: Add openrisc support Stafford Horne <shorne@gmail.com> - 2017-02-24 20:40 +0100
[PATCH v4 22/24] arch/openrisc/lib/memcpy.c: use correct OR1200 option Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 22/24] arch/openrisc/lib/memcpy.c: use correct OR1200 option Jonas Bonn <jonas@southpole.se> - 2017-02-24 10:50 +0100
Re: [PATCH v4 22/24] arch/openrisc/lib/memcpy.c: use correct OR1200 option Stafford Horne <shorne@gmail.com> - 2017-02-24 21:00 +0100
[PATCH v4 08/24] openrisc: add cmpxchg and xchg implementations Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 20/24] openrisc: entry: Fix delay slot detection Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 20/24] openrisc: entry: Fix delay slot detection Jonas Bonn <jonas@southpole.se> - 2017-02-24 11:00 +0100
Re: [PATCH v4 20/24] openrisc: entry: Fix delay slot detection Stafford Horne <shorne@gmail.com> - 2017-02-24 21:20 +0100
[PATCH v4 09/24] openrisc: add optimized atomic operations Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 04/24] openrisc: head: use THREAD_SIZE instead of magic constant Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 07/24] openrisc: add atomic bitops Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 07/24] openrisc: add atomic bitops Peter Zijlstra <peterz@infradead.org> - 2017-02-24 12:00 +0100
Re: [PATCH v4 07/24] openrisc: add atomic bitops Stafford Horne <shorne@gmail.com> - 2017-02-24 20:30 +0100
[PATCH v4 11/24] openrisc: remove unnecessary stddef.h include Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 01/24] openrisc: use SPARSE_IRQ Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 21/24] openrisc: head: Move init strings to rodata section Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 21/24] openrisc: head: Move init strings to rodata section Jonas Bonn <jonas@southpole.se> - 2017-02-24 11:00 +0100
Re: [PATCH v4 21/24] openrisc: head: Move init strings to rodata section Stafford Horne <shorne@gmail.com> - 2017-02-24 21:20 +0100
[PATCH v4 05/24] openrisc: head: refactor out tlb flush into it's own function Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 05/24] openrisc: head: refactor out tlb flush into it's own function Jonas Bonn <jonas@southpole.se> - 2017-02-24 11:00 +0100
Re: [PATCH v4 05/24] openrisc: head: refactor out tlb flush into it's own function Stefan Kristiansson <stefan.kristiansson@saunalahti.fi> - 2017-02-24 12:10 +0100
Re: [PATCH v4 05/24] openrisc: head: refactor out tlb flush into it's own function Jonas Bonn <jonas@southpole.se> - 2017-02-24 13:50 +0100
Re: [PATCH v4 05/24] openrisc: head: refactor out tlb flush into it's own function Stefan Kristiansson <stefan.kristiansson@saunalahti.fi> - 2017-02-24 15:00 +0100
Re: [PATCH v4 05/24] openrisc: head: refactor out tlb flush into it's own function Stafford Horne <shorne@gmail.com> - 2017-02-24 20:30 +0100
[PATCH v4 24/24] openrisc: head: Init r0 to 0 on start Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 24/24] openrisc: head: Init r0 to 0 on start Jonas Bonn <jonas@southpole.se> - 2017-02-24 11:10 +0100
Re: [PATCH v4 24/24] openrisc: head: Init r0 to 0 on start Stafford Horne <shorne@gmail.com> - 2017-02-24 20:40 +0100
[PATCH v4 17/24] MAINTAINERS: Add the openrisc official repository Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 12/24] openrisc: Fix the bitmask for the unit present register Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 06/24] openrisc: add l.lwa/l.swa emulation Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 06/24] openrisc: add l.lwa/l.swa emulation Jonas Bonn <jonas@southpole.se> - 2017-02-24 10:50 +0100
Re: [PATCH v4 06/24] openrisc: add l.lwa/l.swa emulation Stafford Horne <shorne@gmail.com> - 2017-02-24 21:00 +0100
[PATCH v4 10/24] openrisc: add futex_atomic_* implementations Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 03/24] openrisc: tlb miss handler optimizations Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
[PATCH v4 19/24] openrisc: entry: Whitespace and comment cleanups Stafford Horne <shorne@gmail.com> - 2017-02-24 05:40 +0100
Re: [PATCH v4 19/24] openrisc: entry: Whitespace and comment cleanups Jonas Bonn <jonas@southpole.se> - 2017-02-24 10:50 +0100
Re: [PATCH v4 19/24] openrisc: entry: Whitespace and comment cleanups Stafford Horne <shorne@gmail.com> - 2017-02-24 21:40 +0100
[PATCH v4 15/24] openrisc: Add optimized memcpy routine Stafford Horne <shorne@gmail.com> - 2017-02-24 05:50 +0100
[PATCH v4 16/24] openrisc: Add .gitignore Stafford Horne <shorne@gmail.com> - 2017-02-24 05:50 +0100
[PATCH v4 14/24] openrisc: Add optimized memset Stafford Horne <shorne@gmail.com> - 2017-02-24 05:50 +0100
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Jonas Bonn <jonas@southpole.se> |
|---|---|
| Date | 2017-02-24 10:50 +0100 |
| Subject | Re: [PATCH v4 19/24] openrisc: entry: Whitespace and comment cleanups |
| Message-ID | <tekFI-6GV-5@gated-at.bofh.it> |
| In reply to | #1587298 |
On 02/24/2017 05:32 AM, Stafford Horne wrote: > Cleanups to whitespace and add some comments. Reading through the delay > slot logic I noticed some things: > - Delay slot instructions were not indented > - Some comments are not lined up > - Use tabs and spaces consistent with other code > > No functional change No, don't do this. Whitespace cleanups like this make life difficult for people rebasing on your tree, as well as blunting useful tools like git blame. I'm not against the indentation of the delay slot instructions; that seems sane and should be pretty transparent in a merge conflict. The whitespace cleanup and tab-space toggling needs to go, though. These sorts of things are better fixed up when the code lines they apply to are changed for other, functional reasons. I suggest you pull out the delay slot fixups into a separate patch and then just sit on the rest until you've discovered the pain of whitespace cleanup-inflicted merge conflicts and decide to just ditch them altogether. /Jonas > > Signed-off-by: Stafford Horne <shorne@gmail.com> > --- > arch/openrisc/kernel/entry.S | 38 ++++++++++++++++++-------------------- > 1 file changed, 18 insertions(+), 20 deletions(-) > > diff --git a/arch/openrisc/kernel/entry.S b/arch/openrisc/kernel/entry.S > index ba1a361..daae2a4 100644 > --- a/arch/openrisc/kernel/entry.S > +++ b/arch/openrisc/kernel/entry.S > @@ -228,7 +228,7 @@ EXCEPTION_ENTRY(_data_page_fault_handler) > * DTLB miss handler in the CONFIG_GUARD_PROTECTED_CORE part > */ > #ifdef CONFIG_OPENRISC_NO_SPR_SR_DSX > - l.lwz r6,PT_PC(r3) // address of an offending insn > + l.lwz r6,PT_PC(r3) // address of an offending insn > l.lwz r6,0(r6) // instruction that caused pf > > l.srli r6,r6,26 // check opcode for jump insn > @@ -244,49 +244,47 @@ EXCEPTION_ENTRY(_data_page_fault_handler) > l.bf 8f > l.sfeqi r6,0x12 // l.jalr > l.bf 8f > - > - l.nop > + l.nop > > l.j 9f > - l.nop > -8: > + l.nop > > - l.lwz r6,PT_PC(r3) // address of an offending insn > +8: // offending insn is in delay slot > + l.lwz r6,PT_PC(r3) // address of an offending insn > l.addi r6,r6,4 > l.lwz r6,0(r6) // instruction that caused pf > l.srli r6,r6,26 // get opcode > -9: > +9: // offending instruction opcode loaded in r6 > > #else > > - l.mfspr r6,r0,SPR_SR // SR > -// l.lwz r6,PT_SR(r3) // ESR > - l.andi r6,r6,SPR_SR_DSX // check for delay slot exception > - l.sfeqi r6,0x1 // exception happened in delay slot > - l.bnf 7f > - l.lwz r6,PT_PC(r3) // address of an offending insn > + l.mfspr r6,r0,SPR_SR // SR > + l.andi r6,r6,SPR_SR_DSX // check for delay slot exception > + l.sfeqi r6,0x1 // exception happened in delay slot > + l.bnf 7f > + l.lwz r6,PT_PC(r3) // address of an offending insn > > - l.addi r6,r6,4 // offending insn is in delay slot > + l.addi r6,r6,4 // offending insn is in delay slot > 7: > l.lwz r6,0(r6) // instruction that caused pf > l.srli r6,r6,26 // check opcode for write access > #endif > > - l.sfgeui r6,0x33 // check opcode for write access > + l.sfgeui r6,0x33 // check opcode for write access > l.bnf 1f > l.sfleui r6,0x37 > l.bnf 1f > l.ori r6,r0,0x1 // write access > l.j 2f > - l.nop > + l.nop > 1: l.ori r6,r0,0x0 // !write access > 2: > > /* call fault.c handler in or32/mm/fault.c */ > l.jal do_page_fault > - l.nop > + l.nop > l.j _ret_from_exception > - l.nop > + l.nop > > /* ---[ 0x400: Insn Page Fault exception ]------------------------------- */ > EXCEPTION_ENTRY(_itlb_miss_page_fault_handler) > @@ -306,9 +304,9 @@ EXCEPTION_ENTRY(_insn_page_fault_handler) > > /* call fault.c handler in or32/mm/fault.c */ > l.jal do_page_fault > - l.nop > + l.nop > l.j _ret_from_exception > - l.nop > + l.nop > > > /* ---[ 0x500: Timer exception ]----------------------------------------- */
[toc] | [prev] | [next] | [standalone]
| From | Stafford Horne <shorne@gmail.com> |
|---|---|
| Date | 2017-02-24 21:40 +0100 |
| Subject | Re: [PATCH v4 19/24] openrisc: entry: Whitespace and comment cleanups |
| Message-ID | <teuOJ-5to-11@gated-at.bofh.it> |
| In reply to | #1587470 |
On Fri, Feb 24, 2017 at 10:45:08AM +0100, Jonas Bonn wrote: > On 02/24/2017 05:32 AM, Stafford Horne wrote: > > Cleanups to whitespace and add some comments. Reading through the delay > > slot logic I noticed some things: > > - Delay slot instructions were not indented > > - Some comments are not lined up > > - Use tabs and spaces consistent with other code > > > > No functional change > > No, don't do this. Whitespace cleanups like this make life difficult for > people rebasing on your tree, as well as blunting useful tools like git > blame. > > I'm not against the indentation of the delay slot instructions; that seems > sane and should be pretty transparent in a merge conflict. > > The whitespace cleanup and tab-space toggling needs to go, though. These > sorts of things are better fixed up when the code lines they apply to are > changed for other, functional reasons. > > I suggest you pull out the delay slot fixups into a separate patch and then > just sit on the rest until you've discovered the pain of whitespace > cleanup-inflicted merge conflicts and decide to just ditch them altogether. Hi Jonas, I agree here and definitely if I knew others were actively working on this code I would be a bit more careful. But the truth is I think its only me right now. If you have patches that will get a conflict due to this let me know. Also, all of our out of tree patches for this file are actually upstream (as of this series now). Also, for me, this was needed in order for me to read the code (i.e. comments in the right place, and removing commented out lines) and fix the bug in the next patch. Nevertheless, if someone else thinks this should be left out Ill revert. -Stafford > > > > Signed-off-by: Stafford Horne <shorne@gmail.com> > > --- > > arch/openrisc/kernel/entry.S | 38 ++++++++++++++++++-------------------- > > 1 file changed, 18 insertions(+), 20 deletions(-) > > > > diff --git a/arch/openrisc/kernel/entry.S b/arch/openrisc/kernel/entry.S > > index ba1a361..daae2a4 100644 > > --- a/arch/openrisc/kernel/entry.S > > +++ b/arch/openrisc/kernel/entry.S > > @@ -228,7 +228,7 @@ EXCEPTION_ENTRY(_data_page_fault_handler) > > * DTLB miss handler in the CONFIG_GUARD_PROTECTED_CORE part > > */ > > #ifdef CONFIG_OPENRISC_NO_SPR_SR_DSX > > - l.lwz r6,PT_PC(r3) // address of an offending insn > > + l.lwz r6,PT_PC(r3) // address of an offending insn > > l.lwz r6,0(r6) // instruction that caused pf > > l.srli r6,r6,26 // check opcode for jump insn > > @@ -244,49 +244,47 @@ EXCEPTION_ENTRY(_data_page_fault_handler) > > l.bf 8f > > l.sfeqi r6,0x12 // l.jalr > > l.bf 8f > > - > > - l.nop > > + l.nop > > l.j 9f > > - l.nop > > -8: > > + l.nop > > - l.lwz r6,PT_PC(r3) // address of an offending insn > > +8: // offending insn is in delay slot > > + l.lwz r6,PT_PC(r3) // address of an offending insn > > l.addi r6,r6,4 > > l.lwz r6,0(r6) // instruction that caused pf > > l.srli r6,r6,26 // get opcode > > -9: > > +9: // offending instruction opcode loaded in r6 > > #else > > - l.mfspr r6,r0,SPR_SR // SR > > -// l.lwz r6,PT_SR(r3) // ESR > > - l.andi r6,r6,SPR_SR_DSX // check for delay slot exception > > - l.sfeqi r6,0x1 // exception happened in delay slot > > - l.bnf 7f > > - l.lwz r6,PT_PC(r3) // address of an offending insn > > + l.mfspr r6,r0,SPR_SR // SR > > + l.andi r6,r6,SPR_SR_DSX // check for delay slot exception > > + l.sfeqi r6,0x1 // exception happened in delay slot > > + l.bnf 7f > > + l.lwz r6,PT_PC(r3) // address of an offending insn > > - l.addi r6,r6,4 // offending insn is in delay slot > > + l.addi r6,r6,4 // offending insn is in delay slot > > 7: > > l.lwz r6,0(r6) // instruction that caused pf > > l.srli r6,r6,26 // check opcode for write access > > #endif > > - l.sfgeui r6,0x33 // check opcode for write access > > + l.sfgeui r6,0x33 // check opcode for write access > > l.bnf 1f > > l.sfleui r6,0x37 > > l.bnf 1f > > l.ori r6,r0,0x1 // write access > > l.j 2f > > - l.nop > > + l.nop > > 1: l.ori r6,r0,0x0 // !write access > > 2: > > /* call fault.c handler in or32/mm/fault.c */ > > l.jal do_page_fault > > - l.nop > > + l.nop > > l.j _ret_from_exception > > - l.nop > > + l.nop > > /* ---[ 0x400: Insn Page Fault exception ]------------------------------- */ > > EXCEPTION_ENTRY(_itlb_miss_page_fault_handler) > > @@ -306,9 +304,9 @@ EXCEPTION_ENTRY(_insn_page_fault_handler) > > /* call fault.c handler in or32/mm/fault.c */ > > l.jal do_page_fault > > - l.nop > > + l.nop > > l.j _ret_from_exception > > - l.nop > > + l.nop > > /* ---[ 0x500: Timer exception ]----------------------------------------- */ > >
[toc] | [prev] | [next] | [standalone]
| From | Stafford Horne <shorne@gmail.com> |
|---|---|
| Date | 2017-02-24 05:50 +0100 |
| Subject | [PATCH v4 15/24] openrisc: Add optimized memcpy routine |
| Message-ID | <tefZn-3l4-3@gated-at.bofh.it> |
| In reply to | #1587277 |
The generic memcpy routine provided in kernel does only byte copies.
Using word copies we can lower boot time and cycles spend in memcpy
quite significantly.
Booting on my de0 nano I see boot times go from 7.2 to 5.6 seconds.
The avg cycles in memcpy during boot go from 6467 to 1887.
I tested several algorithms (see code in previous patch mails)
The implementations I tested and avg cycles:
- Word Copies + Loop Unrolls + Non Aligned 1882
- Word Copies + Loop Unrolls 1887
- Word Copies 2441
- Byte Copies + Loop Unrolls 6467
- Byte Copies 7600
In the end I ended up going with Word Copies + Loop Unrolls as it
provides best tradeoff between simplicity and boot speedups.
Signed-off-by: Stafford Horne <shorne@gmail.com>
---
arch/openrisc/TODO.openrisc | 1 -
arch/openrisc/include/asm/string.h | 3 +
arch/openrisc/lib/Makefile | 2 +-
arch/openrisc/lib/memcpy.c | 124 +++++++++++++++++++++++++++++++++++++
4 files changed, 128 insertions(+), 2 deletions(-)
create mode 100644 arch/openrisc/lib/memcpy.c
diff --git a/arch/openrisc/TODO.openrisc b/arch/openrisc/TODO.openrisc
index 0eb04c8..c43d4e1 100644
--- a/arch/openrisc/TODO.openrisc
+++ b/arch/openrisc/TODO.openrisc
@@ -10,4 +10,3 @@ that are due for investigation shortly, i.e. our TODO list:
or1k and this change is slowly trickling through the stack. For the time
being, or32 is equivalent to or1k.
--- Implement optimized version of memcpy and memset
diff --git a/arch/openrisc/include/asm/string.h b/arch/openrisc/include/asm/string.h
index 33470d4..64939cc 100644
--- a/arch/openrisc/include/asm/string.h
+++ b/arch/openrisc/include/asm/string.h
@@ -4,4 +4,7 @@
#define __HAVE_ARCH_MEMSET
extern void *memset(void *s, int c, __kernel_size_t n);
+#define __HAVE_ARCH_MEMCPY
+extern void *memcpy(void *dest, __const void *src, __kernel_size_t n);
+
#endif /* __ASM_OPENRISC_STRING_H */
diff --git a/arch/openrisc/lib/Makefile b/arch/openrisc/lib/Makefile
index 67c583e..17d9d37 100644
--- a/arch/openrisc/lib/Makefile
+++ b/arch/openrisc/lib/Makefile
@@ -2,4 +2,4 @@
# Makefile for or32 specific library files..
#
-obj-y = memset.o string.o delay.o
+obj-y := delay.o string.o memset.o memcpy.o
diff --git a/arch/openrisc/lib/memcpy.c b/arch/openrisc/lib/memcpy.c
new file mode 100644
index 0000000..4706f01
--- /dev/null
+++ b/arch/openrisc/lib/memcpy.c
@@ -0,0 +1,124 @@
+/*
+ * arch/openrisc/lib/memcpy.c
+ *
+ * Optimized memory copy routines for openrisc. These are mostly copied
+ * from ohter sources but slightly entended based on ideas discuassed in
+ * #openrisc.
+ *
+ * The word unroll implementation is an extension to the arm byte
+ * unrolled implementation, but using word copies (if things are
+ * properly aligned)
+ *
+ * The great arm loop unroll algorithm can be found at:
+ * arch/arm/boot/compressed/string.c
+ */
+
+#include <linux/export.h>
+
+#include <linux/string.h>
+
+#ifdef CONFIG_OR1200
+/*
+ * Do memcpy with word copies and loop unrolling. This gives the
+ * best performance on the OR1200 and MOR1KX archirectures
+ */
+void *memcpy(void *dest, __const void *src, __kernel_size_t n)
+{
+ int i = 0;
+ unsigned char *d, *s;
+ uint32_t *dest_w = (uint32_t *)dest, *src_w = (uint32_t *)src;
+
+ /* If both source and dest are word aligned copy words */
+ if (!((unsigned int)dest_w & 3) && !((unsigned int)src_w & 3)) {
+ /* Copy 32 bytes per loop */
+ for (i = n >> 5; i > 0; i--) {
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ }
+
+ if (n & 1 << 4) {
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ }
+
+ if (n & 1 << 3) {
+ *dest_w++ = *src_w++;
+ *dest_w++ = *src_w++;
+ }
+
+ if (n & 1 << 2)
+ *dest_w++ = *src_w++;
+
+ d = (unsigned char *)dest_w;
+ s = (unsigned char *)src_w;
+
+ } else {
+ d = (unsigned char *)dest_w;
+ s = (unsigned char *)src_w;
+
+ for (i = n >> 3; i > 0; i--) {
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ }
+
+ if (n & 1 << 2) {
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ *d++ = *s++;
+ }
+ }
+
+ if (n & 1 << 1) {
+ *d++ = *s++;
+ *d++ = *s++;
+ }
+
+ if (n & 1)
+ *d++ = *s++;
+
+ return dest;
+}
+#else
+/*
+ * Use word copies but no loop unrolling as we cannot assume there
+ * will be benefits on the archirecture
+ */
+void *memcpy(void *dest, __const void *src, __kernel_size_t n)
+{
+ unsigned char *d = (unsigned char *)dest, *s = (unsigned char *)src;
+ uint32_t *dest_w = (uint32_t *)dest, *src_w = (uint32_t *)src;
+
+ /* If both source and dest are word aligned copy words */
+ if (!((unsigned int)dest_w & 3) && !((unsigned int)src_w & 3)) {
+ for (; n >= 4; n -= 4)
+ *dest_w++ = *src_w++;
+ }
+
+ d = (unsigned char *)dest_w;
+ s = (unsigned char *)src_w;
+
+ /* For remaining or if not aligned, copy bytes */
+ for (; n >= 1; n -= 1)
+ *d++ = *s++;
+
+ return dest;
+
+}
+#endif
+
+EXPORT_SYMBOL(memcpy);
--
2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Stafford Horne <shorne@gmail.com> |
|---|---|
| Date | 2017-02-24 05:50 +0100 |
| Subject | [PATCH v4 16/24] openrisc: Add .gitignore |
| Message-ID | <tefZn-3l4-5@gated-at.bofh.it> |
| In reply to | #1587277 |
This helps to suppress the vmlinux.lds file. Signed-off-by: Stafford Horne <shorne@gmail.com> --- arch/openrisc/kernel/.gitignore | 1 + 1 file changed, 1 insertion(+) create mode 100644 arch/openrisc/kernel/.gitignore diff --git a/arch/openrisc/kernel/.gitignore b/arch/openrisc/kernel/.gitignore new file mode 100644 index 0000000..c5f676c --- /dev/null +++ b/arch/openrisc/kernel/.gitignore @@ -0,0 +1 @@ +vmlinux.lds -- 2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Stafford Horne <shorne@gmail.com> |
|---|---|
| Date | 2017-02-24 05:50 +0100 |
| Subject | [PATCH v4 14/24] openrisc: Add optimized memset |
| Message-ID | <tefZn-3l4-9@gated-at.bofh.it> |
| In reply to | #1587277 |
From: Olof Kindgren <olof.kindgren@gmail.com> This adds a hand-optimized assembler version of memset and sets __HAVE_ARCH_MEMSET to use this version instead of the generic C routine Signed-off-by: Olof Kindgren <olof.kindgren@gmail.com> Signed-off-by: Stafford Horne <shorne@gmail.com> --- arch/openrisc/include/asm/string.h | 7 +++ arch/openrisc/kernel/or32_ksyms.c | 1 + arch/openrisc/lib/Makefile | 2 +- arch/openrisc/lib/memset.S | 98 ++++++++++++++++++++++++++++++++++++++ 4 files changed, 107 insertions(+), 1 deletion(-) create mode 100644 arch/openrisc/include/asm/string.h create mode 100644 arch/openrisc/lib/memset.S diff --git a/arch/openrisc/include/asm/string.h b/arch/openrisc/include/asm/string.h new file mode 100644 index 0000000..33470d4 --- /dev/null +++ b/arch/openrisc/include/asm/string.h @@ -0,0 +1,7 @@ +#ifndef __ASM_OPENRISC_STRING_H +#define __ASM_OPENRISC_STRING_H + +#define __HAVE_ARCH_MEMSET +extern void *memset(void *s, int c, __kernel_size_t n); + +#endif /* __ASM_OPENRISC_STRING_H */ diff --git a/arch/openrisc/kernel/or32_ksyms.c b/arch/openrisc/kernel/or32_ksyms.c index 86e31cf..5c4695d 100644 --- a/arch/openrisc/kernel/or32_ksyms.c +++ b/arch/openrisc/kernel/or32_ksyms.c @@ -44,3 +44,4 @@ DECLARE_EXPORT(__ashldi3); DECLARE_EXPORT(__lshrdi3); EXPORT_SYMBOL(__copy_tofrom_user); +EXPORT_SYMBOL(memset); diff --git a/arch/openrisc/lib/Makefile b/arch/openrisc/lib/Makefile index 966f65d..67c583e 100644 --- a/arch/openrisc/lib/Makefile +++ b/arch/openrisc/lib/Makefile @@ -2,4 +2,4 @@ # Makefile for or32 specific library files.. # -obj-y = string.o delay.o +obj-y = memset.o string.o delay.o diff --git a/arch/openrisc/lib/memset.S b/arch/openrisc/lib/memset.S new file mode 100644 index 0000000..92cc2ea --- /dev/null +++ b/arch/openrisc/lib/memset.S @@ -0,0 +1,98 @@ +/* + * OpenRISC memset.S + * + * Hand-optimized assembler version of memset for OpenRISC. + * Algorithm inspired by several other arch-specific memset routines + * in the kernel tree + * + * Copyright (C) 2015 Olof Kindgren <olof.kindgren@gmail.com> + * + * This program is free software; you can redistribute it and/or + * modify it under the terms of the GNU General Public License + * as published by the Free Software Foundation; either version + * 2 of the License, or (at your option) any later version. + */ + + .global memset + .type memset, @function +memset: + /* arguments: + * r3 = *s + * r4 = c + * r5 = n + * r13, r15, r17, r19 used as temp regs + */ + + /* Exit if n == 0 */ + l.sfeqi r5, 0 + l.bf 4f + + /* Truncate c to char */ + l.andi r13, r4, 0xff + + /* Skip word extension if c is 0 */ + l.sfeqi r13, 0 + l.bf 1f + /* Check for at least two whole words (8 bytes) */ + l.sfleui r5, 7 + + /* Extend char c to 32-bit word cccc in r13 */ + l.slli r15, r13, 16 // r13 = 000c, r15 = 0c00 + l.or r13, r13, r15 // r13 = 0c0c, r15 = 0c00 + l.slli r15, r13, 8 // r13 = 0c0c, r15 = c0c0 + l.or r13, r13, r15 // r13 = cccc, r15 = c0c0 + +1: l.addi r19, r3, 0 // Set r19 = src + /* Jump to byte copy loop if less than two words */ + l.bf 3f + l.or r17, r5, r0 // Set r17 = n + + /* Mask out two LSBs to check alignment */ + l.andi r15, r3, 0x3 + + /* lsb == 00, jump to word copy loop */ + l.sfeqi r15, 0 + l.bf 2f + l.addi r19, r3, 0 // Set r19 = src + + /* lsb == 01,10 or 11 */ + l.sb 0(r3), r13 // *src = c + l.addi r17, r17, -1 // Decrease n + + l.sfeqi r15, 3 + l.bf 2f + l.addi r19, r3, 1 // src += 1 + + /* lsb == 01 or 10 */ + l.sb 1(r3), r13 // *(src+1) = c + l.addi r17, r17, -1 // Decrease n + + l.sfeqi r15, 2 + l.bf 2f + l.addi r19, r3, 2 // src += 2 + + /* lsb == 01 */ + l.sb 2(r3), r13 // *(src+2) = c + l.addi r17, r17, -1 // Decrease n + l.addi r19, r3, 3 // src += 3 + + /* Word copy loop */ +2: l.sw 0(r19), r13 // *src = cccc + l.addi r17, r17, -4 // Decrease n + l.sfgeui r17, 4 + l.bf 2b + l.addi r19, r19, 4 // Increase src + + /* When n > 0, copy the remaining bytes, otherwise jump to exit */ + l.sfeqi r17, 0 + l.bf 4f + + /* Byte copy loop */ +3: l.addi r17, r17, -1 // Decrease n + l.sb 0(r19), r13 // *src = cccc + l.sfnei r17, 0 + l.bf 3b + l.addi r19, r19, 1 // Increase src + +4: l.jr r9 + l.ori r11, r3, 0 -- 2.9.3
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web