Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1509816 > unrolled thread

Getting interrupt every million cache misses

Started byPavel Machek <pavel@ucw.cz>
First post2016-10-26 23:00 +0200
Last post2016-10-27 17:20 +0200
Articles 20 on this page of 44 — 8 participants

Back to article view | Back to linux.kernel


Contents

  Getting interrupt every million cache misses Pavel Machek <pavel@ucw.cz> - 2016-10-26 23:00 +0200
    Re: Getting interrupt every million cache misses Peter Zijlstra <peterz@infradead.org> - 2016-10-27 11:20 +0200
    Re: Getting interrupt every million cache misses Pavel Machek <pavel@ucw.cz> - 2016-10-27 16:40 +0200
      Re: Getting interrupt every million cache misses Peter Zijlstra <peterz@infradead.org> - 2016-10-27 16:40 +0200
        Re: Getting interrupt every million cache misses Kees Cook <keescook@chromium.org> - 2016-10-27 22:50 +0200
          rowhammer protection [was Re: Getting interrupt every million cache  misses] Pavel Machek <pavel@ucw.cz> - 2016-10-27 23:30 +0200
            Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Ingo Molnar <mingo@kernel.org> - 2016-10-28 09:10 +0200
              Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-28 11:00 +0200
                Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Ingo Molnar <mingo@kernel.org> - 2016-10-28 11:00 +0200
                  Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-28 14:00 +0200
                Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Peter Zijlstra <peterz@infradead.org> - 2016-10-28 11:10 +0200
                  Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Vegard Nossum <vegard.nossum@gmail.com> - 2016-10-28 11:30 +0200
                    Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Ingo Molnar <mingo@kernel.org> - 2016-10-28 11:40 +0200
                      Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Vegard Nossum <vegard.nossum@gmail.com> - 2016-10-28 11:50 +0200
                      Re: [kernel-hardening] Re: rowhammer protection [was Re: Getting  interrupt every million cache misses] Mark Rutland <mark.rutland@arm.com> - 2016-10-28 12:00 +0200
                  Re: rowhammer protection [was Re: Getting interrupt every million  cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-28 13:30 +0200
            Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Mark Rutland <mark.rutland@arm.com> - 2016-10-28 12:00 +0200
              Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-28 13:30 +0200
                Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Mark Rutland <mark.rutland@arm.com> - 2016-10-28 16:10 +0200
                  Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Peter Zijlstra <peterz@infradead.org> - 2016-10-28 16:20 +0200
                    Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-28 20:40 +0200
                      Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Peter Zijlstra <peterz@infradead.org> - 2016-10-28 20:50 +0200
                    Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-11-02 19:20 +0100
                  Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-28 19:30 +0200
                    Re: Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Gruss <daniel@gruss.cc> - 2016-10-29 15:20 +0200
                      Re: Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-29 21:50 +0200
                        Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Gruss <daniel@gruss.cc> - 2016-10-29 22:10 +0200
                          Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-29 23:10 +0200
                            Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Gruss <daniel@gruss.cc> - 2016-10-29 23:10 +0200
                              Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-29 23:50 +0200
                                Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Gruss <daniel@gruss.cc> - 2016-10-30 00:00 +0200
                                  Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-30 00:10 +0200
                                    Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Gruss <daniel@gruss.cc> - 2016-10-30 00:10 +0200
                  Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-31 09:30 +0100
                    Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Mark Rutland <mark.rutland@arm.com> - 2016-10-31 15:50 +0100
                      Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-10-31 22:20 +0100
                        Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Mark Rutland <mark.rutland@arm.com> - 2016-10-31 23:10 +0100
                    Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Ingo Molnar <mingo@kernel.org> - 2016-11-01 07:40 +0100
                      Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Micay <danielmicay@gmail.com> - 2016-11-01 08:30 +0100
                      Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Gruss <daniel@gruss.cc> - 2016-11-01 09:00 +0100
                      Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Daniel Gruss <daniel@gruss.cc> - 2016-11-01 09:20 +0100
                      Re: [kernel-hardening] rowhammer protection [was Re: Getting  interrupt every million cache misses] Pavel Machek <pavel@ucw.cz> - 2016-11-01 09:20 +0100
    Re: Getting interrupt every million cache misses Pavel Machek <pavel@ucw.cz> - 2016-10-27 16:50 +0200
    Re: Getting interrupt every million cache misses Peter Zijlstra <peterz@infradead.org> - 2016-10-27 17:20 +0200

Page 1 of 3  [1] 2 3  Next page →


#1509816 — Getting interrupt every million cache misses

FromPavel Machek <pavel@ucw.cz>
Date2016-10-26 23:00 +0200
SubjectGetting interrupt every million cache misses
Message-ID<swDsL-uO-41@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hi!

I'd like to get an interrupt every million cache misses... to do a
printk() or something like that. As far as I can tell, modern hardware
should allow me to do that. AFAICT performance events subsystem can do
something like that, but I can't figure out where the code is / what I
should call.

Can someone help?

Thanks,
									Pavel
-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [next] | [standalone]


#1510128

FromPeter Zijlstra <peterz@infradead.org>
Date2016-10-27 11:20 +0200
Message-ID<swP0R-8pq-1@gated-at.bofh.it>
In reply to#1509816
On Thu, Oct 27, 2016 at 10:46:38AM +0200, Pavel Machek wrote:

> And actually, printk() is not needed, udelay(50msec) is. Reason is,
> that DRAM becomes unreliable if about milion cache misses happen in
> under 64msec -- so I'd like to slow the system down in such cases to
> prevent bug from biting me.
> 
> (Details are here
> https://googleprojectzero.blogspot.cz/2015/03/exploiting-dram-rowhammer-bug-to-gain.html
> ). Bug is exploitable to get local root; it is also exploitable to
> gain local code execution from javascript... so it is rather severe.

Cute, a rowhammer defence.

So we can do in-kernel perf events too, see for example
kernel/watchdog.c:wd_hw_attr and its users.

I suppose you want PERF_COUNT_HW_CACHE_MISSES as config, although
depending on platform you could use better (u-arch specific) events.

[toc] | [prev] | [next] | [standalone]


#1510291

FromPavel Machek <pavel@ucw.cz>
Date2016-10-27 16:40 +0200
Message-ID<swU0z-38X-43@gated-at.bofh.it>
In reply to#1509816

[Multipart message — attachments visible in raw view] — view raw

On Thu 2016-10-27 10:28:01, Peter Zijlstra wrote:
> On Wed, Oct 26, 2016 at 10:54:16PM +0200, Pavel Machek wrote:
> > Hi!
> > 
> > I'd like to get an interrupt every million cache misses... to do a
> > printk() or something like that. As far as I can tell, modern hardware
> > should allow me to do that. AFAICT performance events subsystem can do
> > something like that, but I can't figure out where the code is / what I
> > should call.
> > 
> > Can someone help?
> 
> Can you go back one step and explain why you would want this? What use
> is a printk() on every 1e6-th cache miss.
> 
> That is, why doesn't:
> 
>  $ perf record -e cache-misses -c 1000000 -a -- sleep 5
> 
> suffice?

How to work around rowhammer, break my system _and_ make kernel perf
maintainers scream at the same time: (:-) )

I think I got the place now. Let me try...

Thanks,
								Pavel


diff --git a/arch/x86/events/core.c b/arch/x86/events/core.c
index d31735f..ce83f5e 100644
--- a/arch/x86/events/core.c
+++ b/arch/x86/events/core.c
@@ -1495,6 +1495,11 @@ perf_event_nmi_handler(unsigned int cmd, struct pt_regs *regs)
 
 	perf_sample_event_took(finish_clock - start_clock);
 
+	/* Here */
+	{
+		udelay(58000);
+	}
+
 	return ret;
 }
 NOKPROBE_SYMBOL(perf_event_nmi_handler);


-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1510298

FromPeter Zijlstra <peterz@infradead.org>
Date2016-10-27 16:40 +0200
Message-ID<swU0z-38X-55@gated-at.bofh.it>
In reply to#1510291
On Thu, Oct 27, 2016 at 11:11:04AM +0200, Pavel Machek wrote:
> How to work around rowhammer, break my system _and_ make kernel perf
> maintainers scream at the same time: (:-) )
> 
> I think I got the place now. Let me try...

Lol ;-)

> 
> diff --git a/arch/x86/events/core.c b/arch/x86/events/core.c
> index d31735f..ce83f5e 100644
> --- a/arch/x86/events/core.c
> +++ b/arch/x86/events/core.c
> @@ -1495,6 +1495,11 @@ perf_event_nmi_handler(unsigned int cmd, struct pt_regs *regs)
>  
>  	perf_sample_event_took(finish_clock - start_clock);
>  
> +	/* Here */
> +	{
> +		udelay(58000);
> +	}
> +
>  	return ret;
>  }
>  NOKPROBE_SYMBOL(perf_event_nmi_handler);

Like you guess, not quite ;-)


I think you want to register a custom overflow handler with your event.

So you get something like:


struct perf_event_attr rh_attr = {
	.type	= PERF_TYPE_HARDWARE,
	.config = PERF_COUNT_HW_CACHE_MISSES,
	.size	= sizeof(struct perf_event_attr),
	.pinned	= 1,
	.sample_period = 1000000,
};

static DEFINE_PER_CPU(struct perf_event *, rh_event);
static DEFINE_PER_CPU(u64, rh_timestamp);

static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
{
	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
	u64 now = ktime_get_mono_fast_ns();
	s64 delta = now - *ts;

	*ts = now;

	if (delta > 64 * NSEC_PER_USEC)
		udelay(58000);
}

__init int my_module_init()
{
	int cpu;

	/* XXX borken vs hotplug */

	for_each_online_cpu(cpu) {
		struct perf_event *event = per_cpu(event, cpu);

		event = perf_event_create_kernel_counter(&rh_attr, cpu, NULL, rh_overflow, NULL);
		if (!event)
			/* meh */
			;

	}
}

__exit void my_module_exit()
{
	int cpu;

	for_each_online_cpu(cpu) {
		struct perf_event *event = per_cpu(event, cpu);

		if (event)
			perf_event_release_kernel(event);
	}
}

[toc] | [prev] | [next] | [standalone]


#1510660

FromKees Cook <keescook@chromium.org>
Date2016-10-27 22:50 +0200
Message-ID<swZMC-6Nm-71@gated-at.bofh.it>
In reply to#1510298
On Thu, Oct 27, 2016 at 2:33 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> On Thu, Oct 27, 2016 at 11:11:04AM +0200, Pavel Machek wrote:
>> How to work around rowhammer, break my system _and_ make kernel perf
>> maintainers scream at the same time: (:-) )
>>
>> I think I got the place now. Let me try...
>
> Lol ;-)
>
>>
>> diff --git a/arch/x86/events/core.c b/arch/x86/events/core.c
>> index d31735f..ce83f5e 100644
>> --- a/arch/x86/events/core.c
>> +++ b/arch/x86/events/core.c
>> @@ -1495,6 +1495,11 @@ perf_event_nmi_handler(unsigned int cmd, struct pt_regs *regs)
>>
>>       perf_sample_event_took(finish_clock - start_clock);
>>
>> +     /* Here */
>> +     {
>> +             udelay(58000);
>> +     }
>> +
>>       return ret;
>>  }
>>  NOKPROBE_SYMBOL(perf_event_nmi_handler);
>
> Like you guess, not quite ;-)
>
>
> I think you want to register a custom overflow handler with your event.
>
> So you get something like:
>
>
> struct perf_event_attr rh_attr = {
>         .type   = PERF_TYPE_HARDWARE,
>         .config = PERF_COUNT_HW_CACHE_MISSES,
>         .size   = sizeof(struct perf_event_attr),
>         .pinned = 1,
>         .sample_period = 1000000,
> };
>
> static DEFINE_PER_CPU(struct perf_event *, rh_event);
> static DEFINE_PER_CPU(u64, rh_timestamp);
>
> static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
> {
>         u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
>         u64 now = ktime_get_mono_fast_ns();
>         s64 delta = now - *ts;
>
>         *ts = now;
>
>         if (delta > 64 * NSEC_PER_USEC)
>                 udelay(58000);
> }
>
> __init int my_module_init()
> {
>         int cpu;
>
>         /* XXX borken vs hotplug */
>
>         for_each_online_cpu(cpu) {
>                 struct perf_event *event = per_cpu(event, cpu);
>
>                 event = perf_event_create_kernel_counter(&rh_attr, cpu, NULL, rh_overflow, NULL);
>                 if (!event)
>                         /* meh */
>                         ;
>
>         }
> }
>
> __exit void my_module_exit()
> {
>         int cpu;
>
>         for_each_online_cpu(cpu) {
>                 struct perf_event *event = per_cpu(event, cpu);
>
>                 if (event)
>                         perf_event_release_kernel(event);
>         }
> }

This is pretty cool. Are there workloads other than rowhammer that
could trip this, and if so, how bad would this delay be for them?

At the very least, this could be behind a CONFIG for people that don't
have a way to fix their RAM refresh timings, etc.

-Kees

-- 
Kees Cook
Nexus Security

[toc] | [prev] | [next] | [standalone]


#1510683 — rowhammer protection [was Re: Getting interrupt every million cache misses]

FromPavel Machek <pavel@ucw.cz>
Date2016-10-27 23:30 +0200
Subjectrowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sx0pj-7fL-1@gated-at.bofh.it>
In reply to#1510660

[Multipart message — attachments visible in raw view] — view raw

Hi!

> >                 if (event)
> >                         perf_event_release_kernel(event);
> >         }
> > }
> 
> This is pretty cool. Are there workloads other than rowhammer that
> could trip this, and if so, how bad would this delay be for them?
> 
> At the very least, this could be behind a CONFIG for people that don't
> have a way to fix their RAM refresh timings, etc.

Yes, CONFIG_ is next.

Here's the patch, notice that I reversed the time handling logic -- it
should be correct now.

We can't tell cache misses on different addresses from cache misses on
same address (rowhammer), so this will have false positive. But so
far, my machine seems to work.

Unfortunately, I don't have machine suitable for testing nearby. Can
someone help with testing? [On the other hand... testing this is not
going to be easy. This will probably make problem way harder to
reproduce it in any case...]

I did run rowhammer, and yes, this did trigger and it was getting
delayed -- by factor of 2. That is slightly low -- delay should be
factor of 8 to get guarantees, if I understand things correctly.

Oh and NMI gets quite angry, but that was to be expected.

[  112.476009] perf: interrupt took too long (23660454 > 23654965),
lowering ker
nel.perf_event_max_sample_rate to 250
[  170.224007] INFO: NMI handler (perf_event_nmi_handler) took too
long to run: 55.844 msecs
[  191.872007] INFO: NMI handler (perf_event_nmi_handler) took too
long to run: 55.845 msecs

Best regards,
								Pavel

diff --git a/kernel/events/Makefile b/kernel/events/Makefile
index 2925188..130a185 100644
--- a/kernel/events/Makefile
+++ b/kernel/events/Makefile
@@ -2,7 +2,7 @@ ifdef CONFIG_FUNCTION_TRACER
 CFLAGS_REMOVE_core.o = $(CC_FLAGS_FTRACE)
 endif
 
-obj-y := core.o ring_buffer.o callchain.o
+obj-y := core.o ring_buffer.o callchain.o nohammer.o
 
 obj-$(CONFIG_HAVE_HW_BREAKPOINT) += hw_breakpoint.o
 obj-$(CONFIG_UPROBES) += uprobes.o
diff --git a/kernel/events/nohammer.c b/kernel/events/nohammer.c
new file mode 100644
index 0000000..01844d2
--- /dev/null
+++ b/kernel/events/nohammer.c
@@ -0,0 +1,66 @@
+/*
+ * Thanks to Peter Zijlstra <peterz@infradead.org>.
+ */
+
+#include <linux/perf_event.h>
+#include <linux/module.h>
+#include <linux/delay.h>
+
+struct perf_event_attr rh_attr = {
+	.type	= PERF_TYPE_HARDWARE,
+	.config = PERF_COUNT_HW_CACHE_MISSES,
+	.size	= sizeof(struct perf_event_attr),
+	.pinned	= 1,
+	/* FIXME: it is 1000000 per cpu. */
+	.sample_period = 500000,
+};
+
+static DEFINE_PER_CPU(struct perf_event *, rh_event);
+static DEFINE_PER_CPU(u64, rh_timestamp);
+
+static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
+{
+	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
+	u64 now = ktime_get_mono_fast_ns();
+	s64 delta = now - *ts;
+
+	*ts = now;
+
+	/* FIXME msec per usec, reverse logic? */
+	if (delta < 64 * NSEC_PER_MSEC)
+		mdelay(56);
+}
+
+static __init int my_module_init(void)
+{
+	int cpu;
+
+	/* XXX borken vs hotplug */
+
+	for_each_online_cpu(cpu) {
+		struct perf_event *event = per_cpu(rh_event, cpu);
+
+		event = perf_event_create_kernel_counter(&rh_attr, cpu, NULL, rh_overflow, NULL);
+		if (!event)
+			pr_err("Not enough resources to initialize nohammer on cpu %d\n", cpu);
+		pr_info("Nohammer initialized on cpu %d\n", cpu);
+		
+	}
+	return 0;
+}
+
+static __exit void my_module_exit(void)
+{
+	int cpu;
+
+	for_each_online_cpu(cpu) {
+		struct perf_event *event = per_cpu(rh_event, cpu);
+
+		if (event)
+			perf_event_release_kernel(event);
+	}
+	return;
+}
+
+module_init(my_module_init);
+module_exit(my_module_exit);


-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1510887 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromIngo Molnar <mingo@kernel.org>
Date2016-10-28 09:10 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sx9sB-56a-3@gated-at.bofh.it>
In reply to#1510683
* Pavel Machek <pavel@ucw.cz> wrote:

> +static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
> +{
> +	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
> +	u64 now = ktime_get_mono_fast_ns();
> +	s64 delta = now - *ts;
> +
> +	*ts = now;
> +
> +	/* FIXME msec per usec, reverse logic? */
> +	if (delta < 64 * NSEC_PER_MSEC)
> +		mdelay(56);
> +}

I'd suggest making the absolute delay sysctl tunable, because 'wait 56 msecs' is 
very magic, and do we know it 100% that 56 msecs is what is needed everywhere?

Plus I'd also suggest exposing an 'NMI rowhammer delay count' in /proc/interrupts, 
to make it easier to debug this. (Perhaps only show the line if the count is 
nonzero.)

Finally, could we please also add a sysctl and Kconfig that allows this feature to 
be turned on/off, with the default bootup value determined by the Kconfig value 
(i.e. by the distribution)? Similar to CONFIG_SECURITY_SELINUX_BOOTPARAM_VALUE.

Thanks,

	Ingo

[toc] | [prev] | [next] | [standalone]


#1510968 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromPavel Machek <pavel@ucw.cz>
Date2016-10-28 11:00 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxbb3-5ZM-11@gated-at.bofh.it>
In reply to#1510887

[Multipart message — attachments visible in raw view] — view raw

On Fri 2016-10-28 09:07:01, Ingo Molnar wrote:
> 
> * Pavel Machek <pavel@ucw.cz> wrote:
> 
> > +static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
> > +{
> > +	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
> > +	u64 now = ktime_get_mono_fast_ns();
> > +	s64 delta = now - *ts;
> > +
> > +	*ts = now;
> > +
> > +	/* FIXME msec per usec, reverse logic? */
> > +	if (delta < 64 * NSEC_PER_MSEC)
> > +		mdelay(56);
> > +}
> 
> I'd suggest making the absolute delay sysctl tunable, because 'wait 56 msecs' is 
> very magic, and do we know it 100% that 56 msecs is what is needed
> everywhere?

I agree this needs to be tunable (and with the other suggestions). But
this is actually not the most important tunable: the detection
threshold (rh_attr.sample_period) should be way more important.

And yes, this will all need to be tunable, somehow. But lets verify
that this works, first :-).

Thanks and best regards,
								Pavel

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1510970 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromIngo Molnar <mingo@kernel.org>
Date2016-10-28 11:00 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxbb3-5ZM-9@gated-at.bofh.it>
In reply to#1510968
* Pavel Machek <pavel@ucw.cz> wrote:

> On Fri 2016-10-28 09:07:01, Ingo Molnar wrote:
> > 
> > * Pavel Machek <pavel@ucw.cz> wrote:
> > 
> > > +static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
> > > +{
> > > +	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
> > > +	u64 now = ktime_get_mono_fast_ns();
> > > +	s64 delta = now - *ts;
> > > +
> > > +	*ts = now;
> > > +
> > > +	/* FIXME msec per usec, reverse logic? */
> > > +	if (delta < 64 * NSEC_PER_MSEC)
> > > +		mdelay(56);
> > > +}
> > 
> > I'd suggest making the absolute delay sysctl tunable, because 'wait 56 msecs' is 
> > very magic, and do we know it 100% that 56 msecs is what is needed
> > everywhere?
> 
> I agree this needs to be tunable (and with the other suggestions). But
> this is actually not the most important tunable: the detection
> threshold (rh_attr.sample_period) should be way more important.
> 
> And yes, this will all need to be tunable, somehow. But lets verify
> that this works, first :-).

Yeah.

Btw., a 56 NMI delay is pretty brutal in terms of latencies - it might
result in a smoother system to detect 100,000 cache misses and do a
~5.6 msecs delay instead?

(Assuming the shorter threshold does not trigger too often, of course.)

With all the tunables and statistics it would be possible to enumerate how 
frequently the protection mechanism kicks in during regular workloads.

Thanks,

	Ingo

[toc] | [prev] | [next] | [standalone]


#1511080 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromPavel Machek <pavel@ucw.cz>
Date2016-10-28 14:00 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxdZf-7Oz-7@gated-at.bofh.it>
In reply to#1510970

[Multipart message — attachments visible in raw view] — view raw

Hi!

> > I agree this needs to be tunable (and with the other suggestions). But
> > this is actually not the most important tunable: the detection
> > threshold (rh_attr.sample_period) should be way more important.
> > 
> > And yes, this will all need to be tunable, somehow. But lets verify
> > that this works, first :-).
> 
> Yeah.
> 
> Btw., a 56 NMI delay is pretty brutal in terms of latencies - it might
> result in a smoother system to detect 100,000 cache misses and do a
> ~5.6 msecs delay instead?
> 
> (Assuming the shorter threshold does not trigger too often, of
> course.)

Yeah, it is brutal workaround for a nasty bug. Slowdown depends on maximum utilization:

+/*
+ * Maximum permitted utilization of DRAM. Setting this to f will mean that
+ * when more than 1/f of maximum cache-miss performance is used, delay will
+ * be inserted, and will have similar effect on rowhammer as refreshing memory
+ * f times more often.
+ *
+ * Setting this to 8 should prevent the rowhammer attack.
+ */
+       int dram_max_utilization_factor = 8;

|                               | no prot. | fact. 1 | fact. 2 | fact. 8 |
| linux-n900$ time ./mkit       | 1m35     | 1m47    | 2m07    | 6m37    |
| rowhammer-test (for 43200000) | 2.86     | 9.75    | 16.7307 | 59.3738 |

(With factor 1 and 2 cpu attacker, we don't guarantee any protection.)

Best regards,
									Pavel
-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1510973 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromPeter Zijlstra <peterz@infradead.org>
Date2016-10-28 11:10 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxbkK-6ib-15@gated-at.bofh.it>
In reply to#1510968
On Fri, Oct 28, 2016 at 10:50:39AM +0200, Pavel Machek wrote:
> On Fri 2016-10-28 09:07:01, Ingo Molnar wrote:
> > 
> > * Pavel Machek <pavel@ucw.cz> wrote:
> > 
> > > +static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
> > > +{
> > > +	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
> > > +	u64 now = ktime_get_mono_fast_ns();
> > > +	s64 delta = now - *ts;
> > > +
> > > +	*ts = now;
> > > +
> > > +	/* FIXME msec per usec, reverse logic? */
> > > +	if (delta < 64 * NSEC_PER_MSEC)
> > > +		mdelay(56);
> > > +}
> > 
> > I'd suggest making the absolute delay sysctl tunable, because 'wait 56 msecs' is 
> > very magic, and do we know it 100% that 56 msecs is what is needed
> > everywhere?
> 
> I agree this needs to be tunable (and with the other suggestions). But
> this is actually not the most important tunable: the detection
> threshold (rh_attr.sample_period) should be way more important.

So being totally ignorant of the detail of how rowhammer abuses the DDR
thing, would it make sense to trigger more often and delay shorter? Or
is there some minimal delay required for things to settle or something.

[toc] | [prev] | [next] | [standalone]


#1510987 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromVegard Nossum <vegard.nossum@gmail.com>
Date2016-10-28 11:30 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxbE6-6p4-21@gated-at.bofh.it>
In reply to#1510973
On 28 October 2016 at 11:04, Peter Zijlstra <peterz@infradead.org> wrote:
> On Fri, Oct 28, 2016 at 10:50:39AM +0200, Pavel Machek wrote:
>> On Fri 2016-10-28 09:07:01, Ingo Molnar wrote:
>> >
>> > * Pavel Machek <pavel@ucw.cz> wrote:
>> >
>> > > +static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
>> > > +{
>> > > + u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
>> > > + u64 now = ktime_get_mono_fast_ns();
>> > > + s64 delta = now - *ts;
>> > > +
>> > > + *ts = now;
>> > > +
>> > > + /* FIXME msec per usec, reverse logic? */
>> > > + if (delta < 64 * NSEC_PER_MSEC)
>> > > +         mdelay(56);
>> > > +}
>> >
>> > I'd suggest making the absolute delay sysctl tunable, because 'wait 56 msecs' is
>> > very magic, and do we know it 100% that 56 msecs is what is needed
>> > everywhere?
>>
>> I agree this needs to be tunable (and with the other suggestions). But
>> this is actually not the most important tunable: the detection
>> threshold (rh_attr.sample_period) should be way more important.
>
> So being totally ignorant of the detail of how rowhammer abuses the DDR
> thing, would it make sense to trigger more often and delay shorter? Or
> is there some minimal delay required for things to settle or something.

Would it make sense to sample the counter on context switch, do some
accounting on a per-task cache miss counter, and slow down just the
single task(s) with a too high cache miss rate? That way there's no
global slowdown (which I assume would be the case here). The task's
slice of CPU would have to be taken into account because otherwise you
could have multiple cooperating tasks that each escape the limit but
taken together go above it.


Vegard

[toc] | [prev] | [next] | [standalone]


#1511000 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromIngo Molnar <mingo@kernel.org>
Date2016-10-28 11:40 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxbNL-6sc-29@gated-at.bofh.it>
In reply to#1510987
* Vegard Nossum <vegard.nossum@gmail.com> wrote:

> Would it make sense to sample the counter on context switch, do some
> accounting on a per-task cache miss counter, and slow down just the
> single task(s) with a too high cache miss rate? That way there's no
> global slowdown (which I assume would be the case here). The task's
> slice of CPU would have to be taken into account because otherwise you
> could have multiple cooperating tasks that each escape the limit but
> taken together go above it.

Attackers could work this around by splitting the rowhammer workload between 
multiple threads/processes.

I.e. the problem is that the risk may come from any 'unprivileged user-space 
code', where the rowhammer workload might be spread over multiple threads, 
processes or even users.

Thanks,

	Ingo

[toc] | [prev] | [next] | [standalone]


#1511003 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromVegard Nossum <vegard.nossum@gmail.com>
Date2016-10-28 11:50 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxbXr-6vs-17@gated-at.bofh.it>
In reply to#1511000
On 28 October 2016 at 11:35, Ingo Molnar <mingo@kernel.org> wrote:
>
> * Vegard Nossum <vegard.nossum@gmail.com> wrote:
>
>> Would it make sense to sample the counter on context switch, do some
>> accounting on a per-task cache miss counter, and slow down just the
>> single task(s) with a too high cache miss rate? That way there's no
>> global slowdown (which I assume would be the case here). The task's
>> slice of CPU would have to be taken into account because otherwise you
>> could have multiple cooperating tasks that each escape the limit but
>> taken together go above it.
>
> Attackers could work this around by splitting the rowhammer workload between
> multiple threads/processes.
>
> I.e. the problem is that the risk may come from any 'unprivileged user-space
> code', where the rowhammer workload might be spread over multiple threads,
> processes or even users.

That's why I emphasised the number of misses per CPU slice rather than
just the total number of misses. I assumed there must be at least one
task continuously hammering memory for a successful attack, in which
case it should be observable with as little as 1 slice of CPU (however
long that is), no?


Vegard

[toc] | [prev] | [next] | [standalone]


#1511016 — Re: [kernel-hardening] Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromMark Rutland <mark.rutland@arm.com>
Date2016-10-28 12:00 +0200
SubjectRe: [kernel-hardening] Re: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxc78-6yE-27@gated-at.bofh.it>
In reply to#1511000
On Fri, Oct 28, 2016 at 11:35:47AM +0200, Ingo Molnar wrote:
> 
> * Vegard Nossum <vegard.nossum@gmail.com> wrote:
> 
> > Would it make sense to sample the counter on context switch, do some
> > accounting on a per-task cache miss counter, and slow down just the
> > single task(s) with a too high cache miss rate? That way there's no
> > global slowdown (which I assume would be the case here). The task's
> > slice of CPU would have to be taken into account because otherwise you
> > could have multiple cooperating tasks that each escape the limit but
> > taken together go above it.
> 
> Attackers could work this around by splitting the rowhammer workload between 
> multiple threads/processes.

With the proposed approach, they could split across multiple CPUs
instead, no?

... or was that covered in a prior thread?

Thanks,
Mark.

[toc] | [prev] | [next] | [standalone]


#1511073 — Re: rowhammer protection [was Re: Getting interrupt every million cache misses]

FromPavel Machek <pavel@ucw.cz>
Date2016-10-28 13:30 +0200
SubjectRe: rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxdwe-7Es-11@gated-at.bofh.it>
In reply to#1510973

[Multipart message — attachments visible in raw view] — view raw

Hi!

> > I agree this needs to be tunable (and with the other suggestions). But
> > this is actually not the most important tunable: the detection
> > threshold (rh_attr.sample_period) should be way more important.
> 
> So being totally ignorant of the detail of how rowhammer abuses the DDR
> thing, would it make sense to trigger more often and delay shorter? Or
> is there some minimal delay required for things to settle or
 > something.

We can trigger more often and delay shorter, but it will mean that
protection will trigger with more false positives. I guess I'll play
with constants too see how big the effect is.

BTW...

[ 6267.180092] INFO: NMI handler (perf_event_nmi_handler) took too
long to run: 63.501 msecs

but I'm doing mdelay(64). .5 msec is not big difference, but...

Best regards,
									Pavel

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1511019 — Re: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]

FromMark Rutland <mark.rutland@arm.com>
Date2016-10-28 12:00 +0200
SubjectRe: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxc78-6yE-29@gated-at.bofh.it>
In reply to#1510683
Hi,

I missed the original, so I've lost some context.

Has this been tested on a system vulnerable to rowhammer, and if so, was
it reliable in mitigating the issue?

Which particular attack codebase was it tested against?

On Thu, Oct 27, 2016 at 11:27:47PM +0200, Pavel Machek wrote:
> --- /dev/null
> +++ b/kernel/events/nohammer.c
> @@ -0,0 +1,66 @@
> +/*
> + * Thanks to Peter Zijlstra <peterz@infradead.org>.
> + */
> +
> +#include <linux/perf_event.h>
> +#include <linux/module.h>
> +#include <linux/delay.h>
> +
> +struct perf_event_attr rh_attr = {
> +	.type	= PERF_TYPE_HARDWARE,
> +	.config = PERF_COUNT_HW_CACHE_MISSES,
> +	.size	= sizeof(struct perf_event_attr),
> +	.pinned	= 1,
> +	/* FIXME: it is 1000000 per cpu. */
> +	.sample_period = 500000,
> +};

I'm not sure that this is general enough to live in core code, because:

* there are existing ways around this (e.g. in the drammer case, using a
  non-cacheable mapping, which I don't believe would count as a cache
  miss).

  Given that, I'm very worried that this gives the false impression of
  protection in cases where a software workaround of this sort is
  insufficient or impossible.

* the precise semantics of performance counter events varies drastically
  across implementations. PERF_COUNT_HW_CACHE_MISSES, might only map to
  one particular level of cache, and/or may not be implemented on all
  cores.

* On some implementations, it may be that the counters are not
  interchangeable, and for those this would take away
  PERF_COUNT_HW_CACHE_MISSES from existing users.

> +static DEFINE_PER_CPU(struct perf_event *, rh_event);
> +static DEFINE_PER_CPU(u64, rh_timestamp);
> +
> +static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
> +{
> +	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
> +	u64 now = ktime_get_mono_fast_ns();
> +	s64 delta = now - *ts;
> +
> +	*ts = now;
> +
> +	/* FIXME msec per usec, reverse logic? */
> +	if (delta < 64 * NSEC_PER_MSEC)
> +		mdelay(56);
> +}

If I round-robin my attack across CPUs, how much does this help?

Thanks,
Mark.

[toc] | [prev] | [next] | [standalone]


#1511072 — Re: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]

FromPavel Machek <pavel@ucw.cz>
Date2016-10-28 13:30 +0200
SubjectRe: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxdwe-7Es-7@gated-at.bofh.it>
In reply to#1511019

[Multipart message — attachments visible in raw view] — view raw

Hi!

> I missed the original, so I've lost some context.

You can read it on lkml, but I guess you did not lose anything
important.

> Has this been tested on a system vulnerable to rowhammer, and if so, was
> it reliable in mitigating the issue?
> 
> Which particular attack codebase was it tested against?

I have rowhammer-test here,

commit 9824453fff76e0a3f5d1ac8200bc6c447c4fff57
Author: Mark Seaborn <mseaborn@chromium.org>

. I do not have vulnerable machine near me, so no "real" tests, but
I'm pretty sure it will make the error no longer reproducible with the
newer version. [Help welcome ;-)]

> > +struct perf_event_attr rh_attr = {
> > +	.type	= PERF_TYPE_HARDWARE,
> > +	.config = PERF_COUNT_HW_CACHE_MISSES,
> > +	.size	= sizeof(struct perf_event_attr),
> > +	.pinned	= 1,
> > +	/* FIXME: it is 1000000 per cpu. */
> > +	.sample_period = 500000,
> > +};
> 
> I'm not sure that this is general enough to live in core code, because:

Well, I'd like to postpone debate 'where does it live' to the later
stage. The problem is not arch-specific, the solution is not too
arch-specific either. I believe we can use Kconfig to hide it from
users where it does not apply. Anyway, lets decide if it works and
where, first.

> * the precise semantics of performance counter events varies drastically
>   across implementations. PERF_COUNT_HW_CACHE_MISSES, might only map to
>   one particular level of cache, and/or may not be implemented on all
>   cores.

If it maps to one particular cache level, we are fine (or maybe will
trigger protection too often). If some cores are not counted, that's
bad.

> * On some implementations, it may be that the counters are not
>   interchangeable, and for those this would take away
>   PERF_COUNT_HW_CACHE_MISSES from existing users.

Yup. Note that with this kind of protection, one missing performance
counter is likely to be small problem.

> > +	*ts = now;
> > +
> > +	/* FIXME msec per usec, reverse logic? */
> > +	if (delta < 64 * NSEC_PER_MSEC)
> > +		mdelay(56);
> > +}
> 
> If I round-robin my attack across CPUs, how much does this help?

See below for new explanation. With 2 CPUs, we are fine. On monster
big-little 8-core machines, we'd probably trigger protection too
often.

								Pavel

diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig
index e24e981..c6ffcaf 100644
--- a/arch/x86/Kconfig
+++ b/arch/x86/Kconfig
@@ -315,6 +315,7 @@ config PGTABLE_LEVELS
 
 source "init/Kconfig"
 source "kernel/Kconfig.freezer"
+source "kernel/events/Kconfig"
 
 menu "Processor type and features"
 
diff --git a/kernel/events/Kconfig b/kernel/events/Kconfig
new file mode 100644
index 0000000..7359427
--- /dev/null
+++ b/kernel/events/Kconfig
@@ -0,0 +1,9 @@
+config NOHAMMER
+        tristate "Rowhammer protection"
+        help
+	  Enable rowhammer attack prevention. Will degrade system
+	  performance under attack so much that attack should not
+	  be feasible.
+
+	  To compile this driver as a module, choose M here: the
+	  module will be called nohammer.
diff --git a/kernel/events/Makefile b/kernel/events/Makefile
index 2925188..03a2785 100644
--- a/kernel/events/Makefile
+++ b/kernel/events/Makefile
@@ -4,6 +4,8 @@ endif
 
 obj-y := core.o ring_buffer.o callchain.o
 
+obj-$(CONFIG_NOHAMMER) += nohammer.o
+
 obj-$(CONFIG_HAVE_HW_BREAKPOINT) += hw_breakpoint.o
 obj-$(CONFIG_UPROBES) += uprobes.o
 
diff --git a/kernel/events/nohammer.c b/kernel/events/nohammer.c
new file mode 100644
index 0000000..d96bacd
--- /dev/null
+++ b/kernel/events/nohammer.c
@@ -0,0 +1,140 @@
+/*
+ * Attempt to prevent rowhammer attack.
+ *
+ * On many new DRAM chips, repeated read access to nearby cells can cause
+ * victim cell to flip bits. Unfortunately, that can be used to gain root
+ * on affected machine, or to execute native code from javascript, escaping
+ * the sandbox.
+ *
+ * Fortunately, a lot of memory accesses is needed between DRAM refresh
+ * cycles. This is rather unusual workload, and we can detect it, and
+ * prevent the DRAM accesses, before bit flips happen.
+ *
+ * Thanks to Peter Zijlstra <peterz@infradead.org>.
+ * Thanks to presentation at blackhat.
+ */
+
+#include <linux/perf_event.h>
+#include <linux/module.h>
+#include <linux/delay.h>
+
+static struct perf_event_attr rh_attr = {
+	.type	= PERF_TYPE_HARDWARE,
+	.config = PERF_COUNT_HW_CACHE_MISSES,
+	.size	= sizeof(struct perf_event_attr),
+	.pinned	= 1,
+	.sample_period = 10000,
+};
+
+/*
+ * How often is the DRAM refreshed. Setting it too high is safe.
+ */
+static int dram_refresh_msec = 64;
+
+static DEFINE_PER_CPU(struct perf_event *, rh_event);
+static DEFINE_PER_CPU(u64, rh_timestamp);
+
+static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
+{
+	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
+	u64 now = ktime_get_mono_fast_ns();
+	s64 delta = now - *ts;
+
+	*ts = now;
+
+	if (delta < dram_refresh_msec * NSEC_PER_MSEC)
+		mdelay(dram_refresh_msec);
+}
+
+static __init int rh_module_init(void)
+{
+	int cpu;
+
+/*
+ * DRAM refresh is every 64 msec. That is not enough to prevent rowhammer.
+ * Some vendors doubled the refresh rate to 32 msec, that helps a lot, but
+ * does not close the attack completely. 8 msec refresh would probably do
+ * that on almost all chips.
+ *
+ * Thinkpad X60 can produce cca 12,200,000 cache misses a second, that's
+ * 780,800 cache misses per 64 msec window.
+ *
+ * X60 is from generation that is not yet vulnerable from rowhammer, and
+ * is pretty slow machine. That means that this limit is probably very
+ * safe on newer machines.
+ */
+	int cache_misses_per_second = 12200000;
+
+/*
+ * Maximum permitted utilization of DRAM. Setting this to f will mean that
+ * when more than 1/f of maximum cache-miss performance is used, delay will
+ * be inserted, and will have similar effect on rowhammer as refreshing memory
+ * f times more often.
+ *
+ * Setting this to 8 should prevent the rowhammer attack.
+ */
+	int dram_max_utilization_factor = 8;
+
+	/*
+	 * Hardware should be able to do approximately this many
+	 * misses per refresh
+	 */
+	int cache_miss_per_refresh = (cache_misses_per_second * dram_refresh_msec)/1000;
+
+	/*
+	 * So we do not want more than this many accesses to DRAM per
+	 * refresh.
+	 */
+	int cache_miss_limit = cache_miss_per_refresh / dram_max_utilization_factor;
+
+/*
+ * DRAM is shared between CPUs, but these performance counters are per-CPU.
+ */
+	int max_attacking_cpus = 2;
+
+	/*
+	 * We ignore counter overflows "too far away", but some of the
+	 * events might have actually occurent recently. Thus additional
+	 * factor of 2
+	 */
+
+	rh_attr.sample_period = cache_miss_limit / (2*max_attacking_cpus);
+
+	printk("Rowhammer protection limit is set to %d cache misses per %d msec\n",
+	       (int) rh_attr.sample_period, dram_refresh_msec);
+
+	/* XXX borken vs hotplug */
+
+	for_each_online_cpu(cpu) {
+		struct perf_event *event;
+
+		event = perf_event_create_kernel_counter(&rh_attr, cpu, NULL, rh_overflow, NULL);
+		per_cpu(rh_event, cpu) = event;		
+		if (!event) {
+			pr_err("Not enough resources to initialize nohammer on cpu %d\n", cpu);
+			continue;
+		}
+		pr_info("Nohammer initialized on cpu %d\n", cpu);
+	}
+	return 0;
+}
+
+static __exit void rh_module_exit(void)
+{
+	int cpu;
+
+	for_each_online_cpu(cpu) {
+		struct perf_event *event = per_cpu(rh_event, cpu);
+
+		if (event)
+			perf_event_release_kernel(event);
+	}
+	return;
+}
+
+module_init(rh_module_init);
+module_exit(rh_module_exit);
+
+MODULE_DESCRIPTION("Rowhammer protection");
+//MODULE_LICENSE("GPL v2+");
+MODULE_LICENSE("GPL");


-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1511151 — Re: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]

FromMark Rutland <mark.rutland@arm.com>
Date2016-10-28 16:10 +0200
SubjectRe: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxg13-Tz-11@gated-at.bofh.it>
In reply to#1511072
Hi,

On Fri, Oct 28, 2016 at 01:21:36PM +0200, Pavel Machek wrote:
> > Has this been tested on a system vulnerable to rowhammer, and if so, was
> > it reliable in mitigating the issue?
> > 
> > Which particular attack codebase was it tested against?
> 
> I have rowhammer-test here,
> 
> commit 9824453fff76e0a3f5d1ac8200bc6c447c4fff57
> Author: Mark Seaborn <mseaborn@chromium.org>

... from which repo?

> I do not have vulnerable machine near me, so no "real" tests, but
> I'm pretty sure it will make the error no longer reproducible with the
> newer version. [Help welcome ;-)]

Even if we hope this works, I think we have to be very careful with that
kind of assertion. Until we have data is to its efficacy, I don't think
we should claim that this is an effective mitigation.

> > > +struct perf_event_attr rh_attr = {
> > > +	.type	= PERF_TYPE_HARDWARE,
> > > +	.config = PERF_COUNT_HW_CACHE_MISSES,
> > > +	.size	= sizeof(struct perf_event_attr),
> > > +	.pinned	= 1,
> > > +	/* FIXME: it is 1000000 per cpu. */
> > > +	.sample_period = 500000,
> > > +};
> > 
> > I'm not sure that this is general enough to live in core code, because:
> 
> Well, I'd like to postpone debate 'where does it live' to the later
> stage. The problem is not arch-specific, the solution is not too
> arch-specific either. I believe we can use Kconfig to hide it from
> users where it does not apply. Anyway, lets decide if it works and
> where, first.

You seem to have forgotten the drammer case here, which this would not
have protected against. I'm not sure, but I suspect that we could have
similar issues with mappings using other attributes (e.g write-through),
as these would cause the memory traffic without cache miss events.

> > * the precise semantics of performance counter events varies drastically
> >   across implementations. PERF_COUNT_HW_CACHE_MISSES, might only map to
> >   one particular level of cache, and/or may not be implemented on all
> >   cores.
> 
> If it maps to one particular cache level, we are fine (or maybe will
> trigger protection too often). If some cores are not counted, that's bad.

Perhaps, but that depends on a number of implementation details. If "too
often" means "all the time", people will turn this off when they could
otherwise have been protected (e.g. if we can accurately monitor the
last level of cache).

> > * On some implementations, it may be that the counters are not
> >   interchangeable, and for those this would take away
> >   PERF_COUNT_HW_CACHE_MISSES from existing users.
> 
> Yup. Note that with this kind of protection, one missing performance
> counter is likely to be small problem.

That depends. Who chooses when to turn this on? If it's down to the
distro, this can adversely affect users with perfectly safe DRAM.

> > > +	/* FIXME msec per usec, reverse logic? */
> > > +	if (delta < 64 * NSEC_PER_MSEC)
> > > +		mdelay(56);
> > > +}
> > 
> > If I round-robin my attack across CPUs, how much does this help?
> 
> See below for new explanation. With 2 CPUs, we are fine. On monster
> big-little 8-core machines, we'd probably trigger protection too
> often.

We see larger core counts in mobile devices these days. In China,
octa-core phones are popular, for example. Servers go much larger.

> diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig
> index e24e981..c6ffcaf 100644
> --- a/arch/x86/Kconfig
> +++ b/arch/x86/Kconfig
> @@ -315,6 +315,7 @@ config PGTABLE_LEVELS
>  
>  source "init/Kconfig"
>  source "kernel/Kconfig.freezer"
> +source "kernel/events/Kconfig"
>  
>  menu "Processor type and features"
>  
> diff --git a/kernel/events/Kconfig b/kernel/events/Kconfig
> new file mode 100644
> index 0000000..7359427
> --- /dev/null
> +++ b/kernel/events/Kconfig
> @@ -0,0 +1,9 @@
> +config NOHAMMER
> +        tristate "Rowhammer protection"
> +        help
> +	  Enable rowhammer attack prevention. Will degrade system
> +	  performance under attack so much that attack should not
> +	  be feasible.


I think that this must make it clear that this is a best-effort approach
(i.e. it does not guarantee that an attack is not possible), and also
should make clear that said penalty may occur in other situations.

[...]

> +static struct perf_event_attr rh_attr = {
> +	.type	= PERF_TYPE_HARDWARE,
> +	.config = PERF_COUNT_HW_CACHE_MISSES,
> +	.size	= sizeof(struct perf_event_attr),
> +	.pinned	= 1,
> +	.sample_period = 10000,
> +};

What kind of overhead (just from taking the interrupt) will this come
with?

> +/*
> + * How often is the DRAM refreshed. Setting it too high is safe.
> + */

Stale comment? Given the check against delta below, this doesn't look to
be true.

> +static int dram_refresh_msec = 64;
> +
> +static DEFINE_PER_CPU(struct perf_event *, rh_event);
> +static DEFINE_PER_CPU(u64, rh_timestamp);
> +
> +static void rh_overflow(struct perf_event *event, struct perf_sample_data *data, struct pt_regs *regs)
> +{
> +	u64 *ts = this_cpu_ptr(&rh_timestamp); /* this is NMI context */
> +	u64 now = ktime_get_mono_fast_ns();
> +	s64 delta = now - *ts;
> +
> +	*ts = now;
> +
> +	if (delta < dram_refresh_msec * NSEC_PER_MSEC)
> +		mdelay(dram_refresh_msec);
> +}

[...]

> +/*
> + * DRAM is shared between CPUs, but these performance counters are per-CPU.
> + */
> +	int max_attacking_cpus = 2;

As above, many systems today have more than two CPUs. In the drammmer
paper, it looks like the majority had four.

Thanks
Mark.

[toc] | [prev] | [next] | [standalone]


#1511153 — Re: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]

FromPeter Zijlstra <peterz@infradead.org>
Date2016-10-28 16:20 +0200
SubjectRe: [kernel-hardening] rowhammer protection [was Re: Getting interrupt every million cache misses]
Message-ID<sxgaJ-ZQ-13@gated-at.bofh.it>
In reply to#1511151
On Fri, Oct 28, 2016 at 03:05:22PM +0100, Mark Rutland wrote:
> 
> > > * the precise semantics of performance counter events varies drastically
> > >   across implementations. PERF_COUNT_HW_CACHE_MISSES, might only map to
> > >   one particular level of cache, and/or may not be implemented on all
> > >   cores.
> > 
> > If it maps to one particular cache level, we are fine (or maybe will
> > trigger protection too often). If some cores are not counted, that's bad.
> 
> Perhaps, but that depends on a number of implementation details. If "too
> often" means "all the time", people will turn this off when they could
> otherwise have been protected (e.g. if we can accurately monitor the
> last level of cache).

Right, so one of the things mentioned in the paper is x86 NT stores.
Those are not cached and I'm not at all sure they're accounted in the
event we use for cache misses.

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | linux.kernel


csiph-web