Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1423623 > unrolled thread

Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power reporting mechanism

Started byVince Weaver <vincent.weaver@maine.edu>
First post2016-06-16 03:20 +0200
Last post2016-06-16 23:40 +0200
Articles 7 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power  reporting mechanism Vince Weaver <vincent.weaver@maine.edu> - 2016-06-16 03:20 +0200
    Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power  reporting mechanism Borislav Petkov <bp@alien8.de> - 2016-06-16 18:50 +0200
    Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power  reporting mechanism Vince Weaver <vincent.weaver@maine.edu> - 2016-06-16 22:50 +0200
      Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power  reporting mechanism Vince Weaver <vincent.weaver@maine.edu> - 2016-06-17 18:00 +0200
    Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power  reporting mechanism Vince Weaver <vincent.weaver@maine.edu> - 2016-06-16 23:20 +0200
      Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power  reporting mechanism Vince Weaver <vincent.weaver@maine.edu> - 2016-06-16 23:20 +0200
        Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power  reporting mechanism Borislav Petkov <bp@alien8.de> - 2016-06-16 23:40 +0200

#1423623 — Re: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power reporting mechanism

FromVince Weaver <vincent.weaver@maine.edu>
Date2016-06-16 03:20 +0200
SubjectRe: [REDO PATCH v7] perf/x86/amd/power: Add AMD accumulated power reporting mechanism
Message-ID<rKu8q-1IJ-9@gated-at.bofh.it>
three questions about this functionality:

1.  In theory this should also work on an amd fam16h model 30h
    processor too, correct?  The current code limits things to fam15h
    even though the fam16mod30h has all the proper cpuid flags.

    I've tested the functionality a bit and it seems to work but for
    some reason the ptsc seems to occasionally count backwards
    on my machine.  Any reason that would be?  (It doesn't seem to be
    an overflow, just reading the ptsc 5ms apart and the values are 
    slightly lower after than before).

2.  Unless I'm misunderstanding things, the code seems to be accumulating 
	Power. (see chunk below) Power is an instantaneous measurement, it 
	makes no sense to add values.  If you use 5W for 1ms and 10W for
	1ms, the average power across the 2ms interval is not 15W.

	You can add energy, but not power.

> +	delta *= cpu_pwr_sample_ratio * 1000;
> +	tdelta = new_ptsc - prev_ptsc;
> +
> +	do_div(delta, tdelta);
> +	local64_add(delta, &event->count);

3.  The actual results gathered seem rediculously low.  341 seconds of
    calculation and only using 183 mWatts of power?

>    Performance counter stats for 'system wide':
> 
>               183.44 mWatts power/power-pkg/
> 
>        341.837270111 seconds time elapsed
> 
>   root@hr-zp:/home/ray/tip# ./tools/perf/perf stat -a -e 'power/power-pkg/' sleep 10

Vince

[toc] | [next] | [standalone]


#1424262

FromBorislav Petkov <bp@alien8.de>
Date2016-06-16 18:50 +0200
Message-ID<rKIEp-2jH-5@gated-at.bofh.it>
In reply to#1423623
On Thu, Jun 16, 2016 at 01:38:14PM +0800, Huang Rui wrote:
> I was told this feature would be supported on fam15h 60h, 70h and
> later processors before. Just checked the fam16h model 30h BKDG, yes,
> it should be also supported. But I didn't test that platform, if you
> confirm it works in your side. We can enable it.

You might want to ask around first whether F16M30's acc power machinery
is even usable? I.e., no errata and whatnot...

-- 
Regards/Gruss,
    Boris.

ECO tip #101: Trim your mails when you reply.

[toc] | [prev] | [next] | [standalone]


#1424387

FromVince Weaver <vincent.weaver@maine.edu>
Date2016-06-16 22:50 +0200
Message-ID<rKMoF-4xR-5@gated-at.bofh.it>
In reply to#1423623
On Thu, 16 Jun 2016, Huang Rui wrote:

> > 1.  In theory this should also work on an amd fam16h model 30h
> >     processor too, correct?  The current code limits things to fam15h
> >     even though the fam16mod30h has all the proper cpuid flags.
> > 
> 
> I was told this feature would be supported on fam15h 60h, 70h and
> later processors before. Just checked the fam16h model 30h BKDG, yes,
> it should be also supported. But I didn't test that platform, if you
> confirm it works in your side. We can enable it.

I can confirm I get power readings on my fam16hmod30h board once I apply a 
trivial patch to the driver.  I'll send the patch in a separate e-mail.

> PTSC's frequency is about 100Mhz, it shouldn't be overflow.

That's what I thought.  I'm trying to read the value using the /dev/msr 
interface from userspace and I get weird results.

i.e.:
	Jx: read 62d299b84
	PTSC MSR: read 72fe92
	
	sleep 5ms

	Jy: read 631b453b9
	PTSC MSR: read 46b25

this happens about half the time (PTSC going backwards).  Though 
admittedly the problem could somehow be in the MSR code I'm using.

> mWatts are for processor power not system power. Below data is
> calculated on fam15h model 60h which is low power platform. Even
> though the method has a minor mistake, the processor power should be
> in mWatts field.

I have an actual wall-mounted power meter hooked up to my system and the 
difference from idle to all-cores-busy is 20W, so I would think that that 
the results we find with perf should be >1W at least.

Vince

[toc] | [prev] | [next] | [standalone]


#1425268

FromVince Weaver <vincent.weaver@maine.edu>
Date2016-06-17 18:00 +0200
Message-ID<rL4lz-8sx-17@gated-at.bofh.it>
In reply to#1424387
On Fri, 17 Jun 2016, Huang Rui wrote:

> Can you try to read the MSR value two times with the same core
> (rdmsrl_on_cpu)?

I'm reading from userspace using the /dev/cpu/0/msr device so it should 
always be reading from cpu0.

I guess I could code up a custom kernel module to debug this if necessary.

It does look that for some reason the 0xc0010280 MSR is only returning the 
lower 24 bits of the PTSC, rather than the 40 bits that 
cpuid 80000008:ecx seems to think it should have.

Vince

[toc] | [prev] | [next] | [standalone]


#1424404

FromVince Weaver <vincent.weaver@maine.edu>
Date2016-06-16 23:20 +0200
Message-ID<rKMRH-4YX-7@gated-at.bofh.it>
In reply to#1423623
On Thu, 16 Jun 2016, Huang Rui wrote:

> On Thu, Jun 16, 2016 at 01:38:13PM +0800, Huang Rui wrote:
> 
> After considering carefully, the original method should be OK. 
> 
>       AMD nomenclature for CMT systems:
> 
>         [node 0] -> [Compute Unit 0] -> [Compute Unit Core 0] -> Linux CPU 0
>                                      -> [Compute Unit Core 1] -> Linux CPU 1
>                  -> [Compute Unit 1] -> [Compute Unit Core 0] -> Linux CPU 2
>                                      -> [Compute Unit Core 1] -> Linux CPU 3
> 
> The deltaN is power per compute unit. Current one package has two CUs.
> In the *same* interval, CU0's power is 10W, CU1's power is 15W. The
> package (CU0 + CU1) power should be 15W, right? Because the interval
> is the same.
> 
> Q = Q1 + Q2.  P = Q/t = (Q1 + Q2)/t = Q1/t + Q2/t = P1 + P2.
> 
> Is that clear?

OK, I was misunderstanding.  I somehow thought there was a periodic timer 
that was adding accumulating power over time.
But no, the driver just assumes the PTSC does not overflow?  And that 
addition is just there to handle adding all the cores together?

If so, then I agree that the addition makes sense, sorry for confusing 
things.

Although I think it would be better if we reported Joules (like 
RAPL does) rather than average power, but too late to change that now.


Also, on my machine I get results that make no physical sense, such as:

sudo perf stat -a -e power/power-pkg/  /usr/games/primes 1 500000000 > /dev/null

 Performance counter stats for 'system wide':

      4,472,401.06 mWatts power/power-pkg/                                            

       6.956135769 seconds time elapsed

I somehow don't think the CPU is really burning 4kW of Power.

Vince

[toc] | [prev] | [next] | [standalone]


#1424408

FromVince Weaver <vincent.weaver@maine.edu>
Date2016-06-16 23:20 +0200
Message-ID<rKMRI-4YX-21@gated-at.bofh.it>
In reply to#1424404
On Thu, 16 Jun 2016, Vince Weaver wrote:

One more followup, if I run the benchmark a bunch of times I get this:

      4,472,401.06 mWatts power/power-pkg/                                            
     50,886,303.28 mWatts power/power-pkg/                                            
     81,737,001.44 mWatts power/power-pkg/                                            
          6,525.89 mWatts power/power-pkg/                                            
          6,522.04 mWatts power/power-pkg/                                            
          6,505.68 mWatts power/power-pkg/                                            
      4,938,855.83 mWatts power/power-pkg/                                            
      4,614,620.11 mWatts power/power-pkg/                                            
     79,679,069.41 mWatts power/power-pkg/                                            
    152,794,060.83 mWatts power/power-pkg/                                            
      3,942,429.02 mWatts power/power-pkg/                                            
          6,506.73 mWatts power/power-pkg/                                            
     60,198,884.39 mWatts power/power-pkg/                                            

I'd believe the 6W report as a value for how much the CPU is using.  The 
others seem spurious.  I guess I should go check the Errata for this chip.

Vince

[toc] | [prev] | [next] | [standalone]


#1424438

FromBorislav Petkov <bp@alien8.de>
Date2016-06-16 23:40 +0200
Message-ID<rKNb3-56X-7@gated-at.bofh.it>
In reply to#1424408
On Thu, Jun 16, 2016 at 05:16:04PM -0400, Vince Weaver wrote:
> I'd believe the 6W report as a value for how much the CPU is using.
> The others seem spurious. I guess I should go check the Errata for
> this chip.

Maybe this is the reason why it got enabled on F15 only :-)

-- 
Regards/Gruss,
    Boris.

ECO tip #101: Trim your mails when you reply.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web