Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1690604 > unrolled thread
| Started by | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| First post | 2017-07-18 22:00 +0200 |
| Last post | 2017-07-19 08:00 +0200 |
| Articles | 20 on this page of 43 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-18 22:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-18 23:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-19 08:00 +0200
RE: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Luck, Tony" <tony.luck@intel.com> - 2017-07-19 17:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-19 18:00 +0200
RE: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Luck, Tony" <tony.luck@intel.com> - 2017-07-19 20:10 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-19 18:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-20 06:40 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-20 22:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-20 22:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-20 23:10 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-21 15:40 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-21 15:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-21 17:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-21 17:40 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-21 17:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-21 18:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-21 19:10 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-21 19:30 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-21 20:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-22 08:30 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-24 17:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-24 17:10 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-24 17:30 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-24 17:40 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-24 18:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-24 18:40 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-24 19:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Boris Petkov <bp@alien8.de> - 2017-07-24 20:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-24 20:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-24 20:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-24 20:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-24 20:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-24 18:10 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-24 18:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-24 20:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-24 20:40 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-26 01:10 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-21 18:00 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-21 18:40 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2017-07-21 17:20 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Mauro Carvalho Chehab <mchehab@s-opensource.com> - 2017-07-21 15:50 +0200
Re: [PATCH 3/3] ghes_edac: add platform check to enable ghes_edac Borislav Petkov <bp@alien8.de> - 2017-07-19 08:00 +0200
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-22 08:30 +0200 |
| Message-ID | <u5W5k-6QR-1@gated-at.bofh.it> |
| In reply to | #1693923 |
On Fri, Jul 21, 2017 at 06:38:52PM +0000, Kani, Toshimitsu wrote:
> Enterprise platforms have very different model (I do not say it's
> better for everyone from the cost perspective). Typically, such
But you do tell your customers that the error counts they see are not
really what *actually* happens, right?
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-24 17:00 +0200 |
| Message-ID | <u6MZY-6gF-23@gated-at.bofh.it> |
| In reply to | #1694112 |
On Sat, 2017-07-22 at 08:28 +0200, Borislav Petkov wrote: > On Fri, Jul 21, 2017 at 06:38:52PM +0000, Kani, Toshimitsu wrote: > > Enterprise platforms have very different model (I do not say it's > > better for everyone from the cost perspective). Typically, such > > But you do tell your customers that the error counts they see are not > really what *actually* happens, right? We do not tell the error counts to customers. We tell customers when they need attention and have actionable items, and we provide support for that. Support gets all info necessary. There are multiple models for multiple types of customers. I am not saying one model is better than the other. Thanks, -Toshi
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-24 17:10 +0200 |
| Message-ID | <u6N9H-6A9-57@gated-at.bofh.it> |
| In reply to | #1694781 |
On Mon, Jul 24, 2017 at 02:49:30PM +0000, Kani, Toshimitsu wrote:
> We do not tell the error counts to customers.
Please read what I said: do you tell your customers that the error
counts they're seeing (or are *not* seeing) is bogus because the BIOS is
hiding them? Not the *actual* numbers!
> We tell customers when they need attention and have actionable items,
> and we provide support for that. Support gets all info necessary.
Ok, good to know. I'll make sure to bounce such issues to you guys in
the future.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-24 17:30 +0200 |
| Message-ID | <u6Nt2-6HD-47@gated-at.bofh.it> |
| In reply to | #1694799 |
On Mon, 2017-07-24 at 17:04 +0200, Borislav Petkov wrote: > On Mon, Jul 24, 2017 at 02:49:30PM +0000, Kani, Toshimitsu wrote: > > We do not tell the error counts to customers. > > Please read what I said: do you tell your customers that the error > counts they're seeing (or are *not* seeing) is bogus because the BIOS > is hiding them? Not the *actual* numbers! Customers do not see error counts. I do not think it's bogus. This model is basically the same as your car. You do not see error counts or periodical normal errors from all kinds of controllers in the car while you are driving. You get an attention lamp lit when you need to bring it to a car dealer. > > We tell customers when they need attention and have actionable > > items, and we provide support for that. Support gets all info > > necessary. > > Ok, good to know. I'll make sure to bounce such issues to you guys in > the future. We've been providing this model for many years now. I am just trying to enable OS error reporting with ghes_edac. Thanks, -Toshi
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-24 17:40 +0200 |
| Message-ID | <u6NCF-6Lv-7@gated-at.bofh.it> |
| In reply to | #1694818 |
On Mon, Jul 24, 2017 at 03:25:34PM +0000, Kani, Toshimitsu wrote:
> Customers do not see error counts. I do not think it's bogus.
Not showing the real error error counts but something contrived is the
definition of bogus numbers. But you're not showing anything - only when
some thresholds are being hit.
> This model is basically the same as your car. You do not see error
Oh jeez, we're talking about cars now.
> We've been providing this model for many years now.
Dude, relax, I'm only trying to point out to you that there are
customers who want to see *every* error and thus track how their
hardware behaves. And that for those customers it is probably worth
considering exposing that info and providing a switch to disable that
dumbing of the RAS functionality in the BIOS so that people can decide
for themselves. That's all.
I'm not questioning your model - I'm just saying that it could be
improved for certain customers. Do me a favor and this time *actually*
*read* my reply.
> I am just trying to enable OS error reporting with ghes_edac.
I know, you don't have to state the obvious constantly.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-24 18:00 +0200 |
| Message-ID | <u6NW2-6Tk-25@gated-at.bofh.it> |
| In reply to | #1694821 |
On Mon, 2017-07-24 at 17:37 +0200, Borislav Petkov wrote: > On Mon, Jul 24, 2017 at 03:25:34PM +0000, Kani, Toshimitsu wrote: : > > > We've been providing this model for many years now. > > Dude, relax, I'm only trying to point out to you that there are > customers who want to see *every* error and thus track how their > hardware behaves. And that for those customers it is probably worth > considering exposing that info and providing a switch to disable that > dumbing of the RAS functionality in the BIOS so that people can > decide for themselves. That's all. Yes, Mauro has already pointed this out. As I replied to him, we do have a separate series of platforms that do not have built-in RAS, and report all errors. Such customers can simply choose them. They do not need to pay for built-in RAS. The model w/ built-in RAS provides warranty & full support. As I said, it's a different model. Thanks, -Toshi
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-24 18:40 +0200 |
| Message-ID | <u6OyK-7ps-19@gated-at.bofh.it> |
| In reply to | #1694856 |
On Mon, Jul 24, 2017 at 03:56:27PM +0000, Kani, Toshimitsu wrote:
> Yes, Mauro has already pointed this out. As I replied to him, we do
> have a separate series of platforms that do not have built-in RAS, and
So this whitelist entry
+static struct acpi_oemlist oemlist[] = {
+ {"HPE ", "Server ", 0, ACPI_SIG_FADT, all_versions},
+ { } /* End */
+};
looks like it'll match every HP server platform not only the ones with
built-in RAS.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-24 19:50 +0200 |
| Message-ID | <u6PEv-87k-35@gated-at.bofh.it> |
| In reply to | #1694882 |
On Mon, 2017-07-24 at 18:37 +0200, Borislav Petkov wrote:
> On Mon, Jul 24, 2017 at 03:56:27PM +0000, Kani, Toshimitsu wrote:
> > Yes, Mauro has already pointed this out. As I replied to him, we
> > do have a separate series of platforms that do not have built-in
> > RAS, and
>
> So this whitelist entry
>
> +static struct acpi_oemlist oemlist[] = {
> + {"HPE ", "Server ", 0, ACPI_SIG_FADT, all_versions},
> + { } /* End */
> +};
>
> looks like it'll match every HP server platform not only the ones
> with built-in RAS.
I assumed our platforms w/o build-in RAS do not implement GHES, but I
will check for sure. Also, all our previous/current platforms have
"HP".
Thanks,
-Toshi
[toc] | [prev] | [next] | [standalone]
| From | Boris Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-24 20:00 +0200 |
| Message-ID | <u6POa-8bd-9@gated-at.bofh.it> |
| In reply to | #1694955 |
On July 24, 2017 8:44:03 PM GMT+03:00, "Kani, Toshimitsu" <toshi.kani@hpe.com> wrote: >I assumed our platforms w/o build-in RAS do not implement GHES, If we make it a normal module, it will be decoupled from GHES and it will rely only on the whitelist to load. -- Sent from a small device: formatting sux and brevity is inevitable.
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-24 20:00 +0200 |
| Message-ID | <u6POb-8bd-41@gated-at.bofh.it> |
| In reply to | #1694958 |
On Mon, 2017-07-24 at 20:50 +0300, Boris Petkov wrote: > On July 24, 2017 8:44:03 PM GMT+03:00, "Kani, Toshimitsu" <toshi.kani > @hpe.com> wrote: > > I assumed our platforms w/o build-in RAS do not implement GHES, > > If we make it a normal module, it will be decoupled from GHES and it > will rely only on the whitelist to load. Umm... I was under impression that we are adding the OSC bit check in addition to the current GHES filtering. Thanks, -Toshi
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-24 20:20 +0200 |
| Message-ID | <u6Q7w-6n-19@gated-at.bofh.it> |
| In reply to | #1694964 |
On Mon, Jul 24, 2017 at 05:54:52PM +0000, Kani, Toshimitsu wrote:
> Umm... I was under impression that we are adding the OSC bit check in
> addition to the current GHES filtering.
Read the parallel subthread again.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | Mauro Carvalho Chehab <mchehab@s-opensource.com> |
|---|---|
| Date | 2017-07-24 20:00 +0200 |
| Message-ID | <u6POb-8bd-31@gated-at.bofh.it> |
| In reply to | #1694856 |
Em Mon, 24 Jul 2017 15:56:27 +0000 "Kani, Toshimitsu" <toshi.kani@hpe.com> escreveu: > On Mon, 2017-07-24 at 17:37 +0200, Borislav Petkov wrote: > > On Mon, Jul 24, 2017 at 03:25:34PM +0000, Kani, Toshimitsu wrote: > : > > > > > We've been providing this model for many years now. > > > > Dude, relax, I'm only trying to point out to you that there are > > customers who want to see *every* error and thus track how their > > hardware behaves. And that for those customers it is probably worth > > considering exposing that info and providing a switch to disable that > > dumbing of the RAS functionality in the BIOS so that people can > > decide for themselves. That's all. > > Yes, Mauro has already pointed this out. As I replied to him, we do > have a separate series of platforms that do not have built-in RAS, and > report all errors. Such customers can simply choose them. They do not > need to pay for built-in RAS. That's probably too late for me as I received a new HP machine we bought just last week, but for the next time I would need to get a new hardware, what would be the non-RAS equivalent to a ML 350 G9 tower-mounted machine with two Xeon v4 CPUs and iLO? Regards, Mauro > > The model w/ built-in RAS provides warranty & full support. As I said, > it's a different model. > > Thanks, > -Toshi Thanks, Mauro
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-24 20:20 +0200 |
| Message-ID | <u6Q7w-6n-13@gated-at.bofh.it> |
| In reply to | #1694962 |
On Mon, 2017-07-24 at 14:56 -0300, Mauro Carvalho Chehab wrote: > Em Mon, 24 Jul 2017 15:56:27 +0000 : > That's probably too late for me as I received a new HP machine > we bought just last week, but for the next time I would need to > get a new hardware, what would be the non-RAS equivalent to > a ML 350 G9 tower-mounted machine with two Xeon v4 CPUs and iLO? Such servers are called "HPE Cloudline". But I think they are all rack-mounted, not tower-mounted machines. HP Inc. (which is now a separate company for consumer-oriented products) probably has such machine. Thanks, -Toshi
[toc] | [prev] | [next] | [standalone]
| From | Mauro Carvalho Chehab <mchehab@s-opensource.com> |
|---|---|
| Date | 2017-07-24 18:10 +0200 |
| Message-ID | <u6O5I-7cq-21@gated-at.bofh.it> |
| In reply to | #1694821 |
Em Mon, 24 Jul 2017 17:37:16 +0200 Borislav Petkov <bp@alien8.de> escreveu: > > Customers do not see error counts. I do not think it's bogus. > > I am just trying to enable OS error reporting with ghes_edac. > > I know, you don't have to state the obvious constantly. The problem I see is that, currently, on users that have EDAC already enabled, the users gets the errors directly from the hardware. If the Kernel force those users to use ghes_edac by default, they they won't see the error counts anymore, but, instead, hardware reports that the memories need to be replaced. Well, if such users are handling thresholds themselves, they won't see those errors anymore, as the errors will be masked. That's a regression. So, the right solution would be to keep hardware first, but providing a modprobe parameter to let them switch to software first. Thanks, Mauro
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-24 18:50 +0200 |
| Message-ID | <u6OIp-7tA-5@gated-at.bofh.it> |
| In reply to | #1694863 |
On Mon, Jul 24, 2017 at 01:04:13PM -0300, Mauro Carvalho Chehab wrote:
> If the Kernel force those users to use ghes_edac by default,
> they they won't see the error counts anymore, but, instead,
> hardware reports that the memories need to be replaced.
This is exactly why I'm trying to load ghes_edac only on those platforms
which would really want it.
> So, the right solution would be to keep hardware first, but
> providing a modprobe parameter to let them switch to software
> first.
That's exactly the issue: if we make it spec-conform and adhere to FF
setting, then it'll be clean. BUT(!), we will force ghes_edac on those
platforms which potentially are using the platform-specific drivers
until now. Not good.
If we do the whitelisting, then we're stuck with maintaining a yucky
whitelist and have to keep updating ghes_edac with it.
So we're basically between a rock and a hard place.
If I had to choose *right* *now*, I'd probably lean slightly towards the
whitelist as it won't break existing users.
A big grumpfy-grumbly hmmm. :-\
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | Mauro Carvalho Chehab <mchehab@s-opensource.com> |
|---|---|
| Date | 2017-07-24 20:20 +0200 |
| Message-ID | <u6Q7w-6n-23@gated-at.bofh.it> |
| In reply to | #1694887 |
Em Mon, 24 Jul 2017 18:44:00 +0200 Borislav Petkov <bp@alien8.de> escreveu: > On Mon, Jul 24, 2017 at 01:04:13PM -0300, Mauro Carvalho Chehab wrote: > > If the Kernel force those users to use ghes_edac by default, > > they they won't see the error counts anymore, but, instead, > > hardware reports that the memories need to be replaced. > > This is exactly why I'm trying to load ghes_edac only on those platforms > which would really want it. > > > So, the right solution would be to keep hardware first, but > > providing a modprobe parameter to let them switch to software > > first. > > That's exactly the issue: if we make it spec-conform and adhere to FF > setting, then it'll be clean. BUT(!), we will force ghes_edac on those > platforms which potentially are using the platform-specific drivers > until now. Not good. > > If we do the whitelisting, then we're stuck with maintaining a yucky > whitelist and have to keep updating ghes_edac with it. Yeah, having a whitelist is a maintainership's burden, but, on the other hand, I suspect that there aren't many systems that implement FF, have a reliable BIOS mapping of MB's silkscreen and doesn't filters out corrected errors using some sort of undocumented mechanism. So, I guess it is doable. Another alternative, with, IMO, is better would be to add a parameter like: edac=FF - firmware first; edac=hw - hardware first; edac=auto - honors FF if set in BIOS. Otherwise, hardware first. In order to avoid regressions, and to avoid the need of a whitelist, I would keep "edac=hw" as default. Thanks, Mauro
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-24 20:40 +0200 |
| Message-ID | <u6QqS-dR-15@gated-at.bofh.it> |
| In reply to | #1694976 |
(Sending to your other mail address because there's some temporary resolution
issue:
msmtp: recipient address mchehab@s-opensource.com not accepted by the server
msmtp: server message: 451 4.3.0 <mchehab@s-opensource.com>: Temporary lookup failure
msmtp: could not send mail (account alien8.de from /home/boris/.msmtprc)
Maybe the problem is on my end.)
On Mon, Jul 24, 2017 at 03:10:13PM -0300, Mauro Carvalho Chehab wrote:
> Yeah, having a whitelist is a maintainership's burden, but, on
> the other hand, I suspect that there aren't many systems that
> implement FF, have a reliable BIOS mapping of MB's silkscreen
> and doesn't filters out corrected errors using some sort of
> undocumented mechanism.
>
> So, I guess it is doable.
Right, let's hope.
> Another alternative, with, IMO, is better would be to add a parameter like:
>
> edac=FF - firmware first;
> edac=hw - hardware first;
> edac=auto - honors FF if set in BIOS. Otherwise, hardware first.
Or maybe edac=try_FF or so. But yeah, I guess we'll need something to
tell the EDAC core to try FF first.
> In order to avoid regressions, and to avoid the need of a whitelist,
> I would keep "edac=hw" as default.
So I don't want to break existing users and thus make only explicitly
known platforms load ghes_edac. In the current case, the HPE machines.
All the rest will simply use the platform drivers and nothing will
change for them.
Later we'll probably need to revisit this decision but right now and
with all things considered, the whitelist seems - as ugly as it is - the
most workable solution for all the different use cases and machines...
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-26 01:10 +0200 |
| Message-ID | <u7h7I-Pv-15@gated-at.bofh.it> |
| In reply to | #1694991 |
On Mon, 2017-07-24 at 20:30 +0200, Borislav Petkov wrote: : > > So I don't want to break existing users and thus make only explicitly > known platforms load ghes_edac. In the current case, the HPE > machines. All the rest will simply use the platform drivers and > nothing will change for them. > > Later we'll probably need to revisit this decision but right now and > with all things considered, the whitelist seems - as ugly as it is - > the most workable solution for all the different use cases and > machines... Agreed. I will verify OEMID info of our other platforms, and add APEI OSC check before calling ghes_edac_register(). Thanks, -Toshi
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-07-21 18:00 +0200 |
| Message-ID | <u5Ivo-6HH-3@gated-at.bofh.it> |
| In reply to | #1693801 |
On Fri, Jul 21, 2017 at 03:34:50PM +0000, Kani, Toshimitsu wrote:
> I suppose it'd depend on vendors, but I do not think users can do it
> properly unless they have depth knowledge about the hardware.
I'm talking about a menu in the BIOS where you can set the thresholding
levels on the system. Does your BIOS have that?
> Corrected errors are normal and expected to occur on healthy hardware.
> They do not need user's attention until they repeatedly occurred at a
> same place.
Apparently, you haven't been on enough maintanance calls, trying to calm
down the customer about the hardware error he sees in his logs...
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | "Kani, Toshimitsu" <toshi.kani@hpe.com> |
|---|---|
| Date | 2017-07-21 18:40 +0200 |
| Message-ID | <u5J85-7bl-1@gated-at.bofh.it> |
| In reply to | #1693811 |
On Fri, 2017-07-21 at 17:53 +0200, Borislav Petkov wrote: > On Fri, Jul 21, 2017 at 03:34:50PM +0000, Kani, Toshimitsu wrote: > > I suppose it'd depend on vendors, but I do not think users can do > > it properly unless they have depth knowledge about the hardware. > > I'm talking about a menu in the BIOS where you can set the > thresholding levels on the system. Does your BIOS have that? No, we don't offer such settings. > > Corrected errors are normal and expected to occur on healthy > > hardware. They do not need user's attention until they repeatedly > > occurred at a same place. > > Apparently, you haven't been on enough maintanance calls, trying to > calm down the customer about the hardware error he sees in his > logs... Actually, that's why. Reporting all corrected errors make users worried, call support, and asking to replace healthy hardware... Thanks, -Toshi
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web