Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1325845 > unrolled thread
| Started by | Laura Abbott <labbott@redhat.com> |
|---|---|
| First post | 2016-02-03 19:50 +0100 |
| Last post | 2016-02-04 04:30 +0100 |
| Articles | 6 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [RFC][PATCH 0/3] Speed up SLUB poisoning + disable checks Laura Abbott <labbott@redhat.com> - 2016-02-03 19:50 +0100
Re: [RFC][PATCH 0/3] Speed up SLUB poisoning + disable checks Kees Cook <keescook@chromium.org> - 2016-02-03 22:10 +0100
Re: [RFC][PATCH 0/3] Speed up SLUB poisoning + disable checks Laura Abbott <labbott@redhat.com> - 2016-02-03 22:40 +0100
Re: [RFC][PATCH 0/3] Speed up SLUB poisoning + disable checks Christoph Lameter <cl@linux.com> - 2016-02-04 00:10 +0100
Re: [RFC][PATCH 0/3] Speed up SLUB poisoning + disable checks Laura Abbott <labbott@redhat.com> - 2016-02-04 01:50 +0100
Re: [RFC][PATCH 0/3] Speed up SLUB poisoning + disable checks Christoph Lameter <cl@linux.com> - 2016-02-04 04:30 +0100
| From | Laura Abbott <labbott@redhat.com> |
|---|---|
| Date | 2016-02-03 19:50 +0100 |
| Subject | Re: [RFC][PATCH 0/3] Speed up SLUB poisoning + disable checks |
| Message-ID | <qYaF6-5Ec-51@gated-at.bofh.it> |
On 01/25/2016 11:03 PM, Joonsoo Kim wrote: > On Mon, Jan 25, 2016 at 05:15:10PM -0800, Laura Abbott wrote: >> Hi, >> >> Based on the discussion from the series to add slab sanitization >> (lkml.kernel.org/g/<1450755641-7856-1-git-send-email-laura@labbott.name>) >> the existing SLAB_POISON mechanism already covers similar behavior. >> The performance of SLAB_POISON isn't very good. With hackbench -g 20 -l 1000 >> on QEMU with one cpu: > > I doesn't follow up that discussion, but, I think that reusing > SLAB_POISON for slab sanitization needs more changes. I assume that > completeness and performance is matter for slab sanitization. > > 1) SLAB_POISON isn't applied to specific kmem_cache which has > constructor or SLAB_DESTROY_BY_RCU flag. For debug, it's not necessary > to be applied, but, for slab sanitization, it is better to apply it to > all caches. The grsecurity patches get around this by calling the constructor again after poisoning. It could be worth investigating doing that as well although my focus was on the cases without the constructor. > > 2) SLAB_POISON makes object size bigger so natural alignment will be > broken. For example, kmalloc(256) cache's size is 256 in normal > case but it would be 264 when SLAB_POISON is enabled. This causes > memory waste. The grsecurity patches also bump the size up to put the free pointer outside the object. For sanitization purposes it is cleaner to have no pointers in the object after free > > In fact, I'd prefer not reusing SLAB_POISON. It would make thing > simpler. But, it's up to Christoph. > > Thanks. > It basically looks like trying to poison on the fast path at all will have a negative impact even with the feature is turned off. Christoph has indicated this is not acceptable so we are forced to limit it to the slow path only if we want runtime enablement. If we're limited to the slow path only, we might as well work with SLAB_POISON to make it faster. We can reevaluate if it turns out the poisoning isn't fast enough to be useful. Thanks, Laura
[toc] | [next] | [standalone]
| From | Kees Cook <keescook@chromium.org> |
|---|---|
| Date | 2016-02-03 22:10 +0100 |
| Message-ID | <qYcQz-7bF-19@gated-at.bofh.it> |
| In reply to | #1325845 |
On Wed, Feb 3, 2016 at 10:46 AM, Laura Abbott <labbott@redhat.com> wrote: > On 01/25/2016 11:03 PM, Joonsoo Kim wrote: >> >> On Mon, Jan 25, 2016 at 05:15:10PM -0800, Laura Abbott wrote: >>> >>> Hi, >>> >>> Based on the discussion from the series to add slab sanitization >>> (lkml.kernel.org/g/<1450755641-7856-1-git-send-email-laura@labbott.name>) >>> the existing SLAB_POISON mechanism already covers similar behavior. >>> The performance of SLAB_POISON isn't very good. With hackbench -g 20 -l >>> 1000 >>> on QEMU with one cpu: >> >> >> I doesn't follow up that discussion, but, I think that reusing >> SLAB_POISON for slab sanitization needs more changes. I assume that >> completeness and performance is matter for slab sanitization. >> >> 1) SLAB_POISON isn't applied to specific kmem_cache which has >> constructor or SLAB_DESTROY_BY_RCU flag. For debug, it's not necessary >> to be applied, but, for slab sanitization, it is better to apply it to >> all caches. > > > The grsecurity patches get around this by calling the constructor again > after poisoning. It could be worth investigating doing that as well > although my focus was on the cases without the constructor. >> >> >> 2) SLAB_POISON makes object size bigger so natural alignment will be >> broken. For example, kmalloc(256) cache's size is 256 in normal >> case but it would be 264 when SLAB_POISON is enabled. This causes >> memory waste. > > > The grsecurity patches also bump the size up to put the free pointer > outside the object. For sanitization purposes it is cleaner to have > no pointers in the object after free > >> >> In fact, I'd prefer not reusing SLAB_POISON. It would make thing >> simpler. But, it's up to Christoph. >> >> Thanks. >> > > It basically looks like trying to poison on the fast path at all > will have a negative impact even with the feature is turned off. > Christoph has indicated this is not acceptable so we are forced > to limit it to the slow path only if we want runtime enablement. Is it possible to have both? i.e fast path via CONFIG, and slow path via runtime options? > If we're limited to the slow path only, we might as well work > with SLAB_POISON to make it faster. We can reevaluate if it turns > out the poisoning isn't fast enough to be useful. And since I'm new to this area, I know of fast/slow path in the syscall sense. What happens in the allocation/free fast/slow path that makes it fast or slow? -Kees -- Kees Cook Chrome OS & Brillo Security
[toc] | [prev] | [next] | [standalone]
| From | Laura Abbott <labbott@redhat.com> |
|---|---|
| Date | 2016-02-03 22:40 +0100 |
| Message-ID | <qYdjA-7nl-11@gated-at.bofh.it> |
| In reply to | #1325949 |
On 02/03/2016 01:06 PM, Kees Cook wrote: > On Wed, Feb 3, 2016 at 10:46 AM, Laura Abbott <labbott@redhat.com> wrote: >> On 01/25/2016 11:03 PM, Joonsoo Kim wrote: >>> >>> On Mon, Jan 25, 2016 at 05:15:10PM -0800, Laura Abbott wrote: >>>> >>>> Hi, >>>> >>>> Based on the discussion from the series to add slab sanitization >>>> (lkml.kernel.org/g/<1450755641-7856-1-git-send-email-laura@labbott.name>) >>>> the existing SLAB_POISON mechanism already covers similar behavior. >>>> The performance of SLAB_POISON isn't very good. With hackbench -g 20 -l >>>> 1000 >>>> on QEMU with one cpu: >>> >>> >>> I doesn't follow up that discussion, but, I think that reusing >>> SLAB_POISON for slab sanitization needs more changes. I assume that >>> completeness and performance is matter for slab sanitization. >>> >>> 1) SLAB_POISON isn't applied to specific kmem_cache which has >>> constructor or SLAB_DESTROY_BY_RCU flag. For debug, it's not necessary >>> to be applied, but, for slab sanitization, it is better to apply it to >>> all caches. >> >> >> The grsecurity patches get around this by calling the constructor again >> after poisoning. It could be worth investigating doing that as well >> although my focus was on the cases without the constructor. >>> >>> >>> 2) SLAB_POISON makes object size bigger so natural alignment will be >>> broken. For example, kmalloc(256) cache's size is 256 in normal >>> case but it would be 264 when SLAB_POISON is enabled. This causes >>> memory waste. >> >> >> The grsecurity patches also bump the size up to put the free pointer >> outside the object. For sanitization purposes it is cleaner to have >> no pointers in the object after free >> >>> >>> In fact, I'd prefer not reusing SLAB_POISON. It would make thing >>> simpler. But, it's up to Christoph. >>> >>> Thanks. >>> >> >> It basically looks like trying to poison on the fast path at all >> will have a negative impact even with the feature is turned off. >> Christoph has indicated this is not acceptable so we are forced >> to limit it to the slow path only if we want runtime enablement. > > Is it possible to have both? i.e fast path via CONFIG, and slow path > via runtime options? > That's what this patch series had. A Kconfig to turn the fast path debugging on and off. When the Kconfig is off it reverts back to the existing behavior and there is no fastpath penalty. >> If we're limited to the slow path only, we might as well work >> with SLAB_POISON to make it faster. We can reevaluate if it turns >> out the poisoning isn't fast enough to be useful. > > And since I'm new to this area, I know of fast/slow path in the > syscall sense. What happens in the allocation/free fast/slow path that > makes it fast or slow? The fast path uses the per cpu caches. No locks are taken and there is no IRQ disabling. For concurrency protection this comment explains it best: /* * The cmpxchg will only match if there was no additional * operation and if we are on the right processor. * * The cmpxchg does the following atomically (without lock * semantics!) * 1. Relocate first pointer to the current per cpu area. * 2. Verify that tid and freelist have not been changed * 3. If they were not changed replace tid and freelist * * Since this is without lock semantics the protection is only * against code executing on this cpu *not* from access by * other cpus. */ in the slow path, IRQs and locks have to be taken at the minimum. The debug options disable ever loading the per CPU caches so it always falls back to the slow path. > > -Kees > Thanks, Laura
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-02-04 00:10 +0100 |
| Message-ID | <qYeIH-fd-35@gated-at.bofh.it> |
| In reply to | #1325981 |
> The fast path uses the per cpu caches. No locks are taken and there > is no IRQ disabling. For concurrency protection this comment > explains it best: > > /* > * The cmpxchg will only match if there was no additional > * operation and if we are on the right processor. > * > * The cmpxchg does the following atomically (without lock > * semantics!) > * 1. Relocate first pointer to the current per cpu area. > * 2. Verify that tid and freelist have not been changed > * 3. If they were not changed replace tid and freelist > * > * Since this is without lock semantics the protection is only > * against code executing on this cpu *not* from access by > * other cpus. > */ > > in the slow path, IRQs and locks have to be taken at the minimum. > The debug options disable ever loading the per CPU caches so it > always falls back to the slow path. You could add the use of per cpu lists to the slow paths as well in order to increase performance. Then weave in the debugging options. But the performance of the fast path is critical to the overall performance of the kernel as a whole since this is a heavily used code path for many subsystems.
[toc] | [prev] | [next] | [standalone]
| From | Laura Abbott <labbott@redhat.com> |
|---|---|
| Date | 2016-02-04 01:50 +0100 |
| Message-ID | <qYghr-1kT-5@gated-at.bofh.it> |
| In reply to | #1326130 |
On 02/03/2016 03:02 PM, Christoph Lameter wrote: >> The fast path uses the per cpu caches. No locks are taken and there >> is no IRQ disabling. For concurrency protection this comment >> explains it best: >> >> /* >> * The cmpxchg will only match if there was no additional >> * operation and if we are on the right processor. >> * >> * The cmpxchg does the following atomically (without lock >> * semantics!) >> * 1. Relocate first pointer to the current per cpu area. >> * 2. Verify that tid and freelist have not been changed >> * 3. If they were not changed replace tid and freelist >> * >> * Since this is without lock semantics the protection is only >> * against code executing on this cpu *not* from access by >> * other cpus. >> */ >> >> in the slow path, IRQs and locks have to be taken at the minimum. >> The debug options disable ever loading the per CPU caches so it >> always falls back to the slow path. > > You could add the use of per cpu lists to the slow paths as well in > order > to increase performance. Then weave in the debugging options. > How would that work? The use of the CPU caches is what defines the fast path so I'm not sure how to add them in on the slow path and not affect the fast path. > But the performance of the fast path is critical to the overall > performance of the kernel as a whole since this is a heavily used code > path for many subsystems. > I also notice that __CMPXCHG_DOUBLE is turned off when the debug options are turned on. I don't see any details about why. What's the reason for turning it off when the debug options are enabled? Thanks, Laura
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-02-04 04:30 +0100 |
| Message-ID | <qYiMh-31h-5@gated-at.bofh.it> |
| In reply to | #1326318 |
On Wed, 3 Feb 2016, Laura Abbott wrote: > I also notice that __CMPXCHG_DOUBLE is turned off when the debug > options are turned on. I don't see any details about why. What's > the reason for turning it off when the debug options are enabled? Because operations on the object need to be locked out while the debug code is running. Otherwise concurrent operations from other processors could lead to weird object states. The object needs to be stable for debug checks. Poisoning and the related checks need that otherwise you will get sporadic false positives.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web