Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1730614 > unrolled thread

[PATCH 4/5] mm:swap: respect page_cluster for readahead

Started byMinchan Kim <minchan@kernel.org>
First post2017-09-12 04:40 +0200
Last post2017-09-13 03:00 +0200
Articles 12 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH 4/5] mm:swap: respect page_cluster for readahead Minchan Kim <minchan@kernel.org> - 2017-09-12 04:40 +0200
    Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead "Huang\, Ying" <ying.huang@intel.com> - 2017-09-12 07:30 +0200
      Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead Minchan Kim <minchan@kernel.org> - 2017-09-12 08:30 +0200
        Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead "Huang\, Ying" <ying.huang@intel.com> - 2017-09-12 08:50 +0200
          Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead Minchan Kim <minchan@kernel.org> - 2017-09-12 09:00 +0200
            Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead "Huang\, Ying" <ying.huang@intel.com> - 2017-09-12 09:30 +0200
              Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead Minchan Kim <minchan@kernel.org> - 2017-09-12 10:00 +0200
                Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead "Huang\, Ying" <ying.huang@intel.com> - 2017-09-12 10:10 +0200
                  Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead Minchan Kim <minchan@kernel.org> - 2017-09-12 10:30 +0200
                    Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead "Huang\, Ying" <ying.huang@intel.com> - 2017-09-12 10:40 +0200
                      Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead Minchan Kim <minchan@kernel.org> - 2017-09-13 01:40 +0200
                        Re: [PATCH 4/5] mm:swap: respect page_cluster for readahead "Huang\, Ying" <ying.huang@intel.com> - 2017-09-13 03:00 +0200

#1730614 — [PATCH 4/5] mm:swap: respect page_cluster for readahead

FromMinchan Kim <minchan@kernel.org>
Date2017-09-12 04:40 +0200
Subject[PATCH 4/5] mm:swap: respect page_cluster for readahead
Message-ID<uoJhf-3C9-3@gated-at.bofh.it>
page_cluster 0 means "we don't want readahead" so in the case,
let's skip the readahead detection logic.

Cc: "Huang, Ying" <ying.huang@intel.com>
Signed-off-by: Minchan Kim <minchan@kernel.org>
---
 include/linux/swap.h | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/include/linux/swap.h b/include/linux/swap.h
index 0f54b491e118..739d94397c47 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
 
 static inline bool swap_use_vma_readahead(void)
 {
-	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
+	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
+				&& !atomic_read(&nr_rotate_swap);
 }
 
 /* Swap 50% full? Release swapcache more aggressively.. */
-- 
2.7.4

[toc] | [next] | [standalone]


#1730668

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-09-12 07:30 +0200
Message-ID<uoLVM-5vV-5@gated-at.bofh.it>
In reply to#1730614
Minchan Kim <minchan@kernel.org> writes:

> page_cluster 0 means "we don't want readahead" so in the case,
> let's skip the readahead detection logic.
>
> Cc: "Huang, Ying" <ying.huang@intel.com>
> Signed-off-by: Minchan Kim <minchan@kernel.org>
> ---
>  include/linux/swap.h | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/include/linux/swap.h b/include/linux/swap.h
> index 0f54b491e118..739d94397c47 100644
> --- a/include/linux/swap.h
> +++ b/include/linux/swap.h
> @@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
>  
>  static inline bool swap_use_vma_readahead(void)
>  {
> -	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
> +	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
> +				&& !atomic_read(&nr_rotate_swap);
>  }
>  
>  /* Swap 50% full? Release swapcache more aggressively.. */

Now the readahead window size of the VMA based swap readahead is
controlled by /sys/kernel/mm/swap/vma_ra_max_order, while that of the
original swap readahead is controlled by sysctl page_cluster.  It is
possible for anonymous memory to use VMA based swap readahead and tmpfs
to use original swap readahead algorithm at the same time.  So that, I
think it is necessary to use different control knob to control these two
algorithm.  So if we want to disable readahead for tmpfs, but keep it
for VMA based readahead, we can set 0 to page_cluster but non-zero to
/sys/kernel/mm/swap/vma_ra_max_order.  With your change, this will be
impossible.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1730675

FromMinchan Kim <minchan@kernel.org>
Date2017-09-12 08:30 +0200
Message-ID<uoMRP-68v-1@gated-at.bofh.it>
In reply to#1730668
On Tue, Sep 12, 2017 at 01:23:01PM +0800, Huang, Ying wrote:
> Minchan Kim <minchan@kernel.org> writes:
> 
> > page_cluster 0 means "we don't want readahead" so in the case,
> > let's skip the readahead detection logic.
> >
> > Cc: "Huang, Ying" <ying.huang@intel.com>
> > Signed-off-by: Minchan Kim <minchan@kernel.org>
> > ---
> >  include/linux/swap.h | 3 ++-
> >  1 file changed, 2 insertions(+), 1 deletion(-)
> >
> > diff --git a/include/linux/swap.h b/include/linux/swap.h
> > index 0f54b491e118..739d94397c47 100644
> > --- a/include/linux/swap.h
> > +++ b/include/linux/swap.h
> > @@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
> >  
> >  static inline bool swap_use_vma_readahead(void)
> >  {
> > -	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
> > +	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
> > +				&& !atomic_read(&nr_rotate_swap);
> >  }
> >  
> >  /* Swap 50% full? Release swapcache more aggressively.. */
> 
> Now the readahead window size of the VMA based swap readahead is
> controlled by /sys/kernel/mm/swap/vma_ra_max_order, while that of the
> original swap readahead is controlled by sysctl page_cluster.  It is
> possible for anonymous memory to use VMA based swap readahead and tmpfs
> to use original swap readahead algorithm at the same time.  So that, I
> think it is necessary to use different control knob to control these two
> algorithm.  So if we want to disable readahead for tmpfs, but keep it
> for VMA based readahead, we can set 0 to page_cluster but non-zero to
> /sys/kernel/mm/swap/vma_ra_max_order.  With your change, this will be
> impossible.

For a long time, page-cluster have been used as controlling swap readahead.
One of example, zram users have been disabled readahead via 0 page-cluster.
However, with your change, it would be regressed if it doesn't disable
vma_ra_max_order.

As well, all of swap users should be aware of vma_ra_max_order as well as
page-cluster to control swap readahead but I didn't see any document about
that. Acutaully, I don't like it but want to unify it with page-cluster.

[toc] | [prev] | [next] | [standalone]


#1730684

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-09-12 08:50 +0200
Message-ID<uoNbc-6fX-9@gated-at.bofh.it>
In reply to#1730675
Minchan Kim <minchan@kernel.org> writes:

> On Tue, Sep 12, 2017 at 01:23:01PM +0800, Huang, Ying wrote:
>> Minchan Kim <minchan@kernel.org> writes:
>> 
>> > page_cluster 0 means "we don't want readahead" so in the case,
>> > let's skip the readahead detection logic.
>> >
>> > Cc: "Huang, Ying" <ying.huang@intel.com>
>> > Signed-off-by: Minchan Kim <minchan@kernel.org>
>> > ---
>> >  include/linux/swap.h | 3 ++-
>> >  1 file changed, 2 insertions(+), 1 deletion(-)
>> >
>> > diff --git a/include/linux/swap.h b/include/linux/swap.h
>> > index 0f54b491e118..739d94397c47 100644
>> > --- a/include/linux/swap.h
>> > +++ b/include/linux/swap.h
>> > @@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
>> >  
>> >  static inline bool swap_use_vma_readahead(void)
>> >  {
>> > -	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
>> > +	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
>> > +				&& !atomic_read(&nr_rotate_swap);
>> >  }
>> >  
>> >  /* Swap 50% full? Release swapcache more aggressively.. */
>> 
>> Now the readahead window size of the VMA based swap readahead is
>> controlled by /sys/kernel/mm/swap/vma_ra_max_order, while that of the
>> original swap readahead is controlled by sysctl page_cluster.  It is
>> possible for anonymous memory to use VMA based swap readahead and tmpfs
>> to use original swap readahead algorithm at the same time.  So that, I
>> think it is necessary to use different control knob to control these two
>> algorithm.  So if we want to disable readahead for tmpfs, but keep it
>> for VMA based readahead, we can set 0 to page_cluster but non-zero to
>> /sys/kernel/mm/swap/vma_ra_max_order.  With your change, this will be
>> impossible.
>
> For a long time, page-cluster have been used as controlling swap readahead.
> One of example, zram users have been disabled readahead via 0 page-cluster.
> However, with your change, it would be regressed if it doesn't disable
> vma_ra_max_order.
>
> As well, all of swap users should be aware of vma_ra_max_order as well as
> page-cluster to control swap readahead but I didn't see any document about
> that. Acutaully, I don't like it but want to unify it with page-cluster.

The document is in

Documentation/ABI/testing/sysfs-kernel-mm-swap

The concern of unifying it with page-cluster is as following.

Original swap readahead on tmpfs may not work well because the combined
workload is running, so we want to disable or constrain it.  But at the
same time, the VMA based swap readahead may work better.  So I think it
may be necessary to control them separately.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1730687

FromMinchan Kim <minchan@kernel.org>
Date2017-09-12 09:00 +0200
Message-ID<uoNkS-6jf-9@gated-at.bofh.it>
In reply to#1730684
On Tue, Sep 12, 2017 at 02:44:36PM +0800, Huang, Ying wrote:
> Minchan Kim <minchan@kernel.org> writes:
> 
> > On Tue, Sep 12, 2017 at 01:23:01PM +0800, Huang, Ying wrote:
> >> Minchan Kim <minchan@kernel.org> writes:
> >> 
> >> > page_cluster 0 means "we don't want readahead" so in the case,
> >> > let's skip the readahead detection logic.
> >> >
> >> > Cc: "Huang, Ying" <ying.huang@intel.com>
> >> > Signed-off-by: Minchan Kim <minchan@kernel.org>
> >> > ---
> >> >  include/linux/swap.h | 3 ++-
> >> >  1 file changed, 2 insertions(+), 1 deletion(-)
> >> >
> >> > diff --git a/include/linux/swap.h b/include/linux/swap.h
> >> > index 0f54b491e118..739d94397c47 100644
> >> > --- a/include/linux/swap.h
> >> > +++ b/include/linux/swap.h
> >> > @@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
> >> >  
> >> >  static inline bool swap_use_vma_readahead(void)
> >> >  {
> >> > -	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
> >> > +	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
> >> > +				&& !atomic_read(&nr_rotate_swap);
> >> >  }
> >> >  
> >> >  /* Swap 50% full? Release swapcache more aggressively.. */
> >> 
> >> Now the readahead window size of the VMA based swap readahead is
> >> controlled by /sys/kernel/mm/swap/vma_ra_max_order, while that of the
> >> original swap readahead is controlled by sysctl page_cluster.  It is
> >> possible for anonymous memory to use VMA based swap readahead and tmpfs
> >> to use original swap readahead algorithm at the same time.  So that, I
> >> think it is necessary to use different control knob to control these two
> >> algorithm.  So if we want to disable readahead for tmpfs, but keep it
> >> for VMA based readahead, we can set 0 to page_cluster but non-zero to
> >> /sys/kernel/mm/swap/vma_ra_max_order.  With your change, this will be
> >> impossible.
> >
> > For a long time, page-cluster have been used as controlling swap readahead.
> > One of example, zram users have been disabled readahead via 0 page-cluster.
> > However, with your change, it would be regressed if it doesn't disable
> > vma_ra_max_order.
> >
> > As well, all of swap users should be aware of vma_ra_max_order as well as
> > page-cluster to control swap readahead but I didn't see any document about
> > that. Acutaully, I don't like it but want to unify it with page-cluster.
> 
> The document is in
> 
> Documentation/ABI/testing/sysfs-kernel-mm-swap
> 
> The concern of unifying it with page-cluster is as following.
> 
> Original swap readahead on tmpfs may not work well because the combined
> workload is running, so we want to disable or constrain it.  But at the
> same time, the VMA based swap readahead may work better.  So I think it
> may be necessary to control them separately.

My concern is users have been disabled swap readahead by page-cluster would
be regressed. Please take care of them.

[toc] | [prev] | [next] | [standalone]


#1730710

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-09-12 09:30 +0200
Message-ID<uoNNU-6JC-29@gated-at.bofh.it>
In reply to#1730687
Minchan Kim <minchan@kernel.org> writes:

> On Tue, Sep 12, 2017 at 02:44:36PM +0800, Huang, Ying wrote:
>> Minchan Kim <minchan@kernel.org> writes:
>> 
>> > On Tue, Sep 12, 2017 at 01:23:01PM +0800, Huang, Ying wrote:
>> >> Minchan Kim <minchan@kernel.org> writes:
>> >> 
>> >> > page_cluster 0 means "we don't want readahead" so in the case,
>> >> > let's skip the readahead detection logic.
>> >> >
>> >> > Cc: "Huang, Ying" <ying.huang@intel.com>
>> >> > Signed-off-by: Minchan Kim <minchan@kernel.org>
>> >> > ---
>> >> >  include/linux/swap.h | 3 ++-
>> >> >  1 file changed, 2 insertions(+), 1 deletion(-)
>> >> >
>> >> > diff --git a/include/linux/swap.h b/include/linux/swap.h
>> >> > index 0f54b491e118..739d94397c47 100644
>> >> > --- a/include/linux/swap.h
>> >> > +++ b/include/linux/swap.h
>> >> > @@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
>> >> >  
>> >> >  static inline bool swap_use_vma_readahead(void)
>> >> >  {
>> >> > -	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
>> >> > +	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
>> >> > +				&& !atomic_read(&nr_rotate_swap);
>> >> >  }
>> >> >  
>> >> >  /* Swap 50% full? Release swapcache more aggressively.. */
>> >> 
>> >> Now the readahead window size of the VMA based swap readahead is
>> >> controlled by /sys/kernel/mm/swap/vma_ra_max_order, while that of the
>> >> original swap readahead is controlled by sysctl page_cluster.  It is
>> >> possible for anonymous memory to use VMA based swap readahead and tmpfs
>> >> to use original swap readahead algorithm at the same time.  So that, I
>> >> think it is necessary to use different control knob to control these two
>> >> algorithm.  So if we want to disable readahead for tmpfs, but keep it
>> >> for VMA based readahead, we can set 0 to page_cluster but non-zero to
>> >> /sys/kernel/mm/swap/vma_ra_max_order.  With your change, this will be
>> >> impossible.
>> >
>> > For a long time, page-cluster have been used as controlling swap readahead.
>> > One of example, zram users have been disabled readahead via 0 page-cluster.
>> > However, with your change, it would be regressed if it doesn't disable
>> > vma_ra_max_order.
>> >
>> > As well, all of swap users should be aware of vma_ra_max_order as well as
>> > page-cluster to control swap readahead but I didn't see any document about
>> > that. Acutaully, I don't like it but want to unify it with page-cluster.
>> 
>> The document is in
>> 
>> Documentation/ABI/testing/sysfs-kernel-mm-swap
>> 
>> The concern of unifying it with page-cluster is as following.
>> 
>> Original swap readahead on tmpfs may not work well because the combined
>> workload is running, so we want to disable or constrain it.  But at the
>> same time, the VMA based swap readahead may work better.  So I think it
>> may be necessary to control them separately.
>
> My concern is users have been disabled swap readahead by page-cluster would
> be regressed. Please take care of them.

How about disable VMA based swap readahead if zram used as swap?  Like
we have done for hard disk?

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1730722

FromMinchan Kim <minchan@kernel.org>
Date2017-09-12 10:00 +0200
Message-ID<uoOgW-6V5-11@gated-at.bofh.it>
In reply to#1730710
On Tue, Sep 12, 2017 at 03:29:45PM +0800, Huang, Ying wrote:
> Minchan Kim <minchan@kernel.org> writes:
> 
> > On Tue, Sep 12, 2017 at 02:44:36PM +0800, Huang, Ying wrote:
> >> Minchan Kim <minchan@kernel.org> writes:
> >> 
> >> > On Tue, Sep 12, 2017 at 01:23:01PM +0800, Huang, Ying wrote:
> >> >> Minchan Kim <minchan@kernel.org> writes:
> >> >> 
> >> >> > page_cluster 0 means "we don't want readahead" so in the case,
> >> >> > let's skip the readahead detection logic.
> >> >> >
> >> >> > Cc: "Huang, Ying" <ying.huang@intel.com>
> >> >> > Signed-off-by: Minchan Kim <minchan@kernel.org>
> >> >> > ---
> >> >> >  include/linux/swap.h | 3 ++-
> >> >> >  1 file changed, 2 insertions(+), 1 deletion(-)
> >> >> >
> >> >> > diff --git a/include/linux/swap.h b/include/linux/swap.h
> >> >> > index 0f54b491e118..739d94397c47 100644
> >> >> > --- a/include/linux/swap.h
> >> >> > +++ b/include/linux/swap.h
> >> >> > @@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
> >> >> >  
> >> >> >  static inline bool swap_use_vma_readahead(void)
> >> >> >  {
> >> >> > -	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
> >> >> > +	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
> >> >> > +				&& !atomic_read(&nr_rotate_swap);
> >> >> >  }
> >> >> >  
> >> >> >  /* Swap 50% full? Release swapcache more aggressively.. */
> >> >> 
> >> >> Now the readahead window size of the VMA based swap readahead is
> >> >> controlled by /sys/kernel/mm/swap/vma_ra_max_order, while that of the
> >> >> original swap readahead is controlled by sysctl page_cluster.  It is
> >> >> possible for anonymous memory to use VMA based swap readahead and tmpfs
> >> >> to use original swap readahead algorithm at the same time.  So that, I
> >> >> think it is necessary to use different control knob to control these two
> >> >> algorithm.  So if we want to disable readahead for tmpfs, but keep it
> >> >> for VMA based readahead, we can set 0 to page_cluster but non-zero to
> >> >> /sys/kernel/mm/swap/vma_ra_max_order.  With your change, this will be
> >> >> impossible.
> >> >
> >> > For a long time, page-cluster have been used as controlling swap readahead.
> >> > One of example, zram users have been disabled readahead via 0 page-cluster.
> >> > However, with your change, it would be regressed if it doesn't disable
> >> > vma_ra_max_order.
> >> >
> >> > As well, all of swap users should be aware of vma_ra_max_order as well as
> >> > page-cluster to control swap readahead but I didn't see any document about
> >> > that. Acutaully, I don't like it but want to unify it with page-cluster.
> >> 
> >> The document is in
> >> 
> >> Documentation/ABI/testing/sysfs-kernel-mm-swap
> >> 
> >> The concern of unifying it with page-cluster is as following.
> >> 
> >> Original swap readahead on tmpfs may not work well because the combined
> >> workload is running, so we want to disable or constrain it.  But at the
> >> same time, the VMA based swap readahead may work better.  So I think it
> >> may be necessary to control them separately.
> >
> > My concern is users have been disabled swap readahead by page-cluster would
> > be regressed. Please take care of them.
> 
> How about disable VMA based swap readahead if zram used as swap?  Like
> we have done for hard disk?

It could be with SWP_SYNCHRONOUS_IO flag which indicates super-fast,
no seek cost swap devices if this patchset is merged so VM automatically
disables readahead. It is in my TODO but it's orthogonal work.

The problem I raised is "Why shouldn't we obey user's decision?",
not zram sepcific issue.

A user has used SSD as swap devices decided to disable swap readahead
by some reason(e.g., small memory system). Anyway, it has worked
via page-cluster for a several years but with vma-based swap devices,
it doesn't work any more.

[toc] | [prev] | [next] | [standalone]


#1730726

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-09-12 10:10 +0200
Message-ID<uoOqB-7dv-3@gated-at.bofh.it>
In reply to#1730722
Minchan Kim <minchan@kernel.org> writes:

> On Tue, Sep 12, 2017 at 03:29:45PM +0800, Huang, Ying wrote:
>> Minchan Kim <minchan@kernel.org> writes:
>> 
>> > On Tue, Sep 12, 2017 at 02:44:36PM +0800, Huang, Ying wrote:
>> >> Minchan Kim <minchan@kernel.org> writes:
>> >> 
>> >> > On Tue, Sep 12, 2017 at 01:23:01PM +0800, Huang, Ying wrote:
>> >> >> Minchan Kim <minchan@kernel.org> writes:
>> >> >> 
>> >> >> > page_cluster 0 means "we don't want readahead" so in the case,
>> >> >> > let's skip the readahead detection logic.
>> >> >> >
>> >> >> > Cc: "Huang, Ying" <ying.huang@intel.com>
>> >> >> > Signed-off-by: Minchan Kim <minchan@kernel.org>
>> >> >> > ---
>> >> >> >  include/linux/swap.h | 3 ++-
>> >> >> >  1 file changed, 2 insertions(+), 1 deletion(-)
>> >> >> >
>> >> >> > diff --git a/include/linux/swap.h b/include/linux/swap.h
>> >> >> > index 0f54b491e118..739d94397c47 100644
>> >> >> > --- a/include/linux/swap.h
>> >> >> > +++ b/include/linux/swap.h
>> >> >> > @@ -427,7 +427,8 @@ extern bool has_usable_swap(void);
>> >> >> >  
>> >> >> >  static inline bool swap_use_vma_readahead(void)
>> >> >> >  {
>> >> >> > -	return READ_ONCE(swap_vma_readahead) && !atomic_read(&nr_rotate_swap);
>> >> >> > +	return page_cluster > 0 && READ_ONCE(swap_vma_readahead)
>> >> >> > +				&& !atomic_read(&nr_rotate_swap);
>> >> >> >  }
>> >> >> >  
>> >> >> >  /* Swap 50% full? Release swapcache more aggressively.. */
>> >> >> 
>> >> >> Now the readahead window size of the VMA based swap readahead is
>> >> >> controlled by /sys/kernel/mm/swap/vma_ra_max_order, while that of the
>> >> >> original swap readahead is controlled by sysctl page_cluster.  It is
>> >> >> possible for anonymous memory to use VMA based swap readahead and tmpfs
>> >> >> to use original swap readahead algorithm at the same time.  So that, I
>> >> >> think it is necessary to use different control knob to control these two
>> >> >> algorithm.  So if we want to disable readahead for tmpfs, but keep it
>> >> >> for VMA based readahead, we can set 0 to page_cluster but non-zero to
>> >> >> /sys/kernel/mm/swap/vma_ra_max_order.  With your change, this will be
>> >> >> impossible.
>> >> >
>> >> > For a long time, page-cluster have been used as controlling swap readahead.
>> >> > One of example, zram users have been disabled readahead via 0 page-cluster.
>> >> > However, with your change, it would be regressed if it doesn't disable
>> >> > vma_ra_max_order.
>> >> >
>> >> > As well, all of swap users should be aware of vma_ra_max_order as well as
>> >> > page-cluster to control swap readahead but I didn't see any document about
>> >> > that. Acutaully, I don't like it but want to unify it with page-cluster.
>> >> 
>> >> The document is in
>> >> 
>> >> Documentation/ABI/testing/sysfs-kernel-mm-swap
>> >> 
>> >> The concern of unifying it with page-cluster is as following.
>> >> 
>> >> Original swap readahead on tmpfs may not work well because the combined
>> >> workload is running, so we want to disable or constrain it.  But at the
>> >> same time, the VMA based swap readahead may work better.  So I think it
>> >> may be necessary to control them separately.
>> >
>> > My concern is users have been disabled swap readahead by page-cluster would
>> > be regressed. Please take care of them.
>> 
>> How about disable VMA based swap readahead if zram used as swap?  Like
>> we have done for hard disk?
>
> It could be with SWP_SYNCHRONOUS_IO flag which indicates super-fast,
> no seek cost swap devices if this patchset is merged so VM automatically
> disables readahead. It is in my TODO but it's orthogonal work.
>
> The problem I raised is "Why shouldn't we obey user's decision?",
> not zram sepcific issue.
>
> A user has used SSD as swap devices decided to disable swap readahead
> by some reason(e.g., small memory system). Anyway, it has worked
> via page-cluster for a several years but with vma-based swap devices,
> it doesn't work any more.

Can they add one more line to their configuration scripts?

echo 0 > /sys/kernel/mm/swap/vma_ra_max_order

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1730731

FromMinchan Kim <minchan@kernel.org>
Date2017-09-12 10:30 +0200
Message-ID<uoOJX-7k7-7@gated-at.bofh.it>
In reply to#1730726
On Tue, Sep 12, 2017 at 04:07:01PM +0800, Huang, Ying wrote:
< snip >
> >> > My concern is users have been disabled swap readahead by page-cluster would
> >> > be regressed. Please take care of them.
> >> 
> >> How about disable VMA based swap readahead if zram used as swap?  Like
> >> we have done for hard disk?
> >
> > It could be with SWP_SYNCHRONOUS_IO flag which indicates super-fast,
> > no seek cost swap devices if this patchset is merged so VM automatically
> > disables readahead. It is in my TODO but it's orthogonal work.
> >
> > The problem I raised is "Why shouldn't we obey user's decision?",
> > not zram sepcific issue.
> >
> > A user has used SSD as swap devices decided to disable swap readahead
> > by some reason(e.g., small memory system). Anyway, it has worked
> > via page-cluster for a several years but with vma-based swap devices,
> > it doesn't work any more.
> 
> Can they add one more line to their configuration scripts?
> 
> echo 0 > /sys/kernel/mm/swap/vma_ra_max_order

We call it as "regression", don't we?

[toc] | [prev] | [next] | [standalone]


#1730737

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-09-12 10:40 +0200
Message-ID<uoOTD-7nd-11@gated-at.bofh.it>
In reply to#1730731
Minchan Kim <minchan@kernel.org> writes:

> On Tue, Sep 12, 2017 at 04:07:01PM +0800, Huang, Ying wrote:
> < snip >
>> >> > My concern is users have been disabled swap readahead by page-cluster would
>> >> > be regressed. Please take care of them.
>> >> 
>> >> How about disable VMA based swap readahead if zram used as swap?  Like
>> >> we have done for hard disk?
>> >
>> > It could be with SWP_SYNCHRONOUS_IO flag which indicates super-fast,
>> > no seek cost swap devices if this patchset is merged so VM automatically
>> > disables readahead. It is in my TODO but it's orthogonal work.
>> >
>> > The problem I raised is "Why shouldn't we obey user's decision?",
>> > not zram sepcific issue.
>> >
>> > A user has used SSD as swap devices decided to disable swap readahead
>> > by some reason(e.g., small memory system). Anyway, it has worked
>> > via page-cluster for a several years but with vma-based swap devices,
>> > it doesn't work any more.
>> 
>> Can they add one more line to their configuration scripts?
>> 
>> echo 0 > /sys/kernel/mm/swap/vma_ra_max_order
>
> We call it as "regression", don't we?

I think this always happen when we switch default algorithm.  For
example, if we had switched default IO scheduler, then the user scripts
to configure the parameters of old default IO scheduler will fail.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1731280

FromMinchan Kim <minchan@kernel.org>
Date2017-09-13 01:40 +0200
Message-ID<up2WC-7W2-7@gated-at.bofh.it>
In reply to#1730737
On Tue, Sep 12, 2017 at 04:32:43PM +0800, Huang, Ying wrote:
> Minchan Kim <minchan@kernel.org> writes:
> 
> > On Tue, Sep 12, 2017 at 04:07:01PM +0800, Huang, Ying wrote:
> > < snip >
> >> >> > My concern is users have been disabled swap readahead by page-cluster would
> >> >> > be regressed. Please take care of them.
> >> >> 
> >> >> How about disable VMA based swap readahead if zram used as swap?  Like
> >> >> we have done for hard disk?
> >> >
> >> > It could be with SWP_SYNCHRONOUS_IO flag which indicates super-fast,
> >> > no seek cost swap devices if this patchset is merged so VM automatically
> >> > disables readahead. It is in my TODO but it's orthogonal work.
> >> >
> >> > The problem I raised is "Why shouldn't we obey user's decision?",
> >> > not zram sepcific issue.
> >> >
> >> > A user has used SSD as swap devices decided to disable swap readahead
> >> > by some reason(e.g., small memory system). Anyway, it has worked
> >> > via page-cluster for a several years but with vma-based swap devices,
> >> > it doesn't work any more.
> >> 
> >> Can they add one more line to their configuration scripts?
> >> 
> >> echo 0 > /sys/kernel/mm/swap/vma_ra_max_order
> >
> > We call it as "regression", don't we?
> 
> I think this always happen when we switch default algorithm.  For
> example, if we had switched default IO scheduler, then the user scripts
> to configure the parameters of old default IO scheduler will fail.

I don't follow what you are saying with specific example.
If kernel did it which breaks something on userspace which has worked well,
it should be fixed. No doubt.

Even although it happened by mistakes, it couldn't be a excuse to break
new thing, either.

Simple. Fix the regression. 

If you insist on "swap users should fix it by themselves via modification
of their script or it's not a regression", I don't want to waste my time to
persuade you any more. I will ask reverting your patches to Andrew.

[toc] | [prev] | [next] | [standalone]


#1731314

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-09-13 03:00 +0200
Message-ID<up4c2-bu-3@gated-at.bofh.it>
In reply to#1731280
Minchan Kim <minchan@kernel.org> writes:

> On Tue, Sep 12, 2017 at 04:32:43PM +0800, Huang, Ying wrote:
>> Minchan Kim <minchan@kernel.org> writes:
>> 
>> > On Tue, Sep 12, 2017 at 04:07:01PM +0800, Huang, Ying wrote:
>> > < snip >
>> >> >> > My concern is users have been disabled swap readahead by page-cluster would
>> >> >> > be regressed. Please take care of them.
>> >> >> 
>> >> >> How about disable VMA based swap readahead if zram used as swap?  Like
>> >> >> we have done for hard disk?
>> >> >
>> >> > It could be with SWP_SYNCHRONOUS_IO flag which indicates super-fast,
>> >> > no seek cost swap devices if this patchset is merged so VM automatically
>> >> > disables readahead. It is in my TODO but it's orthogonal work.
>> >> >
>> >> > The problem I raised is "Why shouldn't we obey user's decision?",
>> >> > not zram sepcific issue.
>> >> >
>> >> > A user has used SSD as swap devices decided to disable swap readahead
>> >> > by some reason(e.g., small memory system). Anyway, it has worked
>> >> > via page-cluster for a several years but with vma-based swap devices,
>> >> > it doesn't work any more.
>> >> 
>> >> Can they add one more line to their configuration scripts?
>> >> 
>> >> echo 0 > /sys/kernel/mm/swap/vma_ra_max_order
>> >
>> > We call it as "regression", don't we?
>> 
>> I think this always happen when we switch default algorithm.  For
>> example, if we had switched default IO scheduler, then the user scripts
>> to configure the parameters of old default IO scheduler will fail.
>
> I don't follow what you are saying with specific example.
> If kernel did it which breaks something on userspace which has worked well,
> it should be fixed. No doubt.
>
> Even although it happened by mistakes, it couldn't be a excuse to break
> new thing, either.
>
> Simple. Fix the regression. 
>
> If you insist on "swap users should fix it by themselves via modification
> of their script or it's not a regression", I don't want to waste my time to
> persuade you any more. I will ask reverting your patches to Andrew.

There is no functionality regression definitely.  It may cause some
performance regression for some users, which could be resolved via some
scripts changing.  Please don't mix them.

Best Regards,
Huang, Ying

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web