Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1303675 > unrolled thread

Re: [PATCH] autofs: show pipe inode in mount options

Started byStanislav Kinsburskiy <skinsbursky@odin.com>
First post2016-01-07 16:50 +0100
Last post2016-01-11 12:40 +0100
Articles 7 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH] autofs: show pipe inode in mount options Stanislav Kinsburskiy <skinsbursky@odin.com> - 2016-01-07 16:50 +0100
    Re: [PATCH] autofs: show pipe inode in mount options Ian Kent <raven@themaw.net> - 2016-01-08 08:30 +0100
      Re: [PATCH] autofs: show pipe inode in mount options Stanislav Kinsburskiy <skinsbursky@odin.com> - 2016-01-08 12:30 +0100
        Re: [PATCH] autofs: show pipe inode in mount options Ian Kent <raven@themaw.net> - 2016-01-08 14:00 +0100
          Re: [PATCH] autofs: show pipe inode in mount options Stanislav Kinsburskiy <skinsbursky@odin.com> - 2016-01-08 16:10 +0100
            Re: [PATCH] autofs: show pipe inode in mount options Ian Kent <raven@themaw.net> - 2016-01-09 02:40 +0100
              Re: [PATCH] autofs: show pipe inode in mount options Stanislav Kinsburskiy <skinsbursky@odin.com> - 2016-01-11 12:40 +0100

#1303675 — Re: [PATCH] autofs: show pipe inode in mount options

FromStanislav Kinsburskiy <skinsbursky@odin.com>
Date2016-01-07 16:50 +0100
SubjectRe: [PATCH] autofs: show pipe inode in mount options
Message-ID<qOkZ3-1DY-1@gated-at.bofh.it>
Good day, gentlemen.

Could you update, what's the status with this patch?
Without it it's impossible to match process pipe with kernel pipe, while 
this is "must have" to be able to migrate AutoFS via CRIU.



16.12.2015 13:02, Stanislav Kinsburskiy пишет:
> This is required for CRIU to migrate a mount point, when write end in user
> space is closed.
> To be able to migrate such mount, read end of the pipe have to be searched
> within autofs master process, and pipe inode will be used as a key.
>
> Signed-off-by: Stanislav Kinsburskiy <skinsbursky@virtuozzo.com>
> ---
>   fs/autofs4/inode.c |    4 ++++
>   1 file changed, 4 insertions(+)
>
> diff --git a/fs/autofs4/inode.c b/fs/autofs4/inode.c
> index a3ae0b2..16f875a 100644
> --- a/fs/autofs4/inode.c
> +++ b/fs/autofs4/inode.c
> @@ -77,6 +77,10 @@ static int autofs4_show_options(struct seq_file *m, struct dentry *root)
>   		return 0;
>   
>   	seq_printf(m, ",fd=%d", sbi->pipefd);
> +	if (sbi->pipe)
> +		seq_printf(m, ",pipe_ino=%ld", sbi->pipe->f_inode->i_ino);
> +	else
> +		seq_printf(m, ",pipe_ino=-1");
>   	if (!uid_eq(root_inode->i_uid, GLOBAL_ROOT_UID))
>   		seq_printf(m, ",uid=%u",
>   			from_kuid_munged(&init_user_ns, root_inode->i_uid));
>

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1304228

FromIan Kent <raven@themaw.net>
Date2016-01-08 08:30 +0100
Message-ID<qOzEK-3rg-15@gated-at.bofh.it>
In reply to#1303675
On Thu, 2016-01-07 at 16:46 +0100, Stanislav Kinsburskiy wrote:
> Good day, gentlemen.
> 
> Could you update, what's the status with this patch?
> Without it it's impossible to match process pipe with kernel pipe,
> while 
> this is "must have" to be able to migrate AutoFS via CRIU.

Right, I did mean to reply to this mail but have been distracted by
family stuff.

I don't know what CRIU is and people looking at changelog entries
shouldn't need to do a web search to find out.

Could you change it a little.

I'm also not sure whether to forward this (assuming the description is
updated a little) to Al or to include it in the series to rename
autofs4 to autofs that I'm hoping to ask be included in linux-next
fairly soon.

Passing it on to Al will likely interfere with the series coming from
linux-next so that could be bit of a hassle.

Another thing I'm wondering about is the order this entry will appear
at in the options. You order choice is sensible though and autofs
shouldn't have a problem with the inserted option but other
applications might.

Finally, and perhaps most importantly, I don't get what your trying to
do, you also haven't given any clues to that in the patch dscription.

IOW how do you expect to use this.

> 
> 
> 16.12.2015 13:02, Stanislav Kinsburskiy пишет:
> > This is required for CRIU to migrate a mount point, when write end
> > in user
> > space is closed.

Like I said what does this mean.

autofs doesn't need this when it re-constructs a mount tree from
existing mounts on re-start or after a SIGKILL on the automount
process.

How is this different and how will it be used?

The question to be answered here is "is this the best way to do it and
will it work for the autofs mount types you expect it to"?

> > To be able to migrate such mount, read end of the pipe have to be
> > searched
> > within autofs master process, and pipe inode will be used as a key.
> > 
> > Signed-off-by: Stanislav Kinsburskiy <skinsbursky@virtuozzo.com>
> > ---
> >   fs/autofs4/inode.c |    4 ++++
> >   1 file changed, 4 insertions(+)
> > 
> > diff --git a/fs/autofs4/inode.c b/fs/autofs4/inode.c
> > index a3ae0b2..16f875a 100644
> > --- a/fs/autofs4/inode.c
> > +++ b/fs/autofs4/inode.c
> > @@ -77,6 +77,10 @@ static int autofs4_show_options(struct seq_file
> > *m, struct dentry *root)
> >   		return 0;
> >   
> >   	seq_printf(m, ",fd=%d", sbi->pipefd);
> > +	if (sbi->pipe)
> > +		seq_printf(m, ",pipe_ino=%ld", sbi->pipe->f_inode
> > ->i_ino);
> > +	else
> > +		seq_printf(m, ",pipe_ino=-1");
> >   	if (!uid_eq(root_inode->i_uid, GLOBAL_ROOT_UID))
> >   		seq_printf(m, ",uid=%u",
> >   			from_kuid_munged(&init_user_ns,
> > root_inode->i_uid));
> > 
> 

[toc] | [prev] | [next] | [standalone]


#1304401

FromStanislav Kinsburskiy <skinsbursky@odin.com>
Date2016-01-08 12:30 +0100
Message-ID<qODp0-60g-15@gated-at.bofh.it>
In reply to#1304228

08.01.2016 08:20, Ian Kent пишет:
> On Thu, 2016-01-07 at 16:46 +0100, Stanislav Kinsburskiy wrote:
>> Good day, gentlemen.
>>
>> Could you update, what's the status with this patch?
>> Without it it's impossible to match process pipe with kernel pipe,
>> while
>> this is "must have" to be able to migrate AutoFS via CRIU.
> Right, I did mean to reply to this mail but have been distracted by
> family stuff.
>
> I don't know what CRIU is and people looking at changelog entries
> shouldn't need to do a web search to find out.
>
> Could you change it a little.

Fair enough. I'll resend with more descriptive message.
But first I would like to clarify to you the problem root and why it's 
done like this.

> I'm also not sure whether to forward this (assuming the description is
> updated a little) to Al or to include it in the series to rename
> autofs4 to autofs that I'm hoping to ask be included in linux-next
> fairly soon.

Here I don't know, what's better. Of course Al can take it as well. But, 
probably, first would be nice to make sure, that this solution is the 
best one.
Description of the problem is below.

> Passing it on to Al will likely interfere with the series coming from
> linux-next so that could be bit of a hassle.
>
> Another thing I'm wondering about is the order this entry will appear
> at in the options. You order choice is sensible though and autofs
> shouldn't have a problem with the inserted option but other
> applications might.

I should put it at the end, probably?

> Finally, and perhaps most importantly, I don't get what your trying to
> do, you also haven't given any clues to that in the patch dscription.
>
> IOW how do you expect to use this.
>
>>
>> 16.12.2015 13:02, Stanislav Kinsburskiy пишет:
>>> This is required for CRIU to migrate a mount point, when write end
>>> in user
>>> space is closed.
> Like I said what does this mean.
>
> autofs doesn't need this when it re-constructs a mount tree from
> existing mounts on re-start or after a SIGKILL on the automount
> process.
>
> How is this different and how will it be used?
>
> The question to be answered here is "is this the best way to do it and
> will it work for the autofs mount types you expect it to"?

So, here is a brief description of the problem.
To migrate autofs mount, one have to reconstruct control pipe between 
kernel and autofs master.
There are two cases I'm wiling to support:
1) Automount binary (autofs package). This program is very gentle and it 
doesn't close write end of the pipe after mount.
2) Systemd. This program closes write end of the pipe once the mount is 
done.

The autofs restore concept is the following:
1) Mount autofs from some process with some dummy pipe.
2) Fix it's pgrp, pipe fd, timeout, etc on top of existent mount in the 
right master later (this is because of implementation of CRIU mounts 
restore, where all of them are created by one process).
What is the most important here, is that during pipe reconstruction, 
read end of it have to be placed _exactly_ on the file descriptor, which 
process has before, thus allowing to autofs master still can play it's 
role after migration.

To be able to reconstruct control pipe, one must know _exactly_ on 
_dump_ stage, which descriptor in autofs master corresponds to read end 
of the pipe, because this pipe have to be empty, because we can't (and 
don't want to) transfer some interim state in the kernel via userspace 
migration solution.
In case of systemd, this write end is already closed, so searching for 
the read end is not possible.
In case of automount write end is still there (with the same fd as in 
mount options), but one can't be sure, that this descriptor is connected 
to file structure, which is used by kernel. There can be another pipe 
end (for example, in case of systemd).

Thus, to be able to find read end of the pipe in autofs master, some 
other mark is required.
The best solution is to use pipe inode number, which allows to match 
opened pipes in a process with 100% reliability.
And the easies solution I found is to expose this number is autofs mount 
options.
If you have another better solution, I would be glad to implement it.

Hope, the above explains the problem clearly.

>>> To be able to migrate such mount, read end of the pipe have to be
>>> searched
>>> within autofs master process, and pipe inode will be used as a key.
>>>
>>> Signed-off-by: Stanislav Kinsburskiy <skinsbursky@virtuozzo.com>
>>> ---
>>>    fs/autofs4/inode.c |    4 ++++
>>>    1 file changed, 4 insertions(+)
>>>
>>> diff --git a/fs/autofs4/inode.c b/fs/autofs4/inode.c
>>> index a3ae0b2..16f875a 100644
>>> --- a/fs/autofs4/inode.c
>>> +++ b/fs/autofs4/inode.c
>>> @@ -77,6 +77,10 @@ static int autofs4_show_options(struct seq_file
>>> *m, struct dentry *root)
>>>    		return 0;
>>>    
>>>    	seq_printf(m, ",fd=%d", sbi->pipefd);
>>> +	if (sbi->pipe)
>>> +		seq_printf(m, ",pipe_ino=%ld", sbi->pipe->f_inode
>>> ->i_ino);
>>> +	else
>>> +		seq_printf(m, ",pipe_ino=-1");
>>>    	if (!uid_eq(root_inode->i_uid, GLOBAL_ROOT_UID))
>>>    		seq_printf(m, ",uid=%u",
>>>    			from_kuid_munged(&init_user_ns,
>>> root_inode->i_uid));
>>>

[toc] | [prev] | [next] | [standalone]


#1304513

FromIan Kent <raven@themaw.net>
Date2016-01-08 14:00 +0100
Message-ID<qOEO7-6QV-33@gated-at.bofh.it>
In reply to#1304401
On Fri, 2016-01-08 at 12:29 +0100, Stanislav Kinsburskiy wrote:
> 
> 08.01.2016 08:20, Ian Kent пишет:
> > On Thu, 2016-01-07 at 16:46 +0100, Stanislav Kinsburskiy wrote:
> > > Good day, gentlemen.
> > > 
> > > Could you update, what's the status with this patch?
> > > Without it it's impossible to match process pipe with kernel
> > > pipe,
> > > while
> > > this is "must have" to be able to migrate AutoFS via CRIU.
> > Right, I did mean to reply to this mail but have been distracted by
> > family stuff.
> > 
> > I don't know what CRIU is and people looking at changelog entries
> > shouldn't need to do a web search to find out.
> > 
> > Could you change it a little.
> 
> Fair enough. I'll resend with more descriptive message.
> But first I would like to clarify to you the problem root and why
> it's 
> done like this.
> 
> > I'm also not sure whether to forward this (assuming the description
> > is
> > updated a little) to Al or to include it in the series to rename
> > autofs4 to autofs that I'm hoping to ask be included in linux-next
> > fairly soon.
> 
> Here I don't know, what's better. Of course Al can take it as well.
> But, 
> probably, first would be nice to make sure, that this solution is the
> best one.
> Description of the problem is below.
> 
> > Passing it on to Al will likely interfere with the series coming
> > from
> > linux-next so that could be bit of a hassle.
> > 
> > Another thing I'm wondering about is the order this entry will
> > appear
> > at in the options. You order choice is sensible though and autofs
> > shouldn't have a problem with the inserted option but other
> > applications might.
> 
> I should put it at the end, probably?
> 
> > Finally, and perhaps most importantly, I don't get what your trying
> > to
> > do, you also haven't given any clues to that in the patch
> > dscription.
> > 
> > IOW how do you expect to use this.
> > 
> > > 
> > > 16.12.2015 13:02, Stanislav Kinsburskiy пишет:
> > > > This is required for CRIU to migrate a mount point, when write
> > > > end
> > > > in user
> > > > space is closed.
> > Like I said what does this mean.
> > 
> > autofs doesn't need this when it re-constructs a mount tree from
> > existing mounts on re-start or after a SIGKILL on the automount
> > process.
> > 
> > How is this different and how will it be used?
> > 
> > The question to be answered here is "is this the best way to do it
> > and
> > will it work for the autofs mount types you expect it to"?
> 
> So, here is a brief description of the problem.
> To migrate autofs mount, one have to reconstruct control pipe between
> kernel and autofs master.
> There are two cases I'm wiling to support:
> 1) Automount binary (autofs package). This program is very gentle and
> it 
> doesn't close write end of the pipe after mount.
> 2) Systemd. This program closes write end of the pipe once the mount
> is 
> done.

I must admit I'm having trouble understanding the description.
Give me a little time with it.

I don't know how systemd works with autofs mounts only that it uses the
autofs direct mount type.

autofs uses both indirect and direct mounts and both can have offsets
(from the kernel POV semantically direct mounts). So there is quite a
bit to worry about beside the kernel pipe.

Anyway, it seems your only concern is the kernel pipe and I wonder why
you can't just set the mount catatonic (in autofs speak) on save and
open a new kernel pipe then set the pipefd on the autofs mount on
restore.

But probably my suggestion is far to simplistic as I get the impression
you have a process already in a given state which you want to restore.

One thing to keep in mind is that if an autofs mount is not set
catatonic any access other than the owner process (process group pid)
will hang unless there is an actual user space process to service the
callback.

Although I don't know the flow of things that might be important at
some point.

And if the mount is set catatonic the process needs to set the pipefd
to take "ownership" which also clears the mount catatonic flag.

Anyway, let me think about what you've written for a while.
Ian

[toc] | [prev] | [next] | [standalone]


#1304628

FromStanislav Kinsburskiy <skinsbursky@odin.com>
Date2016-01-08 16:10 +0100
Message-ID<qOGPU-8rM-15@gated-at.bofh.it>
In reply to#1304513

08.01.2016 13:58, Ian Kent пишет:
> On Fri, 2016-01-08 at 12:29 +0100, Stanislav Kinsburskiy wrote:
>> 08.01.2016 08:20, Ian Kent пишет:
>>> On Thu, 2016-01-07 at 16:46 +0100, Stanislav Kinsburskiy wrote:
>>>> Good day, gentlemen.
>>>>
>>>> Could you update, what's the status with this patch?
>>>> Without it it's impossible to match process pipe with kernel
>>>> pipe,
>>>> while
>>>> this is "must have" to be able to migrate AutoFS via CRIU.
>>> Right, I did mean to reply to this mail but have been distracted by
>>> family stuff.
>>>
>>> I don't know what CRIU is and people looking at changelog entries
>>> shouldn't need to do a web search to find out.
>>>
>>> Could you change it a little.
>> Fair enough. I'll resend with more descriptive message.
>> But first I would like to clarify to you the problem root and why
>> it's
>> done like this.
>>
>>> I'm also not sure whether to forward this (assuming the description
>>> is
>>> updated a little) to Al or to include it in the series to rename
>>> autofs4 to autofs that I'm hoping to ask be included in linux-next
>>> fairly soon.
>> Here I don't know, what's better. Of course Al can take it as well.
>> But,
>> probably, first would be nice to make sure, that this solution is the
>> best one.
>> Description of the problem is below.
>>
>>> Passing it on to Al will likely interfere with the series coming
>>> from
>>> linux-next so that could be bit of a hassle.
>>>
>>> Another thing I'm wondering about is the order this entry will
>>> appear
>>> at in the options. You order choice is sensible though and autofs
>>> shouldn't have a problem with the inserted option but other
>>> applications might.
>> I should put it at the end, probably?
>>
>>> Finally, and perhaps most importantly, I don't get what your trying
>>> to
>>> do, you also haven't given any clues to that in the patch
>>> dscription.
>>>
>>> IOW how do you expect to use this.
>>>
>>>> 16.12.2015 13:02, Stanislav Kinsburskiy пишет:
>>>>> This is required for CRIU to migrate a mount point, when write
>>>>> end
>>>>> in user
>>>>> space is closed.
>>> Like I said what does this mean.
>>>
>>> autofs doesn't need this when it re-constructs a mount tree from
>>> existing mounts on re-start or after a SIGKILL on the automount
>>> process.
>>>
>>> How is this different and how will it be used?
>>>
>>> The question to be answered here is "is this the best way to do it
>>> and
>>> will it work for the autofs mount types you expect it to"?
>> So, here is a brief description of the problem.
>> To migrate autofs mount, one have to reconstruct control pipe between
>> kernel and autofs master.
>> There are two cases I'm wiling to support:
>> 1) Automount binary (autofs package). This program is very gentle and
>> it
>> doesn't close write end of the pipe after mount.
>> 2) Systemd. This program closes write end of the pipe once the mount
>> is
>> done.
> I must admit I'm having trouble understanding the description.
> Give me a little time with it.
>
> I don't know how systemd works with autofs mounts only that it uses the
> autofs direct mount type.

Systemd closes write end of the pipe after mount.

> autofs uses both indirect and direct mounts and both can have offsets
> (from the kernel POV semantically direct mounts). So there is quite a
> bit to worry about beside the kernel pipe.

It's not about direct or indirects mounts.
It's about process state restore.
With CRIU migration, any task is frozen, then disassembled into pieces 
(dump files), which are used to assemble task exactly in the same state 
in was before dump.
The technology is very complex and uses a lot a different tricky 
techniques to make this possible in userspace to describe all the 
details here.

But below is a bit more information, which, hopefully, will clarify all 
this a little bit more.
One of a process attributed to migrate is "opened files". Pipes also 
belong to this attribute.

To restore a pipe CRIU does the following (a very simplified description):
1) Creates a new pipe.
2) Writes (previously stores in images) its contents via write end.
3) Duplicate pipe descriptors to the fds of the process, which were used 
before dump, if required
4) Send pipe descriptors to other processes, sharing it, via unix socket.
5) Close those pipe descriptors, which are not required (say, this 
process had only read end, while it's child had write end).

Thus in case of restoring and autofs mount of systemd (which, for 
example, closed write end and has read end on fd 40), one have to create 
a pipe (say, appeared with fd 5 and fd 6), fill it with content via fd 
6, duplicate fd 5 into fd 40, call mount with pipe fd 6 and then close fd 6.
This is, yet again, a very simple explanation.

> Anyway, it seems your only concern is the kernel pipe and I wonder why
> you can't just set the mount catatonic (in autofs speak) on save and
> open a new kernel pipe then set the pipefd on the autofs mount on
> restore.

I can't because of a bunch of reasons:
1) It can be migration, thus I don't have autofs mount on destination 
node at all
2) It can be a container, which is stopped after dump (thus mount point 
is destroyed).

> But probably my suggestion is far to simplistic as I get the impression
> you have a process already in a given state which you want to restore.
>
> One thing to keep in mind is that if an autofs mount is not set
> catatonic any access other than the owner process (process group pid)
> will hang unless there is an actual user space process to service the
> callback.
>
> Although I don't know the flow of things that might be important at
> some point.
>
> And if the mount is set catatonic the process needs to set the pipefd
> to take "ownership" which also clears the mount catatonic flag.

The migration is already implemented and sent to CRIU mailing list.
Here is the list, if you are interesting (I use kernel with this patch 
applied):
https://lists.openvz.org/pipermail/criu/2016-January/024749.html

> Anyway, let me think about what you've written for a while.
> Ian

Sure, take your time.

[toc] | [prev] | [next] | [standalone]


#1305130

FromIan Kent <raven@themaw.net>
Date2016-01-09 02:40 +0100
Message-ID<qOQFz-6Is-15@gated-at.bofh.it>
In reply to#1304628
On Fri, 2016-01-08 at 16:05 +0100, Stanislav Kinsburskiy wrote:
> 
> 08.01.2016 13:58, Ian Kent пишет:
> > On Fri, 2016-01-08 at 12:29 +0100, Stanislav Kinsburskiy wrote:
> > > 08.01.2016 08:20, Ian Kent пишет:
> > > > On Thu, 2016-01-07 at 16:46 +0100, Stanislav Kinsburskiy wrote:
> > > > > Good day, gentlemen.
> > > > > 
> > > > > Could you update, what's the status with this patch?
> > > > > Without it it's impossible to match process pipe with kernel
> > > > > pipe,
> > > > > while
> > > > > this is "must have" to be able to migrate AutoFS via CRIU.
> > > > Right, I did mean to reply to this mail but have been
> > > > distracted by
> > > > family stuff.
> > > > 
> > > > I don't know what CRIU is and people looking at changelog
> > > > entries
> > > > shouldn't need to do a web search to find out.
> > > > 
> > > > Could you change it a little.
> > > Fair enough. I'll resend with more descriptive message.
> > > But first I would like to clarify to you the problem root and why
> > > it's
> > > done like this.
> > > 
> > > > I'm also not sure whether to forward this (assuming the
> > > > description
> > > > is
> > > > updated a little) to Al or to include it in the series to
> > > > rename
> > > > autofs4 to autofs that I'm hoping to ask be included in linux
> > > > -next
> > > > fairly soon.
> > > Here I don't know, what's better. Of course Al can take it as
> > > well.
> > > But,
> > > probably, first would be nice to make sure, that this solution is
> > > the
> > > best one.
> > > Description of the problem is below.
> > > 
> > > > Passing it on to Al will likely interfere with the series
> > > > coming
> > > > from
> > > > linux-next so that could be bit of a hassle.
> > > > 
> > > > Another thing I'm wondering about is the order this entry will
> > > > appear
> > > > at in the options. You order choice is sensible though and
> > > > autofs
> > > > shouldn't have a problem with the inserted option but other
> > > > applications might.
> > > I should put it at the end, probably?
> > > 
> > > > Finally, and perhaps most importantly, I don't get what your
> > > > trying
> > > > to
> > > > do, you also haven't given any clues to that in the patch
> > > > dscription.
> > > > 
> > > > IOW how do you expect to use this.
> > > > 
> > > > > 16.12.2015 13:02, Stanislav Kinsburskiy пишет:
> > > > > > This is required for CRIU to migrate a mount point, when
> > > > > > write
> > > > > > end
> > > > > > in user
> > > > > > space is closed.
> > > > Like I said what does this mean.
> > > > 
> > > > autofs doesn't need this when it re-constructs a mount tree
> > > > from
> > > > existing mounts on re-start or after a SIGKILL on the automount
> > > > process.
> > > > 
> > > > How is this different and how will it be used?
> > > > 
> > > > The question to be answered here is "is this the best way to do
> > > > it
> > > > and
> > > > will it work for the autofs mount types you expect it to"?
> > > So, here is a brief description of the problem.
> > > To migrate autofs mount, one have to reconstruct control pipe
> > > between
> > > kernel and autofs master.
> > > There are two cases I'm wiling to support:
> > > 1) Automount binary (autofs package). This program is very gentle
> > > and
> > > it
> > > doesn't close write end of the pipe after mount.
> > > 2) Systemd. This program closes write end of the pipe once the
> > > mount
> > > is
> > > done.
> > I must admit I'm having trouble understanding the description.
> > Give me a little time with it.
> > 
> > I don't know how systemd works with autofs mounts only that it uses
> > the
> > autofs direct mount type.
> 
> Systemd closes write end of the pipe after mount.
> 
> > autofs uses both indirect and direct mounts and both can have
> > offsets
> > (from the kernel POV semantically direct mounts). So there is quite
> > a
> > bit to worry about beside the kernel pipe.
> 
> It's not about direct or indirects mounts.
> It's about process state restore.
> With CRIU migration, any task is frozen, then disassembled into
> pieces 
> (dump files), which are used to assemble task exactly in the same
> state 
> in was before dump.
> The technology is very complex and uses a lot a different tricky 
> techniques to make this possible in userspace to describe all the 
> details here.
> 
> But below is a bit more information, which, hopefully, will clarify
> all 
> this a little bit more.
> One of a process attributed to migrate is "opened files". Pipes also 
> belong to this attribute.
> 
> To restore a pipe CRIU does the following (a very simplified
> description):
> 1) Creates a new pipe.
> 2) Writes (previously stores in images) its contents via write end.
> 3) Duplicate pipe descriptors to the fds of the process, which were
> used 
> before dump, if required
> 4) Send pipe descriptors to other processes, sharing it, via unix
> socket.
> 5) Close those pipe descriptors, which are not required (say, this 
> process had only read end, while it's child had write end).
> 
> Thus in case of restoring and autofs mount of systemd (which, for 
> example, closed write end and has read end on fd 40), one have to
> create 
> a pipe (say, appeared with fd 5 and fd 6), fill it with content via
> fd 
> 6, duplicate fd 5 into fd 40, call mount with pipe fd 6 and then
> close fd 6.
> This is, yet again, a very simple explanation.

Right, as said initially (more or less), if you need the patch you
posted then it shouldn't cause a problem so it should be ok. Al hasn't
responded so I guess that means I should go the linux-next path
possibly via pull request for the series I have to rename autofs4 to
autofs (along with this one, to prevent merge conflicts).

I haven't asked Steven about this yet so I'm not sure if a pull request
is even the right thing to do.

There is another case I was wondering about.

That's when there is a direct mount that is covered by a real mount.

autofs will have a file handle open to it (on the underlying mount
point path) to use for control purposes like expires. I think you also
need to restore those file handles to restore process state and in this
case the mount point is covered.

> 
> > Anyway, it seems your only concern is the kernel pipe and I wonder
> > why
> > you can't just set the mount catatonic (in autofs speak) on save
> > and
> > open a new kernel pipe then set the pipefd on the autofs mount on
> > restore.
> 
> I can't because of a bunch of reasons:
> 1) It can be migration, thus I don't have autofs mount on destination
> node at all
> 2) It can be a container, which is stopped after dump (thus mount
> point 
> is destroyed).
> 
> > But probably my suggestion is far to simplistic as I get the
> > impression
> > you have a process already in a given state which you want to
> > restore.
> > 
> > One thing to keep in mind is that if an autofs mount is not set
> > catatonic any access other than the owner process (process group
> > pid)
> > will hang unless there is an actual user space process to service
> > the
> > callback.
> > 
> > Although I don't know the flow of things that might be important at
> > some point.
> > 
> > And if the mount is set catatonic the process needs to set the
> > pipefd
> > to take "ownership" which also clears the mount catatonic flag.
> 
> The migration is already implemented and sent to CRIU mailing list.
> Here is the list, if you are interesting (I use kernel with this
> patch 
> applied):
> https://lists.openvz.org/pipermail/criu/2016-January/024749.html

ok, I'll try and have a look although I'm pressed for time so I'm not
sure I'll spend much time on it.

In any case the project needs to do what it thinks best so my only real
concern is to try and alert you to possible problems.

> 
> > Anyway, let me think about what you've written for a while.
> > Ian
> 
> Sure, take your time.

[toc] | [prev] | [next] | [standalone]


#1306101

FromStanislav Kinsburskiy <skinsbursky@odin.com>
Date2016-01-11 12:40 +0100
Message-ID<qPIZk-1I2-13@gated-at.bofh.it>
In reply to#1305130

09.01.2016 02:31, Ian Kent пишет:
> On Fri, 2016-01-08 at 16:05 +0100, Stanislav Kinsburskiy wrote:
>> 08.01.2016 13:58, Ian Kent пишет:
>>> On Fri, 2016-01-08 at 12:29 +0100, Stanislav Kinsburskiy wrote:
>>>> 08.01.2016 08:20, Ian Kent пишет:
>>>>> On Thu, 2016-01-07 at 16:46 +0100, Stanislav Kinsburskiy wrote:
>>>>>> Good day, gentlemen.
>>>>>>
>>>>>> Could you update, what's the status with this patch?
>>>>>> Without it it's impossible to match process pipe with kernel
>>>>>> pipe,
>>>>>> while
>>>>>> this is "must have" to be able to migrate AutoFS via CRIU.
>>>>> Right, I did mean to reply to this mail but have been
>>>>> distracted by
>>>>> family stuff.
>>>>>
>>>>> I don't know what CRIU is and people looking at changelog
>>>>> entries
>>>>> shouldn't need to do a web search to find out.
>>>>>
>>>>> Could you change it a little.
>>>> Fair enough. I'll resend with more descriptive message.
>>>> But first I would like to clarify to you the problem root and why
>>>> it's
>>>> done like this.
>>>>
>>>>> I'm also not sure whether to forward this (assuming the
>>>>> description
>>>>> is
>>>>> updated a little) to Al or to include it in the series to
>>>>> rename
>>>>> autofs4 to autofs that I'm hoping to ask be included in linux
>>>>> -next
>>>>> fairly soon.
>>>> Here I don't know, what's better. Of course Al can take it as
>>>> well.
>>>> But,
>>>> probably, first would be nice to make sure, that this solution is
>>>> the
>>>> best one.
>>>> Description of the problem is below.
>>>>
>>>>> Passing it on to Al will likely interfere with the series
>>>>> coming
>>>>> from
>>>>> linux-next so that could be bit of a hassle.
>>>>>
>>>>> Another thing I'm wondering about is the order this entry will
>>>>> appear
>>>>> at in the options. You order choice is sensible though and
>>>>> autofs
>>>>> shouldn't have a problem with the inserted option but other
>>>>> applications might.
>>>> I should put it at the end, probably?
>>>>
>>>>> Finally, and perhaps most importantly, I don't get what your
>>>>> trying
>>>>> to
>>>>> do, you also haven't given any clues to that in the patch
>>>>> dscription.
>>>>>
>>>>> IOW how do you expect to use this.
>>>>>
>>>>>> 16.12.2015 13:02, Stanislav Kinsburskiy пишет:
>>>>>>> This is required for CRIU to migrate a mount point, when
>>>>>>> write
>>>>>>> end
>>>>>>> in user
>>>>>>> space is closed.
>>>>> Like I said what does this mean.
>>>>>
>>>>> autofs doesn't need this when it re-constructs a mount tree
>>>>> from
>>>>> existing mounts on re-start or after a SIGKILL on the automount
>>>>> process.
>>>>>
>>>>> How is this different and how will it be used?
>>>>>
>>>>> The question to be answered here is "is this the best way to do
>>>>> it
>>>>> and
>>>>> will it work for the autofs mount types you expect it to"?
>>>> So, here is a brief description of the problem.
>>>> To migrate autofs mount, one have to reconstruct control pipe
>>>> between
>>>> kernel and autofs master.
>>>> There are two cases I'm wiling to support:
>>>> 1) Automount binary (autofs package). This program is very gentle
>>>> and
>>>> it
>>>> doesn't close write end of the pipe after mount.
>>>> 2) Systemd. This program closes write end of the pipe once the
>>>> mount
>>>> is
>>>> done.
>>> I must admit I'm having trouble understanding the description.
>>> Give me a little time with it.
>>>
>>> I don't know how systemd works with autofs mounts only that it uses
>>> the
>>> autofs direct mount type.
>> Systemd closes write end of the pipe after mount.
>>
>>> autofs uses both indirect and direct mounts and both can have
>>> offsets
>>> (from the kernel POV semantically direct mounts). So there is quite
>>> a
>>> bit to worry about beside the kernel pipe.
>> It's not about direct or indirects mounts.
>> It's about process state restore.
>> With CRIU migration, any task is frozen, then disassembled into
>> pieces
>> (dump files), which are used to assemble task exactly in the same
>> state
>> in was before dump.
>> The technology is very complex and uses a lot a different tricky
>> techniques to make this possible in userspace to describe all the
>> details here.
>>
>> But below is a bit more information, which, hopefully, will clarify
>> all
>> this a little bit more.
>> One of a process attributed to migrate is "opened files". Pipes also
>> belong to this attribute.
>>
>> To restore a pipe CRIU does the following (a very simplified
>> description):
>> 1) Creates a new pipe.
>> 2) Writes (previously stores in images) its contents via write end.
>> 3) Duplicate pipe descriptors to the fds of the process, which were
>> used
>> before dump, if required
>> 4) Send pipe descriptors to other processes, sharing it, via unix
>> socket.
>> 5) Close those pipe descriptors, which are not required (say, this
>> process had only read end, while it's child had write end).
>>
>> Thus in case of restoring and autofs mount of systemd (which, for
>> example, closed write end and has read end on fd 40), one have to
>> create
>> a pipe (say, appeared with fd 5 and fd 6), fill it with content via
>> fd
>> 6, duplicate fd 5 into fd 40, call mount with pipe fd 6 and then
>> close fd 6.
>> This is, yet again, a very simple explanation.
> Right, as said initially (more or less), if you need the patch you
> posted then it shouldn't cause a problem so it should be ok. Al hasn't
> responded so I guess that means I should go the linux-next path
> possibly via pull request for the series I have to rename autofs4 to
> autofs (along with this one, to prevent merge conflicts).
>
> I haven't asked Steven about this yet so I'm not sure if a pull request
> is even the right thing to do.
>
> There is another case I was wondering about.
>
> That's when there is a direct mount that is covered by a real mount.
>
> autofs will have a file handle open to it (on the underlying mount
> point path) to use for control purposes like expires. I think you also
> need to restore those file handles to restore process state and in this
> case the mount point is covered.
>

This is covered: all the mount points first mounted somewhere to be able 
to reopen files. Then mount order is restored.

>>> Anyway, it seems your only concern is the kernel pipe and I wonder
>>> why
>>> you can't just set the mount catatonic (in autofs speak) on save
>>> and
>>> open a new kernel pipe then set the pipefd on the autofs mount on
>>> restore.
>> I can't because of a bunch of reasons:
>> 1) It can be migration, thus I don't have autofs mount on destination
>> node at all
>> 2) It can be a container, which is stopped after dump (thus mount
>> point
>> is destroyed).
>>
>>> But probably my suggestion is far to simplistic as I get the
>>> impression
>>> you have a process already in a given state which you want to
>>> restore.
>>>
>>> One thing to keep in mind is that if an autofs mount is not set
>>> catatonic any access other than the owner process (process group
>>> pid)
>>> will hang unless there is an actual user space process to service
>>> the
>>> callback.
>>>
>>> Although I don't know the flow of things that might be important at
>>> some point.
>>>
>>> And if the mount is set catatonic the process needs to set the
>>> pipefd
>>> to take "ownership" which also clears the mount catatonic flag.
>> The migration is already implemented and sent to CRIU mailing list.
>> Here is the list, if you are interesting (I use kernel with this
>> patch
>> applied):
>> https://lists.openvz.org/pipermail/criu/2016-January/024749.html
> ok, I'll try and have a look although I'm pressed for time so I'm not
> sure I'll spend much time on it.
>
> In any case the project needs to do what it thinks best so my only real
> concern is to try and alert you to possible problems.

Thanks for the alerts.
Should I move this option to the end of the list to preserve the sequence?

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web