Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1444345

[PATCH v17 net-next 0/1] introduce Hyper-V VM Sockets(hv_sock)

Path csiph.com!news.mixmin.net!aioe.org!bofh.it!news.nic.it!robomod
From Dexuan Cui <decui@microsoft.com>
Newsgroups linux.kernel
Subject [PATCH v17 net-next 0/1] introduce Hyper-V VM Sockets(hv_sock)
Date Fri, 15 Jul 2016 16:30:02 +0200
Message-ID <rVchQ-7ZP-9@gated-at.bofh.it> (permalink)
X-Original-To "davem@davemloft.net" <davem@davemloft.net>, "gregkh@linuxfoundation.org" <gregkh@linuxfoundation.org>, "netdev@vger.kernel.org" <netdev@vger.kernel.org>, "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>, "devel@linuxdriverproject.org" <devel@linuxdriverproject.org>, "olaf@aepfle.de" <olaf@aepfle.de>, "apw@canonical.com" <apw@canonical.com>, "jasowang@redhat.com" <jasowang@redhat.com>, Vitaly Kuznetsov <vkuznets@redhat.com>, Cathy Avery <cavery@redhat.com>, KY Srinivasan <kys@microsoft.com>, "joe@perches.com" <joe@perches.com>, Rolf Neugebauer <rolf.neugebauer@docker.com>, "Michal Kubecek" <mkubecek@suse.cz>
Dkim-Signature v=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=selector1; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version; bh=MAZdVA0VrcltlNYzIy/pXkSjqk0phxxByitPhc12JQg=; b=LbXTDyYP/9bFwAXk0GQwbZ6KuyEqHik7Silb+gyF7OoB+tDsoGjDpEvp7PcKYTFROySK0w4aJ7/Was1LK50Cj9a0MDLtei5BXupDlNzkOJV8QnZTmQpqmMGpvveZf7WdiSRa4SzYGyjAvzcZRGNyGoLB0twSPqI6w0xVEKi5xXE=
Thread-Topic [PATCH v17 net-next 0/1] introduce Hyper-V VM Sockets(hv_sock)
Thread-Index AdHepRgCvbD0kur+T/uKQ8orLSFoFg==
Accept-Language en-US
Content-Language en-US
Authentication-Results spf=none (sender IP is ) smtp.mailfrom=decui@microsoft.com;
X-Originating-IP [2404:f801:9000:19::4cc]
X-Ms-Office365-Filtering-Correlation-ID 01a9bf1f-39a6-433b-4e35-08d3acbc66f8
X-Microsoft-Exchange-Diagnostics 1;CO2PR03MB2279;6:ZOuGatdZ+mK/PR1MNkwd0uVEKF6+rSRnN63qPfeiR+utabJAo4Jl6umvsG5D9TnMXlb7CVKfS356awozY1bRNTYL4U/0m6880c7bF1kcwedljhddNE2yVDn98OIeGElyTXfmETTR0stcN3vJON0DMkBOLTWJnYmyeodmAkKVtFZwqcSRy0ZdESrHBkfNcX+uLLm2HxBkjz2I0dm9Fz5EDsJcX9fAT2vaOpSr4IWLGlVkGs6NDb0aWm8bLhblMEcsZ0i+6FV5xX43T5aOREohnpM5ZjEAZhfvBScSFbvSe9x7Y4a/ky9RmbG3umceH6vY5J1nYT+CEc7wni4f1GXK+g==;5:poxInySJBzGMfRX5VIyodxt3zQdpsGyQgQ4U9JrYkLr8pHlGPf/AY36B/YKALW9vBYNsHKIXIwkMoScLXxkpCLvFacYih5uiPo9IIY5zmD4mrr9OO7I2f3CyqiE7Pq9THLPg2sXV4pf958aY/12dbQ==;24:UWliPVbltQ3MT9Opc4SSauNquhLEqVJtu6OB1Xza1JPWiXNPZljxWAFUe76a+DKYAfXYcubotVHMBINAG+GiD+SYvfvyeB9sSeZG57I5KZk=;7:KphsRuz6eD2qOD+d0sJM6WT4DDM0q5JO+j0JeLjQEhagl1CyLGQS2JHWP1zHMIFco6CDdNn34gKWR87ct5vNrVYMWZTb13vZQMTK/9k1/zh9I/4LCs2Seoo8yW3bpfpP6too14blENhWYcGSsDbbG5/FdaIMdnCZ1i4rX9RsfFrhRKUl0G6G8F8t9E3BLQ/rpcMiB3QqG14pKJrZ9AYPGhuleoFhW168TgU8mUSZT1kZg6gd/rG5ufmwDM74pvkO
X-Microsoft-Antispam UriScan:;BCL:0;PCL:0;RULEID:;SRVR:CO2PR03MB2279;
X-Microsoft-Antispam-Prvs <CO2PR03MB2279ECE65C99D62A698825C7BF330@CO2PR03MB2279.namprd03.prod.outlook.com>
X-Exchange-Antispam-Report-Test UriScan:(158342451672863)(278428928389397)(166708455590820)(36064498253994)(788757137089)(176295241369792)(21532816269658);
X-Exchange-Antispam-Report-Cfa-Test BCL:0;PCL:0;RULEID:(61425038)(601004)(2401047)(8121501046)(5005006)(3002001)(10201501046)(6055026)(61426038)(61427038);SRVR:CO2PR03MB2279;BCL:0;PCL:0;RULEID:;SRVR:CO2PR03MB2279;
X-Forefront-Prvs 00046D390F
X-Forefront-Antispam-Report SFV:NSPM;SFS:(10019020)(6009001)(7916002)(209900001)(45984002)(189002)(3905003)(199003)(122556002)(81156014)(33656002)(6116002)(7736002)(2501003)(8936002)(81166006)(9686002)(97736004)(2421001)(68736007)(105586002)(3660700001)(92566002)(7846002)(2906002)(77096005)(7696003)(8676002)(4326007)(101416001)(229853001)(5003600100003)(15395725005)(305945005)(2900100001)(102836003)(5001770100001)(5002640100001)(3280700002)(87936001)(54356999)(19580395003)(50986999)(106356001)(86362001)(74316002)(8990500004)(11100500001)(586003)(5005710100001)(15975445007)(99286002)(1511001)(10090500001)(2561002)(10290500002)(86612001)(2201001)(189998001)(10400500002)(76576001)(921003)(3826002)(1121003)(6606295002);DIR:OUT;SFP:1102;SCL:1;SRVR:CO2PR03MB2279;H:CO2PR03MB2182.namprd03.prod.outlook.com;FPR:;SPF:None;PTR:InfoNoRecords;MX:1;A:1;LANG:en;
Received-Spf None (protection.outlook.com: microsoft.com does not designate permitted sender hosts)
Spamdiagnosticoutput 1:99
Spamdiagnosticmetadata NSPM
Content-Type text/plain; charset="us-ascii"
Content-Transfer-Encoding 8BIT
MIME-Version 1.0
X-Originatororg microsoft.com
X-Ms-Exchange-Crosstenant-Originalarrivaltime 15 Jul 2016 14:29:16.9879 (UTC)
X-Ms-Exchange-Crosstenant-Fromentityheader Hosted
X-Ms-Exchange-Crosstenant-ID 72f988bf-86f1-41af-91ab-2d7cd011db47
X-Ms-Exchange-Transport-Crosstenantheadersstamped CO2PR03MB2279
Sender robomod@news.nic.it
List-ID <linux-kernel.vger.kernel.org>
X-Mailing-List linux-kernel@vger.kernel.org
Approved robomod@news.nic.it
Lines 211
Organization linux.* mail to news gateway
X-Original-Cc Haiyang Zhang <haiyangz@microsoft.com>, Dave Scott <dave.scott@docker.com>
X-Original-Date Fri, 15 Jul 2016 14:29:16 +0000
X-Original-Message-ID <CO2PR03MB21829C45F7096C2C0373871CBF330@CO2PR03MB2182.namprd03.prod.outlook.com>
X-Original-Sender linux-kernel-owner@vger.kernel.org
Xref csiph.com linux.kernel:1444345

Show key headers only | View raw


Hyper-V Sockets (hv_sock) supplies a byte-stream based communication
mechanism between the host and the guest. It's somewhat like TCP over
VMBus, but the transportation layer (VMBus) is much simpler than IP.

With Hyper-V Sockets, applications between the host and the guest can talk
to each other directly by the traditional BSD-style socket APIs.

Hyper-V Sockets is only available on new Windows hosts, like Windows Server
2016. More info is in this article "Make your own integration services":
https://msdn.microsoft.com/en-us/virtualization/hyperv_on_windows/develop/make_mgmt_service

The patch implements the necessary support in the guest side by
introducing a new socket address family AF_HYPERV.

You can also get the patch by (commit fcf045af6):
https://github.com/dcui/linux/tree/decui/hv_sock/net-next/20160715_v17

Note: the VMBus driver side's supporting patches have been in the mainline
tree.

I know the kernel has already had a VM Sockets driver (AF_VSOCK) based
on VMware VMCI (net/vmw_vsock/, drivers/misc/vmw_vmci), and KVM is
proposing AF_VSOCK of virtio version:
http://marc.info/?l=linux-netdev&m=145952064004765&w=2

However, though Hyper-V Sockets may seem conceptually similar to
AF_VOSCK, there are differences in the transportation layer, and IMO these
make the direct code reusing impractical:

1. In AF_VSOCK, the endpoint type is: <u32 ContextID, u32 Port>, but in
AF_HYPERV, the endpoint type is: <GUID VM_ID, GUID ServiceID>. Here GUID
is 128-bit.

2. AF_VSOCK supports SOCK_DGRAM, while AF_HYPERV doesn't.

3. AF_VSOCK supports some special sock opts, like SO_VM_SOCKETS_BUFFER_SIZE,
SO_VM_SOCKETS_BUFFER_MIN/MAX_SIZE and SO_VM_SOCKETS_CONNECT_TIMEOUT.
These are meaningless to AF_HYPERV.

4. Some AF_VSOCK's VMCI transportation ops are meanless to AF_HYPERV/VMBus,
like .notify_recv_init
.notify_recv_pre_block
.notify_recv_pre_dequeue
.notify_recv_post_dequeue
.notify_send_init
.notify_send_pre_block
.notify_send_pre_enqueue
.notify_send_post_enqueue
etc.

So I think we'd better introduce a new address family: AF_HYPERV.

Please review the patch.

Looking forward to your comments, especially comments from David. :-)

Changes since v1:
- updated "[PATCH 6/7] hvsock: introduce Hyper-V VM Sockets feature"
- added __init and __exit for the module init/exit functions
- net/hv_sock/Kconfig: "default m" -> "default m if HYPERV"
- MODULE_LICENSE: "Dual MIT/GPL" -> "Dual BSD/GPL"

Changes since v2:
- fixed various coding issue pointed out by David Miller
- fixed indentation issues
- removed pr_debug in net/hv_sock/af_hvsock.c
- used reverse-Chrismas-tree style for local variables.
- EXPORT_SYMBOL -> EXPORT_SYMBOL_GPL

Changes since v3:
- fixed a few coding issue pointed by Vitaly Kuznetsov and Dan Carpenter
- fixed the ret value in vmbus_recvpacket_hvsock on error
- fixed the style of multi-line comment: vmbus_get_hvsock_rw_status()

Changes since v4 (https://lkml.org/lkml/2015/7/28/404):
- addressed all the comments about V4.
- treat the hvsock offers/channels as special VMBus devices
- add a mechanism to pass hvsock events to the hvsock driver
- fixed some corner cases with proper locking when a connection is closed
- rebased to the latest Greg's tree

Changes since v5 (https://lkml.org/lkml/2015/12/24/103):
- addressed the coding style issues (Vitaly Kuznetsov & David Miller, thanks!)
- used a better coding for the per-channel rescind callback (Thank Vitaly!)
- avoided the introduction of new VMBUS driver APIs vmbus_sendpacket_hvsock()
and vmbus_recvpacket_hvsock() and used vmbus_sendpacket()/vmbus_recvpacket()
in the higher level (i.e., the vmsock driver). Thank Vitaly!

Changes since v6 (http://lkml.iu.edu/hypermail/linux/kernel/1601.3/01813.html)
- only a few minor changes of coding style and comments

Changes since v7
- a few minor changes of coding style: thanks, Joe Perches!
- added some lines of comments about GUID/UUID before the struct sockaddr_hv.

Changes since v8
- removed the unnecessary __packed for some definitions: thanks, David!
- hvsock_open_connection:  use offer.u.pipe.user_def[0] to know the connection
and reorganized the function
direction 
- reorganized the code according to suggestions from Cathy Avery: split big
functions into small ones, set .setsockopt and getsockopt to
sock_no_setsockopt/sock_no_getsockopt
- inline'd some small list helper functions

Changes since v9
- minimized struct hvsock_sock by making the send/recv buffers pointers.
   the buffers are allocated by kmalloc() in __hvsock_create() now.
- minimized the sizes of the send/recv buffers and the vmbus ringbuffers.

Changes since v10

1) add module params: send_ring_page, recv_ring_page. They can be used to
enlarge the ringbuffer size to get better performance, e.g.,
# modprobe hv_sock  recv_ring_page=16 send_ring_page=16
By default, recv_ring_page is 3 and send_ring_page is 2.

2) add module param max_socket_number (the default is 1024).
A user can enlarge the number to create more than 1024 hv_sock sockets.
By default, 1024 sockets take about 1024 * (3+2+1+1) * 4KB = 28M bytes.
(Here 1+1 means 1 page for send/recv buffers per connection, respectively.)

3) implement the TODO in hvsock_shutdown().

4) fix a bug in hvsock_close_connection():
   I remove "sk->sk_socket->state = SS_UNCONNECTED;" -- actually this line
is not really useful. For a connection triggered by a host app's connect(),
sk->sk_socket remains NULL before the connection is accepted by the server
app (in Linux VM): see hvsock_accept() -> hvsock_accept_wait() ->
sock_graft(connected, newsock). If the host app exits before the server
app's accept() returns, the host can send a rescind-message to close the
connection and later in the Linux VM's message handler 
i.e. vmbus_onoffer_rescind()), Linux will get a NULL de-referencing crash. 

5) fix a bug in hvsock_open_connection()
  I move the vmbus_set_chn_rescind_callback() to a later place, because
when vmbus_open() fails, hvsock_close_connection() can do nothing and we
count on vmbus_onoffer_rescind() -> vmbus_device_unregister() to clean up
the device.

6) some stylistic modificiation.


Changes since v11:
1) remove the module params as David suggested.

2) use 5 exact pages for VMBus send/recv rings, respectively.
The host side's design of the feature requires 5 exact pages for recv/send
rings respectively -- this is suboptimal considering memory consumption,
however unluckily we have to live with it, before the host comes up with
a new design in the future. :-(

3) remove the per-connection static send/recv buffers
Instead, we allocate and free the buffers dynamically only when we recv/send
data. This means: when a connection is idle, no memory is consumed as
recv/send buffers at all.
Dexuan Cui (1):
  hv_sock: introduce Hyper-V Sockets

Changes since v12:
return ENOMEM on buffer alllocation failure

   Actually "man read/write" says "Other errors may occur, depending on the
object connected to fd". "man send/recv" indeed lists ENOMEM.
   Considering AF_HYPERV is a new socket type, ENOMEM seems OK here.
   In the long run, I think we should add a new API in the VMBus driver,
allowing data copy from VMBus ringbuffer into user mode buffer directly.
This way, we can even eliminate this temporary buffer.

Changes since v13:
fix some coding style issues pointed out by David.

Changes since v14:
Just some stylistic changes addressing comments from Joe Perches and
Olaf Hering -- thank you!
- add a GPL blurb.
- define a new macro PAGE_SIZE_4K and use it to replace PAGE_SIZE
- change sk_to_hvsock/hvsock_to_sk() from macros to inline functions
- remove a not-very-useful pr_err()
- fix some typos in comment and coding style issues.

Changes since v15:
Made stylistic changes addressing comments from Vitaly Kuznetsov.
Thank you very much for the detailed comments, Vitaly! 

Changes since v16:
- PAGE_SIZE -> PAGE_SIZE_4K
- allow regular users to use the socket
Thank you Michal Kubecek for the suggestions!

Dexuan Cui (1):
  hv_sock: introduce Hyper-V Sockets

 MAINTAINERS                 |    2 +
 include/linux/hyperv.h      |   13 +
 include/linux/socket.h      |    4 +-
 include/net/af_hvsock.h     |   78 +++
 include/uapi/linux/hyperv.h |   23 +
 net/Kconfig                 |    1 +
 net/Makefile                |    1 +
 net/hv_sock/Kconfig         |   10 +
 net/hv_sock/Makefile        |    3 +
 net/hv_sock/af_hvsock.c     | 1507 +++++++++++++++++++++++++++++++++++++++++++
 10 files changed, 1641 insertions(+), 1 deletion(-)
 create mode 100644 include/net/af_hvsock.h
 create mode 100644 net/hv_sock/Kconfig
 create mode 100644 net/hv_sock/Makefile
 create mode 100644 net/hv_sock/af_hvsock.c

-- 
2.1.0

Back to linux.kernel | Previous | Next | Find similar | Unroll thread


Thread

[PATCH v17 net-next 0/1] introduce Hyper-V VM Sockets(hv_sock) Dexuan Cui <decui@microsoft.com> - 2016-07-15 16:30 +0200

csiph-web