Path: csiph.com!eternal-september.org!feeder.eternal-september.org!aioe.org!bofh.it!news.nic.it!robomod From: Dashi DS1 Cao Newsgroups: linux.kernel Subject: RE: work queue of scsi fc transports should be serialized Date: Sat, 20 May 2017 10:30:01 +0200 Message-ID: References: X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFupkleJIrShJLcpLzFFi42LJePGQSXfaX/l Igx8vTCwO/mxjtLi8aw6bRff1HWwOzB7T1pxn8vi8SS6AKYo1My8pvyKBNePS/HNMBd3yFSsP pjcw/hLrYuTiEBJ4wihxac1MdghnIaPE/Ff/gBwODjYBdYnfJ/i6GDk5RARyJNasaGYFsZkFH CVu733LBGILCzhJrLy8nh2ixlni1ccXjBC2kcTMFe/AbBYBVYm9d5ewgNi8Aj4SP9+fYQaxhQ SKJW6e3QhWwwlU/7TjNpjNKCArMe3RfSaIXeISc6fNAtsrISAgsWTPeWYIW1Ti5eN/rCBnSgj IS2yZJQhRridxY+oUNghbW2LZwtfMEGsFJU7OfMIygVFkFpKps5C0zELSMgtJywJGllWMGsWp RWWpRbqGZnpJRZnpGSW5iZk5uoYGFnq5qcXFiempOYlJxXrJ+bmbGIHxUs/AwLiDccJu90OMk hxMSqK89Q3ykUJ8SfkplRmJxRnxRaU5qcWHGNU5OAS+796QJMWSl5+XqiTBq/MHqEywKDU9tS ItMwcYzzCVEhw8SiK8oSBp3uKCxNzizHSI1ClGRSlx3gSQhABIIqM0D64NlkIuMcpKCfMyMjA wCPEUpBblZpagyr9iFOdgVBLmlQGZwpOZVwI3/RXQYiagxdbPwBaXJCKkpBoYpfwr/Di9nBKu e6V43Hhg5c0RxmQi7nzv2ZbfissemqqJffH9H6Re9izG4vyStTalZ/++kVd9tmz6tnNpqvX+H iF5zvF1h4L1/vXsP8i70fGe8jnuU3OSFCO0AgtyIzSTTfUv3ghrseRuPv5ScF1s/Aami9c3TF A1nSW+eurU6D33b4o7pEgrsRRnJBpqMRcVJwIAQ938UhwDAAA= X-Env-Sender: caods1@lenovo.com X-Msg-Ref: server-5.tower-144.messagelabs.com!1495268758!16726325!1 X-Originating-IP: [104.232.225.2] X-Starscan-Version: 9.4.12; banners=-,-,- X-Viruschecked: Checked Thread-Topic: work queue of scsi fc transports should be serialized Thread-Index: AQHS0O/i4QRdvq/650a75t6/OTGIsKH836bA Accept-Language: zh-CN, en-US Content-Language: zh-CN X-Originating-IP: [10.96.19.89] Content-Type: text/plain; charset="iso-8859-1" Content-Transfer-Encoding: 8BIT MIME-Version: 1.0 Sender: robomod@news.nic.it List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Approved: robomod@news.nic.it Lines: 88 Organization: linux.* mail to news gateway X-Original-Cc: "linux-kernel@vger.kernel.org" X-Original-Date: Sat, 20 May 2017 08:25:09 +0000 X-Original-Message-ID: <23B7B563BA4E9446B962B142C86EF24A088AF8B0@CNMAILEX03.lenovo.com> X-Original-References: <23B7B563BA4E9446B962B142C86EF24A088AE2FB@CNMAILEX03.lenovo.com> <1495233163.2581.5.camel@sandisk.com> X-Original-Sender: linux-kernel-owner@vger.kernel.org Xref: csiph.com linux.kernel:1646081 On Fri, 2017-05-19 at 09:36 +0000, Dashi DS1 Cao wrote: > It seems there is a race of multiple "fc_starget_delete" of the same > rport, thus of the same SCSI host. The race leads to the race of > scsi_remove_target and it cannot be prevented by the code snippet > alone, even of the most recent > version: > spin_lock_irqsave(shost->host_lock, flags); > list_for_each_entry(starget, &shost->__targets, siblings) { > if (starget->state == STARGET_DEL || > starget->state == STARGET_REMOVE) > continue; > If there is a possibility that the starget is under deletion(state == > STARGET_DEL), it should be possible that list_next_entry(starget, > siblings) could cause a read access violation. >Hello Dashi, >Something else must be going on. From scsi_remove_target(): >restart: > spin_lock_irqsave(shost->host_lock, flags); > list_for_each_entry(starget, &shost->__targets, siblings) { > if (starget->state == STARGET_DEL || >     starget->state == STARGET_REMOVE) > continue; > if (starget->dev.parent == dev || &starget->dev == dev) { > kref_get(&starget->reap_ref); > starget->state = STARGET_REMOVE; > spin_unlock_irqrestore(shost->host_lock, flags); > __scsi_remove_target(starget); > scsi_target_reap(starget); > goto restart; > } > } > spin_unlock_irqrestore(shost->host_lock, flags); >In other words, before scsi_remove_target() decides to call __scsi_remove_target(), it changes the target state into STARGET_REMOVE while holding the host lock. >This means that scsi_remove_target() won't call __scsi_remove_target() twice and also that it won't invoke list_next_entry(starget, siblings) after starget has been >freed. >Bart. In the crashes of Suse 12 sp1, the root cause is the deletion of a list node without holding the lock: spin_lock_irqsave(shost->host_lock, flags); list_for_each_entry_safe(starget, tmp, &shost->__targets, siblings) { if (starget->state == STARGET_DEL) continue; if (starget->dev.parent == dev || &starget->dev == dev) { /* assuming new targets arrive at the end */ kref_get(&starget->reap_ref); spin_unlock_irqrestore(shost->host_lock, flags); __scsi_remove_target(starget); list_move_tail(&starget->siblings, &reap_list); --this deletion from shost->__targets list is done without the lock. spin_lock_irqsave(shost->host_lock, flags); } } spin_unlock_irqrestore(shost->host_lock, flags); A better solution is as follows, without introducing more states: restart: spin_lock_irqsave(shost->host_lock, flags); list_for_each_entry_safe(starget, tmp, &shost->__targets, siblings) { if (starget->dev.parent == dev || &starget->dev == dev) { /* assuming new targets arrive at the end */ kref_get(&starget->reap_ref); list_move_tail(&starget->siblings, &reap_list); spin_unlock_irqrestore(shost->host_lock, flags); __scsi_remove_target(starget); goto restart; } } spin_unlock_irqrestore(shost->host_lock, flags); list_for_each_entry_safe(starget, tmp, &reap_list, siblings) scsi_target_reap(starget); Another place that should be modified is the scsi_transport_fc.c: From: if (rport->scsi_target_id != -1) fc_starget_delete(&rport->stgt_delete_work); To: if (rport->scsi_target_id != -1) { fc_flush_work(shost); BUG_ON(ACCESS_ONCE(rport->scsi_target_id) != -1); } Regards, Dashi Cao