Published August 21, 2020 | Version v1
Journal article

Multi-level feature aggregation network for instrument identification of endoscopic images

  • 1. Beijing Engineering Research Center of Mixed Reality and Advanced Display, School of Optics and Photonics, Beijing Institute of Technology, Beijing 100081 (China)
  • 2. School of Computer Science Technology, Beijing Institute of Technology, Beijing 100081 (China)

Description

Identification of surgical instruments is crucial in understanding surgical scenarios and providing an assistive process in endoscopic image-guided surgery. This study proposes a novel multilevel feature-aggregated deep convolutional neural network (MLFA-Net) for identifying surgical instruments in endoscopic images. First, a global feature augmentation layer is created on the top layer of the backbone to improve the localization ability of object identification by boosting the high-level semantic information to the feature flow network. Second, a modified interaction path of cross-channel features is proposed to increase the nonlinear combination of features in the same level and improve the efficiency of information propagation. Third, a multiview fusion branch of features is built to aggregate the location-sensitive information of the same level in different views, increase the information diversity of features, and enhance the localization ability of objects. By utilizing the latent information, the proposed network of multilevel feature aggregation can accomplish multitask instrument identification with a single network. Three tasks are handled by the proposed network, including object detection, which classifies the type of instrument and locates its border; mask segmentation, which detects the instrument shape; and pose estimation, which detects the keypoint of instrument parts. The experiments are performed on laparoscopic images from MICCAI 2017 Endoscopic Vision Challenge, and the mean average precision (AP) and average recall (AR) are utilized to quantify the segmentation and pose estimation results. For the bounding box regression, the AP and AR are 79.1% and 63.2%, respectively, while the AP and AR of mask segmentation are 78.1% and 62.1%, and the AP and AR of the pose estimation achieve 67.1% and 55.7%, respectively. The experiments demonstrate that our method efficiently improves the recognition accuracy of the instrument in endoscopic images, and outperforms the other state-of-the-art methods. (paper)

Availability note (English)

Available from http://dx.doi.org/10.1088/1361-6560/ab8dda

Additional details

Identifiers

Publishing Information

Journal Title
Physics in Medicine and Biology
Journal Volume
65
Journal Issue
16
Journal Page Range
[15 p.]
ISSN
0031-9155
CODEN
PHMBA7

INIS

Country of Publication
United Kingdom
Country of Input or Organization
International Atomic Energy Agency (IAEA)
INIS RN
52074205
Subject category
S62: RADIOLOGY AND NUCLEAR MEDICINE;
Descriptors DEI
ACCURACY; BIOMEDICAL RADIOGRAPHY; IMAGES; NEURAL NETWORKS; SURGERY
Descriptors DEC
DIAGNOSTIC TECHNIQUES; MEDICINE; NUCLEAR MEDICINE; RADIOLOGY