首页 /研究 /Learning Freehand Ultrasound Through Multimodal Representation and Skill Adaptation
MANIPULATION

Learning Freehand Ultrasound Through Multimodal Representation and Skill Adaptation

Xutian Deng, Junnan Jiang, Wen Cheng, Chenguang Yang, Miao Li

发表年份
2024
引用次数
5

摘要

With medical ultrasound becoming one of the most prevalent examination methods, robotic ultrasound systems offer the potential to simplify the scanning process and relieve professional sonographers from repetitive and tedious tasks. Despite recent advances, enabling robots to autonomously perform ultrasound examinations remains a challenge, mainly due to the difficulty in representing and generalizing professional ultrasound skills. In this paper, we present a comprehensive framework for learning autonomous ultrasound skills from freehand demonstrations in clinical settings. Our proposed framework consists of two key stages: offline learning and online adaptation. During the offline learning stage, ultrasound skills are encapsulated into a low-dimensional probabilistic model using a self-supervised architecture. The multimodal signals include ultrasound images, probe orientations, and contact forces. During the online adaptation stage, the model predicts the optimal actions either by direct regression or by using local exploration schemes. We perform clinical demonstrations with 24 volunteers and collect 120 experiences. Our benchmark includes 5 different tasks, including intra-patient, inter-patient, inter-sex, inter-age, and inter-obesity tasks. Both one-step and sequence-based predictions are achieved by using different variants of our framework. Customized and generic representation learning backbones are tested and analyzed. In conclusion, our autonomous ultrasound framework is flexible and robust, and potentially enriches the options for freehand/robotic ultrasound applications. Note to Practitioners—This paper is motivated by the problem of learning multimodal manipulation skills from human demonstrations, with a specific focus on freehand ultrasound skills. Our multimodal fusion framework is effective and compatible with some popular image representation backbones. Our adaptive methods and the variants have satisfactory prediction accuracy, with flexibility achieved by adjusting a few factors. We collect high-quality freehand demonstrations from ultrasound examinations in clinical settings. The data is openly available to facilitate the reproducibility of our work and to support the development of autonomous ultrasound strategies based on imitation learning.

关键词

Computer scienceRepresentation (politics)Adaptation (eye)Human–computer interactionComputer visionArtificial intelligenceUltrasonic imagingUltrasoundPsychologyAcoustics

相关论文

查看 MANIPULATION 分类全部论文