首页 /研究 /Efficient video object segmentation based on frame-wise and segment-wise spatio-temporal interaction memory networks
OTHER

Efficient video object segmentation based on frame-wise and segment-wise spatio-temporal interaction memory networks

Jisheng Dang, Huicheng ZHENG, Bimei Wang, Juncheng Li, Jianhuang Lai

发表年份
2024
引用次数
3

摘要

Video object segmentation aims to automatically segment objects of interest in videos, with wide applications in areas such as video editing, robot navigation, and autonomous driving. Existing methods for video object segmentation mostly rely on independent-frame appearance memory, which often falls short when dealing with complex video scenes with severe occlusions or appearance similarities. To address these challenges, this paper proposes a VOS method based on frame-wise and segment-wise spatio-temporal interaction memory (FSSTIM). FSSTIM introduces frame-wise and segment-wise spatio-temporal interaction memory construction blocks, which extract segment-wise spatio-temporal memory feature maps by constructing spatio-temporal context graph networks and enhance them by interacting with frame-wise memory feature maps, significantly improving the network's ability to handle similar appearances and object occlusions. Furthermore, the introduction of dynamic sampling memory readers achieves efficient multi-granularity historical information retrieval, speeding up inference and improving segmentation accuracy. Experiments on popular VOS datasets such as DAVIS, YouTube-VOS, and MOSE demonstrate that the proposed method achieves state-of-the-art performance while maintaining real-time processing speed and strong generalization capability.

关键词

Computer science

相关论文

查看 OTHER 分类全部论文