首页 /研究 /Learning incipient slip with GelSight sensors: Attention Classification with Video Vision Transformers
MANIPULATION

Learning incipient slip with GelSight sensors: Attention Classification with Video Vision Transformers

Amit Parag, Edward H. Adelson, Ekrem Misimi

发表年份
2024
引用次数
3

摘要

An important aspect of robotic grasping is the ability to detect incipient slip based on real-time information through tactile sensors. In this paper, we propose to use Video Vision Transformers to detect the onset of slip in grasping scenarios. The dynamic nature of slip makes Video Vision Transformers well-suited for capturing temporal correlations with relatively small datasets. The training data is acquired through two GelSight tactile sensors attached to the generic finger grippers of a Panda Franka Emika robot arm that grasps, lifts and shakes 30 everyday objects in order to induce slip. We further conducted an ablation study by considering 5, 4, 3, and 2 frames prior to slip onset, revealing consistent prediction accuracy. Our approach demonstrates the capability to predict slips well in advance, even up to the 5<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">th</sup> frame before the onset. This underscores the predictive capability of our approach, indicating its effectiveness in slip detection well before of its occurrence. This advance prediction capability may be a valuable tool for undertaking preemptive corrective actions, such as implementing a more secure gripper closure. We evaluate the efficiency of our approach to predict onset of slip on 10 previously-unseen objects and achieve a zero-shot mean prediction accuracy of 99%.

关键词

Computer scienceArtificial intelligenceComputer visionTransformerSlip (aerodynamics)EngineeringElectrical engineeringVoltageAerospace engineering

相关论文

查看 MANIPULATION 分类全部论文