TY - GEN
T1 - Hierarchical Multi-Array Tactile Representation Learning and Diffusion Policy for Real-World Dexterous Placement Tasks
AU - Yang, Jiaqi
AU - Peng, Gang
AU - Wang, Chaoze
AU - Cong, Mingjun
AU - Li, Chuangye
AU - Yang, Bingchuan
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - In contact-rich dexterous manipulation tasks, tactile feedback is critical for achieving high-precision alignment and stable physical interaction. However, signals from commonly used distributed tactile sensors are often high-dimensional and sparse, with scattered and non-uniform spatial layouts, which makes it challenging to learn generalizable tactile representations from raw measurements and to leverage them effectively for policy learning. To address these challenges, we propose a hierarchical tactile representation learning framework for real-world dexterous placement. In the local pretraining stage, we reweight zero and non-zero samples in the reconstruction loss to encourage the model to focus on physically meaningful contact regions. In the global representation stage, we further develop a multi-array tactile modeling network that integrates hand-structure priors and incorporates sensor pose information to enable structured cross-array alignment and aggregation, thereby learning stable and consistent global tactile features. For policy learning, visual observations and hierarchical tactile representations are jointly used as conditioning inputs, and a diffusion policy is adopted to model the action distribution. Experimental results demonstrate that the proposed hierarchical tactile representation learning substantially improves the success rate and robustness of policy learning, while simultaneously reducing the number of closedloop action steps required to complete a placement episode, thus enhancing overall execution efficiency.
AB - In contact-rich dexterous manipulation tasks, tactile feedback is critical for achieving high-precision alignment and stable physical interaction. However, signals from commonly used distributed tactile sensors are often high-dimensional and sparse, with scattered and non-uniform spatial layouts, which makes it challenging to learn generalizable tactile representations from raw measurements and to leverage them effectively for policy learning. To address these challenges, we propose a hierarchical tactile representation learning framework for real-world dexterous placement. In the local pretraining stage, we reweight zero and non-zero samples in the reconstruction loss to encourage the model to focus on physically meaningful contact regions. In the global representation stage, we further develop a multi-array tactile modeling network that integrates hand-structure priors and incorporates sensor pose information to enable structured cross-array alignment and aggregation, thereby learning stable and consistent global tactile features. For policy learning, visual observations and hierarchical tactile representations are jointly used as conditioning inputs, and a diffusion policy is adopted to model the action distribution. Experimental results demonstrate that the proposed hierarchical tactile representation learning substantially improves the success rate and robustness of policy learning, while simultaneously reducing the number of closedloop action steps required to complete a placement episode, thus enhancing overall execution efficiency.
KW - Diffusion Policy
KW - Tactile Representation Learning
KW - Vision-tactile Fusion
UR - https://www.scopus.com/pages/publications/105043949193
U2 - 10.1109/CCDC69976.2026.11559719
DO - 10.1109/CCDC69976.2026.11559719
M3 - 会议稿件
AN - SCOPUS:105043949193
T3 - 38th Chinese Control and Decision Conference, CCDC 2026
SP - 5434
EP - 5439
BT - 38th Chinese Control and Decision Conference, CCDC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 38th Chinese Control and Decision Conference, CCDC 2026
Y2 - 15 May 2026 through 18 May 2026
ER -