A Unified and Biologically Plausible Relational Graph Representation of Vision Transformers

Yuzhong Chen; Zhenxiang Xiao; Yu Du; Lin Zhao; Lu Zhang; Zihao Wu; Dajiang Zhu; Tuo Zhang; Dezhong Yao; Xintao Hu; Tianming Liu; Xi Jiang

doi:10.1109/TNNLS.2023.3342810

A Unified and Biologically Plausible Relational Graph Representation of Vision Transformers

Yuzhong Chen, Zhenxiang Xiao, Yu Du, Lin Zhao, Lu Zhang, Zihao Wu, Dajiang Zhu, Tuo Zhang, Dezhong Yao, Xintao Hu, Tianming Liu, Xi Jiang

自动化学院

科研成果: 期刊稿件 › 文章 › 同行评审

1 引用（Scopus）

摘要

Vision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph.

源语言	英语
页（从-至）	3231-3243
页数	13
期刊	IEEE Transactions on Neural Networks and Learning Systems
卷	36
期	2
DOI	https://doi.org/10.1109/TNNLS.2023.3342810
出版状态	已出版 - 2025

访问文件

10.1109/TNNLS.2023.3342810

其它文件与链接

链接到 Scopus 的出版物

引用此

@article{7d0329550d5d4e24919dd2fa2da0250e,

title = "A Unified and Biologically Plausible Relational Graph Representation of Vision Transformers",

abstract = "Vision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph.",

keywords = "Artificial neural network (ANN), biological neural network (BNN), relational graph, vision transformer (ViT)",

author = "Yuzhong Chen and Zhenxiang Xiao and Yu Du and Lin Zhao and Lu Zhang and Zihao Wu and Dajiang Zhu and Tuo Zhang and Dezhong Yao and Xintao Hu and Tianming Liu and Xi Jiang",

note = "Publisher Copyright: {\textcopyright} 2024 IEEE.",

year = "2025",

doi = "10.1109/TNNLS.2023.3342810",

language = "英语",

volume = "36",

pages = "3231--3243",

journal = "IEEE Transactions on Neural Networks and Learning Systems",

issn = "2162-237X",

publisher = "IEEE Computational Intelligence Society",

number = "2",

}

TY - JOUR

T1 - A Unified and Biologically Plausible Relational Graph Representation of Vision Transformers

AU - Chen, Yuzhong

AU - Xiao, Zhenxiang

AU - Du, Yu

AU - Zhao, Lin

AU - Zhang, Lu

AU - Wu, Zihao

AU - Zhu, Dajiang

AU - Zhang, Tuo

AU - Yao, Dezhong

AU - Hu, Xintao

AU - Liu, Tianming

AU - Jiang, Xi

PY - 2025

Y1 - 2025

N2 - Vision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph.

AB - Vision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph.

KW - Artificial neural network (ANN)

KW - biological neural network (BNN)

KW - relational graph

KW - vision transformer (ViT)

UR - http://www.scopus.com/inward/record.url?scp=85181581900&partnerID=8YFLogxK

U2 - 10.1109/TNNLS.2023.3342810

DO - 10.1109/TNNLS.2023.3342810

M3 - 文章

AN - SCOPUS:85181581900

SN - 2162-237X

VL - 36

SP - 3231

EP - 3243

JO - IEEE Transactions on Neural Networks and Learning Systems

JF - IEEE Transactions on Neural Networks and Learning Systems

IS - 2

ER -

A Unified and Biologically Plausible Relational Graph Representation of Vision Transformers

摘要

访问文件

其它文件与链接

指纹

引用此