TY - JOUR
T1 - Chinese Legal Case Similarity Matching Based on Text Importance Extraction
AU - Fan, Aman
AU - Wang, Shaoxi
AU - Wang, Yanchuan
N1 - Publisher Copyright:
© IEEE. 2013 IEEE.
PY - 2025
Y1 - 2025
N2 - Similarity case matching can effectively enhance the efficiency of case adjudication and promote judicial fairness. Recent advances in natural language processing (NLP), particularly those based on deep learning technologies, have significantly enhanced the intelligent development of similar case judgments. The BERT model can efficiently extract features from legal texts by utilizing self-attention mechanisms, thereby facilitating subsequent matching tasks. However, traditional BERT models are often constrained by the input text length. To achieve better comparison results for long case descriptions, an iterative unsupervised clustering method is employed to evaluate the importance of legal case texts during contrastive learning. This results in extracted texts that align more closely with cluster centers in the feature space, thus becoming more representative. Extracting key information and retaining important legal texts can reduce the input length for the BERT model, thereby improving its performance. The selected case statement text is fed into a model based on the BERT framework for similarity case matching. Compared to inputting the original case statement text, the approach proposed in the paper can more effectively retain critical information about the case. The accuracy of our method on the test set is 75.08%, outperforming all existing methods on the public CAIL2019-SCM dataset.
AB - Similarity case matching can effectively enhance the efficiency of case adjudication and promote judicial fairness. Recent advances in natural language processing (NLP), particularly those based on deep learning technologies, have significantly enhanced the intelligent development of similar case judgments. The BERT model can efficiently extract features from legal texts by utilizing self-attention mechanisms, thereby facilitating subsequent matching tasks. However, traditional BERT models are often constrained by the input text length. To achieve better comparison results for long case descriptions, an iterative unsupervised clustering method is employed to evaluate the importance of legal case texts during contrastive learning. This results in extracted texts that align more closely with cluster centers in the feature space, thus becoming more representative. Extracting key information and retaining important legal texts can reduce the input length for the BERT model, thereby improving its performance. The selected case statement text is fed into a model based on the BERT framework for similarity case matching. Compared to inputting the original case statement text, the approach proposed in the paper can more effectively retain critical information about the case. The accuracy of our method on the test set is 75.08%, outperforming all existing methods on the public CAIL2019-SCM dataset.
KW - Similarity case matching
KW - contrastive learning
KW - deep learning
KW - unsupervised clustering
UR - https://www.scopus.com/pages/publications/105009957856
U2 - 10.1109/ACCESS.2025.3585265
DO - 10.1109/ACCESS.2025.3585265
M3 - 文章
AN - SCOPUS:105009957856
SN - 2169-3536
VL - 13
SP - 118745
EP - 118758
JO - IEEE Access
JF - IEEE Access
ER -