论文信息 - Instance Smoothed Contrastive Learning for Unsupervised Sentence Embedding - 字舞流文

Instance Smoothed Contrastive Learning for Unsupervised Sentence Embedding

Contrastive learning-based methods, such as unsup-SimCSE, have achieved state-of-the-art (SOTA) performances in learning unsupervised sentence embeddings. However, in previous studies, each embedding used for contrastive learning only derived from one sentence instance, and we call these embeddings instance-level embeddings. In other words, each embedding is regarded as a unique class of its own, which may hurt the generalization performance. In this study, we propose IS-CSE (instance smoothing contrastive sentence embedding) to smooth the boundaries of embeddings in the feature space. Specifically, we retrieve embeddings from a dynamic memory buffer according to the semantic similarity to get a positive embedding group. Then embeddings in the group are aggregated by a self-attention operation to produce a smoothed instance embedding for further analysis. We evaluate our method on standard semantic text similarity (STS) tasks and achieve an average of 78.30%, 79.47%, 77.73%, and 79.42% Spearman’s correlation on the base of BERT-base, BERT-large, RoBERTa-base, and RoBERTa-large respectively, a 2.05%, 1.06%, 1.16% and 0.52% improvement compared to unsup-SimCSE.

Zhenzhong Lan | Yue Zhang | Hongliang He | Junlei Zhang

[1] T. Klein,et al. miCSE: Mutual Information Contrastive Learning for Low-shot Sentence Embeddings , 2022, ACL.

[2] Wei Wang,et al. Improving Contrastive Learning of Sentence Embeddings with Case-Augmented Positives and Retrieved Negatives , 2022, SIGIR.

[3] Wayne Xin Zhao,et al. Debiased Contrastive Learning of Unsupervised Sentence Representations , 2022, ACL.

[4] Qi Zhang,et al. PromptBERT: Improving BERT Sentence Embeddings with Prompts , 2022, EMNLP.

[5] Junlei Zhang,et al. S-SimCSE: Sampled Sub-networks for Contrastive Learning of Sentence Embedding , 2021, ArXiv.

[6] Wayne Xin Zhao,et al. Contrastive Curriculum Learning for Sequential User Behavior Modeling via Data Augmentation , 2021, CIKM.

[7] Xing Wu,et al. ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding , 2021, COLING.

[8] Fuzheng Zhang,et al. ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer , 2021, ACL.

[9] Danqi Chen,et al. SimCSE: Simple Contrastive Learning of Sentence Embeddings , 2021, EMNLP.

[10] Jiarun Cao,et al. Whitening Sentence Representations for Better Semantics and Faster Retrieval , 2021, ArXiv.

[11] Yiming Yang,et al. On the Sentence Embeddings from BERT for Semantic Textual Similarity , 2020, EMNLP.

[12] Phillip Isola,et al. Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere , 2020, ICML.

[13] Kaiming He,et al. Improved Baselines with Momentum Contrastive Learning , 2020, ArXiv.

[14] Ross B. Girshick,et al. Momentum Contrast for Unsupervised Visual Representation Learning , 2019, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

[15] Lysandre Debut,et al. HuggingFace's Transformers: State-of-the-art Natural Language Processing , 2019, ArXiv.

[16] Omer Levy,et al. RoBERTa: A Robustly Optimized BERT Pretraining Approach , 2019, ArXiv.

[17] Geoffrey E. Hinton,et al. When Does Label Smoothing Help? , 2019, NeurIPS.

[18] Hongxun Yao,et al. Illustrate your travel notes: web-based story visualization , 2018, ICIMCS '18.

[19] Oriol Vinyals,et al. Representation Learning with Contrastive Predictive Coding , 2018, ArXiv.

[20] Eneko Agirre,et al. SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation , 2017, *SEMEVAL.

[21] Holger Schwenk,et al. Supervised Learning of Universal Sentence Representations from Natural Language Inference Data , 2017, EMNLP.

[22] Sergey Ioffe,et al. Rethinking the Inception Architecture for Computer Vision , 2015, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).

[23] Roberto Navigli,et al. From senses to texts: An all-in-one graph-based approach for measuring semantic similarity , 2015, Artif. Intell..

[24] Claire Cardie,et al. SemEval-2015 Task 2: Semantic Textual Similarity, English, Spanish and Pilot on Interpretability , 2015, *SEMEVAL.

[25] Jimmy Ba,et al. Adam: A Method for Stochastic Optimization , 2014, ICLR.

[26] Jeffrey Pennington,et al. GloVe: Global Vectors for Word Representation , 2014, EMNLP.

[27] Claire Cardie,et al. SemEval-2014 Task 10: Multilingual Semantic Textual Similarity , 2014, *SEMEVAL.

[28] Marco Marelli,et al. A SICK cure for the evaluation of compositional distributional semantic models , 2014, LREC.

[29] Christopher Potts,et al. Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank , 2013, EMNLP.

[30] Eneko Agirre,et al. *SEM 2013 shared task: Semantic Textual Similarity , 2013, *SEMEVAL.

[31] Eneko Agirre,et al. SemEval-2012 Task 6: A Pilot on Semantic Textual Similarity , 2012, *SEMEVAL.

[32] Leif E. Peterson. K-nearest neighbor , 2009, Scholarpedia.

[33] B. Pang,et al. Seeing Stars: Exploiting Class Relationships for Sentiment Categorization with Respect to Rating Scales , 2005, ACL.

[34] Bing Liu,et al. Mining and summarizing customer reviews , 2004, KDD.

[35] Ellen M. Voorhees,et al. Building a question answering test collection , 2000, SIGIR '00.

[36] Pengtao Xie,et al. CERT: Contrastive Self-supervised Learning for Language Understanding , 2020, ArXiv.

[37] Jiashi Feng,et al. Residual Distillation: Towards Portable Deep Neural Networks without Shortcuts , 2020, NeurIPS.

[38] Ming-Wei Chang,et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding , 2019, NAACL.

[39] Eneko Agirre,et al. SemEval-2016 Task 1: Semantic Textual Similarity, Monolingual and Cross-Lingual Evaluation , 2016, *SEMEVAL.

[40] Chris Brockett,et al. Automatically Constructing a Corpus of Sentential Paraphrases , 2005, IJCNLP.

[41] Claire Cardie,et al. Annotating Expressions of Opinions and Emotions in Language , 2005, Lang. Resour. Evaluation.