File(s) under permanent embargo
Short Text Similarity Measurement Using Context from Bag of Word Pairs and Word Co-occurrence
conference contribution
posted on 2020-01-01, 00:00 authored by Shuiqiao Yang, Guangyan HuangGuangyan Huang, Bahadorreza OfoghiBahadorreza Ofoghi© Springer Nature Singapore Pte Ltd 2020. With the rapid development of social networks, short texts have become a prevalent form of social communications on the Internet. Measuring the similarity between short texts is a fundamental task to many applications, such as social network text querying, short text clustering and geographical event detection for smart city. However, short texts in social media always show limited contextual information and they are sparse, noisy and ambiguous. Hence, effectively measuring the distance between short texts is a challenging task. In this paper, we propose a new heuristic word pair distance measurement (WPDM) technique for short texts, which exploits the corpus level word relations and enriches the context of each short text with bag of word pairs representation. We first adjust Jaccard similarity to measure the distance between words. Then, words are paired up to capture latent semantics in a short text document and thus transfer short text into a bag of word pairs representation. The similarity between short text documents is finally calculated through averaging the distances of the word pairs. Experimental results on a real-world dataset demonstrate that the proposed WPDM is effective and achieves much better performance than state-of-the-art methods.
History
Event
Data Science. Conference (2019 : 6th : Ningbo, China)Volume
1179Series
Communications in Computer and Information SciencePagination
221 - 231Publisher
SpringerLocation
Ningbo, ChinaPlace of publication
Berlin, GermanyPublisher DOI
Start date
2019-05-15End date
2019-05-20ISSN
1865-0929eISSN
1865-0937ISBN-13
9789811528095Language
engPublication classification
E1 Full written paper - refereedTitle of proceedings
ICDS 2019 : Data science : 6th international conference, ICDS 2019, Ningbo, China, May 15-20, 2019, revised selected papersUsage metrics
Categories
No categories selectedLicence
Exports
RefWorks
BibTeX
Ref. manager
Endnote
DataCite
NLM
DC