Deakin University
Browse

File(s) under permanent embargo

Short Text Similarity Measurement Using Context from Bag of Word Pairs and Word Co-occurrence

conference contribution
posted on 2020-01-01, 00:00 authored by Shuiqiao Yang, Guangyan HuangGuangyan Huang, Bahadorreza OfoghiBahadorreza Ofoghi
© Springer Nature Singapore Pte Ltd 2020. With the rapid development of social networks, short texts have become a prevalent form of social communications on the Internet. Measuring the similarity between short texts is a fundamental task to many applications, such as social network text querying, short text clustering and geographical event detection for smart city. However, short texts in social media always show limited contextual information and they are sparse, noisy and ambiguous. Hence, effectively measuring the distance between short texts is a challenging task. In this paper, we propose a new heuristic word pair distance measurement (WPDM) technique for short texts, which exploits the corpus level word relations and enriches the context of each short text with bag of word pairs representation. We first adjust Jaccard similarity to measure the distance between words. Then, words are paired up to capture latent semantics in a short text document and thus transfer short text into a bag of word pairs representation. The similarity between short text documents is finally calculated through averaging the distances of the word pairs. Experimental results on a real-world dataset demonstrate that the proposed WPDM is effective and achieves much better performance than state-of-the-art methods.

History

Event

Data Science. Conference (2019 : 6th : Ningbo, China)

Volume

1179

Series

Communications in Computer and Information Science

Pagination

221 - 231

Publisher

Springer

Location

Ningbo, China

Place of publication

Berlin, Germany

Start date

2019-05-15

End date

2019-05-20

ISSN

1865-0929

eISSN

1865-0937

ISBN-13

9789811528095

Language

eng

Publication classification

E1 Full written paper - refereed

Title of proceedings

ICDS 2019 : Data science : 6th international conference, ICDS 2019, Ningbo, China, May 15-20, 2019, revised selected papers

Usage metrics

    Research Publications

    Categories

    No categories selected

    Exports

    RefWorks
    BibTeX
    Ref. manager
    Endnote
    DataCite
    NLM
    DC