Efficiently answering top-k frequent term queries in temporal-categorical range

He, Z; Wang, L; Lu, C; Jing, Y; Zhang, K; Han, W; Li, Jianxin; Liu, C; Wang, X S

File(s) under permanent embargo

Efficiently answering top-k frequent term queries in temporal-categorical range

journal contribution

posted on 2021-10-01, 00:00 authored by Z He, L Wang, C Lu, Y Jing, K Zhang, W Han, Jianxin LiJianxin Li, C Liu, X S Wang

In the procedure of extracting hot topics and detecting emerging topic, counting term frequency is one of the most inevitable and time-consuming steps. For the purpose of text exploration, users may change the query range frequently, and the adjustment of ranges would cause recalculation of term frequency when finding hot terms, bringing unacceptable time cost. In addition, real-time update of dimensions is also a challenge. To address these problems, we first propose a novel data structure based on prefix cube to store terms and their frequencies, so that the time for counting term frequency gets a significant reduction. Based on the data structure, we propose an efficient range query algorithm that significantly decreases the number of input word lists involved in top-k queries. Considering the underlying dimension update, we also design an efficient maintenance mechanism to cope with different dimension updates. Finally, we conduct comprehensive experiments to validate the effectiveness of the proposed structure and the efficiency of the optimized query algorithm. We also prove that using the proposed data structure, the time cost of our algorithms in hot topic extraction and emerging topic detection can be reduced by about ten times compared with the previous algorithms.

History

Journal

Information Sciences

Volume

574

Pagination

238 - 258

Publisher

Elsevier

Location

Amsterdam, The Netherlands

Publisher DOI

https://doi.org/10.1016/j.ins.2021.05.081

ISSN

0020-0255

Language

eng

Publication classification

C1 Refereed article in a scholarly journal

Usage metrics

Keywords

Top-k query Temporal-categorical databases Term frequency counting Threshold-style algorithms Range query Science & Technology Technology Computer Science, Information Systems Computer Science

Licence

Exports

RefWorks

BibTeX

Ref. manager

Endnote

DataCite

NLM

DC

File(s) under permanent embargo

Efficiently answering top-k frequent term queries in temporal-categorical range

History

Journal

Volume

Pagination

Publisher

Location

Publisher DOI

ISSN

Language

Publication classification

Usage metrics

Categories

Keywords

Licence

Exports