A heuristic approach to determine an appropriate number of topics in topic modeling

详细信息查看全文

作者：Weizhong Zhao ; James J Chen ; Roger Perkins ; Zhichao Liu ; Weigong Ge…
关键词：Rate of perplexity change (RPC) ; perplexity ; topic number ; latent Dirichlet allocation (LDA)
刊名：BMC Bioinformatics
出版年：2015
出版时间：December 2015
年：2015
卷：16
期：13-supp
全文大小：5,144 KB
刊物主题：Bioinformatics; Microarrays; Computational Biology/Bioinformatics; Computer Appl. in Life Sciences; Combinatorial Libraries; Algorithms;
出版者：BioMed Central
ISSN：1471-2105

文摘

Background Topic modelling is an active research field in machine learning. While mainly used to build models from unstructured textual data, it offers an effective means of data mining where samples represent documents, and different biological endpoints or omics data represent words. Latent Dirichlet Allocation (LDA) is the most commonly used topic modelling method across a wide number of technical fields. However, model development can be arduous and tedious, and requires burdensome and systematic sensitivity studies in order to find the best set of model parameters. Often, time-consuming subjective evaluations are needed to compare models. Currently, research has yielded no easy way to choose the proper number of topics in a model beyond a major iterative approach.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700