Journal article
A model for word clustering
JA Thom, J Zobel
Journal of the American Society for Information Science | Published : 1992
Abstract
It is common to model the distribution of words in text by measures such as the Poisson approximation. However, these measures ignore effects such as clustering: our analysis of document collections demonstrates that the Poisson approximation can significantly overestimate the probability that a document contains a word. Based on our analysis, we propose a new model for distribution of words in text, and show how this model can be used to estimate the probability that a document contains a word and the number of distinct words in a document. © 1992 John Wiley & Sons, Inc. Copyright © 1992 John Wiley & Sons, Inc.