This paper describes the results of applying computational text mining to the Tipiṭaka, or Pali Canon, the canonical scripture of Therevāda Buddhism. Individual volumes of the Tipiṭaka are divided into “clusters” using purely computation tools, and in many cases these clusters appear to match the rough scholarly consensus around the relative age of the volumes. Texts are also summarized into “word clouds” based on relative word frequency, and these also seem to reflect the underlying themes of the texts. While these initial results are essentially confirmational ratherthan novel, they suggest these approaches will be valuable additions to the Pali scholar’s toolbox.
目次
Computational text mining 107 Particular challenges of the Pali Canon 108 Categorizing the Tipiṭaka through k-means clustering 116 Categorizing volumes of the Sutta Piṭaka 121 Summarizing the Sutta Piṭaka 125 Limitations and future directions 128 Tentative conclusions 129 Abbreviations 130