Average-case performance of the Apriori Algorithm

被引：16

作者：

Purdom, PW ^{[1
]}

Van Gucht, D

Groth, DP

机构：

[1] Indiana Univ, Dept Comp Sci, Bloomington, IN 47405 USA

[2] Indiana Univ, Sch Informat, Bloomington, IN 47405 USA

来源：

SIAM JOURNAL ON COMPUTING | 2004年 / 33卷 / 05期

关键词：

data mining; algorithm analysis; Apriori Algorithm;

D O I：

10.1137/S0097539703422881

中图分类号：

TP301 [理论、方法];

学科分类号：

081202 ;

摘要：

The failure rate of the Apriori Algorithm is studied analytically for the case of random shoppers. The time needed by the Apriori Algorithm is determined by the number of item sets that are output (successes: item sets that occur in at least k baskets) and the number of item sets that are counted but not output (failures: item sets where all subsets of the item set occur in at least k baskets but the full set occurs in less than k baskets). The number of successes is a property of the data; no algorithm that is required to output each success can avoid doing work associated with the successes. The number of failures is a property of both the algorithm and the data. We find that under a wide range of conditions the performance of the Apriori Algorithm is almost as bad as is permitted under sophisticated worst-case analyses. In particular, there is usually a bad level with two properties: (1) it is the level where nearly all of the work is done, and (2) nearly all item sets counted are failures. Let l be the level with the most successes, and let the number of successes on level l be approximately ((m)(l)) for some m. Then, typically, the Apriori Algorithm has total output proportional to approximately ((m)(l)) and total work proportional to approximately ((m)(l+1)). In addition m is usually much larger than l, so the ratio of work to output is proportional to approximately m/(l+1). The analytical results for random shoppers are compared against measurements for three data sets. These data sets are more like the usual applications of the algorithm. In particular, the buying patterns of the various shoppers are highly correlated. For most thresholds, these data sets also have a bad level. Thus, under most conditions nearly all of the work done by the Apriori Algorithm consists in counting item sets that fail.

引用

页码：1223 / 1260

页数：38

共 50 条

[31] An average-case sublinear forward algorithm for the haploid Li and Stephens model
Yohei M. Rosen
Benedict J. Paten
Algorithms for Molecular Biology, 14
[32] Average-case analysis of a greedy algorithm for the 0/1 knapsack problem
Calvin, JM
Leung, JYT
OPERATIONS RESEARCH LETTERS, 2003, 31 (03) : 202 - 210
[33] AVERAGE-CASE ANALYSIS OF ROBINSON UNIFICATION ALGORITHM WITH 2 DIFFERENT VARIABLES
CASAS, R
DIAZ, J
STEYAERT, JM
INFORMATION PROCESSING LETTERS, 1989, 31 (05) : 227 - 232
[34] A BIDIRECTIONAL SHORTEST-PATH ALGORITHM WITH GOOD AVERAGE-CASE BEHAVIOR
LUBY, M
RAGDE, P
ALGORITHMICA, 1989, 4 (04) : 551 - 567
[35] Average-case approximation ratio of the 2-opt algorithm for the TSP
Engels, Christian
Manthey, Bodo
OPERATIONS RESEARCH LETTERS, 2009, 37 (02) : 83 - 84
[36] A BIDIRECTIONAL SHORTEST-PATH ALGORITHM WITH GOOD AVERAGE-CASE BEHAVIOR
LUBY, M
RAGDE, P
LECTURE NOTES IN COMPUTER SCIENCE, 1985, 194 : 394 - 403
[37] An average-case sublinear forward algorithm for the haploid Li and Stephens model
Rosen, Yohei M.
Paten, Benedict J.
ALGORITHMS FOR MOLECULAR BIOLOGY, 2019, 14 (1)
[38] Average-case analysis of multiple Quickselect: An algorithm for finding order statistics
Lent, J
Mahmoud, HM
STATISTICS & PROBABILITY LETTERS, 1996, 28 (04) : 299 - 310
[39] On the average-case complexity of underdetermined functions
Chashkin, Aleksandr V.
DISCRETE MATHEMATICS AND APPLICATIONS, 2018, 28 (04): : 201 - 221
[40] AVERAGE-CASE LOWER BOUNDS FOR SEARCHING
MCDIARMID, C
SIAM JOURNAL ON COMPUTING, 1988, 17 (05) : 1044 - 1060

← 1 2 3 4 5 →