This dataset is generated by the method described by Agrawal and Srikant (1994) using the reimplementation in arules which also retains the patterns used in the generation process.
Details
Agrawal.db contains the dataset (1000 items/20000 transactions) and
Agrawal.pat contains the patterns that were used to create the
dataset.
References
Rakesh Agrawal and Ramakrishnan Srikant (1994). Fast algorithms for mining association rules in large databases. In Jorge B. Bocca, Matthias Jarke, and Carlo Zaniolo, editors, Proceedings of the 20th International Conference on Very Large Data Bases, VLDB, pages 487-499, Santiago, Chile.
Examples
data(Agrawal)
summary(Agrawal.pat)
#> set of 2000 itemsets
#>
#> most frequent items:
#> item938 item446 item457 item615 item594 (Other)
#> 38 37 34 29 28 3844
#>
#> element (itemset/transaction) length distribution:sizes
#> 1 2 3 4 5 6
#> 721 759 353 132 26 9
#>
#> Min. 1st Qu. Median Mean 3rd Qu. Max.
#> 1.000 1.000 2.000 2.005 3.000 6.000
#>
#> summary of quality measures:
#> pWeights pCorrupts
#> Min. :8.742e-07 Min. :0.0000
#> 1st Qu.:1.476e-04 1st Qu.:0.2748
#> Median :3.392e-04 Median :0.4881
#> Mean :5.000e-04 Mean :0.4920
#> 3rd Qu.:6.899e-04 3rd Qu.:0.7085
#> Max. :3.150e-03 Max. :1.0000
#>
#> includes transaction ID lists: FALSE
summary(Agrawal.db)
#> transactions as itemMatrix in sparse format with
#> 20000 rows (elements/itemsets/transactions) and
#> 1000 columns (items) and a density of 0.0099933
#>
#> most frequent items:
#> item446 item938 item818 item457 item401 (Other)
#> 1638 1514 1450 1397 1389 192478
#>
#> element (itemset/transaction) length distribution:
#> sizes
#> 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16
#> 16 68 215 427 763 1234 1813 2215 2341 2437 2320 1896 1457 1045 739 447
#> 17 18 19 20 21 22 23 24
#> 260 171 74 25 16 15 2 4
#>
#> Min. 1st Qu. Median Mean 3rd Qu. Max.
#> 1.000 8.000 10.000 9.993 12.000 24.000
#>
#> includes extended item information - examples:
#> labels
#> 1 item1
#> 2 item2
#> 3 item3
#>
#> includes extended transaction information - examples:
#> transactionID
#> 1 trans1
#> 2 trans2
#> 3 trans3
## the data set was generated with the following code
if (FALSE) { # \dontrun{
Agrawal.pat <- random.patterns(1000, nPats = 2000, method = "agrawal",
lPats = 2, corr = 0.5, cmean = 0.5, cvar = 0.1, iWeight = NULL,
verbose = FALSE)
Agrawal.db <- random.transactions(1000, 20000, method="agrawal",
patterns = Agrawal.pat)
} # }