This is an interface to the MOA implementation of streamKM++.
Details
streamKM++ uses a tree-based sampling strategy to build a small weighted sample of the stream (a coreset). The MOA implementation applies k-means++ to find the requested number of centers in the coreset.
Notes
The clusterer can process at most
lengthpoints. Processing more points causes anArrayIndexOutOfBoundsException.The coreset is not exposed as micro-clusters; only macro-clusters can be requested.
References
Marcel R. Ackermann, Christiane Lammersen, Marcus Maertens, Christoph Raupach, Christian Sohler, Kamil Swierkot. StreamKM++: A Clustering Algorithm for Data Streams. In: Proceedings of the 12th Workshop on Algorithm Engineering and Experiments (ALENEX '10), 2010.
See also
Other DSC_MOA:
DSC_BICO_MOA(),
DSC_CluStream(),
DSC_ClusTree(),
DSC_DStream_MOA(),
DSC_DenStream(),
DSC_MCOD(),
DSC_MOA()
Examples
set.seed(1000)
stream <- DSD_Gaussians(k = 3, d = 2, noise = 0.05)
# cluster with streamKM++
streamkm <- DSC_StreamKM(sizeCoreset = 100, numClusters = 3, length = 1000)
update(streamkm, stream, 100)
streamkm
#> StreamKM
#> Class: moa/clusterers/streamkm/StreamKM, DSC_MOA, DSC_Micro, DSC
#> Number of macro-clusters: 3
# plot macro-clusters (no access to micro-clusters)
plot(streamkm, stream)