cvi.PS

class cvi.PS(*, backend='numpy')

Bases: CVI

Partition Separation (PS) Cluster Validity Index.

References

  1. Miin-Shen Yang and Kuo-Lung Wu, “A new validity index for fuzzy clustering,” 10th IEEE International Conference on Fuzzy Systems. (Cat. No.01CH37297), Melbourne, Victoria, Australia, 2001, pp. 89-92, vol.1.

    1. Lughofer, “Extensions of vector quantization for incremental clustering,” Pattern Recognit., vol. 41, no. 3, pp. 995-1011, 2008.

__init__(*, backend='numpy')

Partition Separation (PS) initialization routine.

Parameters:

backend ({"numpy", "numba"}, default="numpy") – Select the numerical backend. Numba is loaded on demand.

Methods

__init__(*[, backend])

Partition Separation (PS) initialization routine.

get_cvi(data, label)

Update the CVI and return its criterion value.

merge(target_label, source_label)

Merge a source cluster into a target cluster.

remove(sample, label)

Remove a sample from an initialized CVI.

split(retained_label, new_label, count, ...)

Split a tracked subset from an existing cluster.

update_many(data, labels, *[, return_history])

Add a chunk to a fixed-capacity JAX stream, atomically on input errors.

Attributes

backend

Selected numerical backend (fixed for this object's lifetime).

capacity

Maximum distinct clusters for JAX streaming, or None for batch only.

info

stream_state

Immutable device state for JAX streaming; None until initialized.

property backend

Selected numerical backend (fixed for this object’s lifetime).

property capacity

Maximum distinct clusters for JAX streaming, or None for batch only.

get_cvi(data: ndarray, label: int | ndarray) → float

Update the CVI and return its criterion value.

Pass a one-dimensional sample and scalar integer label for an incremental update, or a two-dimensional batch and label vector for batch initialization. The object is mutated in both modes. A batch may be followed by incremental updates, but a second batch is not supported. JAX supports incremental additions when capacity is provided; it rejects remove and merge.

Parameters:
  • data (np.ndarray) – The sample(s) of features used for clustering.

  • label (Union[int, np.ndarray]) – The label(s) prescribed to the sample(s) by the clustering algorithm.

Returns:

The CVI’s criterion value.

Return type:

float

Raises:

ValueError – If the input dimensionality is invalid, feature dimensionality changes after initialization, batch labels contain fewer than two distinct values, or a second batch update is requested.

Warns:

RuntimeWarning – If the criterion is undefined after a batch evaluation. The returned value is still numpy.nan.

merge(target_label: int, source_label: int) → float

Merge a source cluster into a target cluster.

The target external label is retained and the source label is removed.

Parameters:
  • target_label (int) – External label of the cluster that remains after the merge.

  • source_label (int) – External label of the cluster merged into the target.

Returns:

The updated CVI criterion value.

Return type:

float

Raises:
  • NotImplementedError – If this index does not implement cluster merging.

  • ValueError – If the index is uninitialized, either label is unknown, or the two labels are equal.

remove(sample: ndarray, label: int) → float

Remove a sample from an initialized CVI.

The caller is responsible for ensuring that the sample belongs to the supplied cluster label. If the sample is the cluster’s final member, the empty cluster and its label are removed.

Parameters:
  • sample (numpy.ndarray) – One sample vector of features.

  • label (int) – External label of the cluster containing the sample.

Returns:

The updated CVI criterion value.

Return type:

float

Raises:
  • NotImplementedError – If this index does not implement removal.

  • ValueError – If the index is uninitialized, the label is unknown, the sample has the wrong shape, or the sample is inconsistent with the stored sufficient statistics.

split(retained_label: int, new_label: int, count: int, centroid: ndarray, *, compactness: float | None = None, covariance: ndarray | None = None) → float

Split a tracked subset from an existing cluster.

The existing external label is retained by the residual cluster. The supplied sufficient statistics are assigned to a new cluster with new_label. The total sample count and global mean do not change.

Parameters:
  • retained_label (int) – External label of the cluster retaining the residual statistics.

  • new_label (int) – Unused external label assigned to the split-off subset.

  • count (int) – Number of samples in the split-off subset.

  • centroid (numpy.ndarray) – Mean vector of the split-off subset.

  • compactness (float, optional) – Sum of squared distances from the subset centroid. Required by compactness-based indices when count is greater than one.

  • covariance (numpy.ndarray, optional) – Unregularized unbiased sample covariance of the subset. Required by rCIP when count is greater than one.

Returns:

The updated CVI criterion value.

Return type:

float

Raises:
  • NotImplementedError – If this index does not implement cluster splitting.

  • ValueError – If the index is uninitialized, labels or statistics are invalid, or the supplied subset is inconsistent with the retained cluster.

property stream_state

Immutable device state for JAX streaming; None until initialized.

update_many(data, labels, *, return_history=True)

Add a chunk to a fixed-capacity JAX stream, atomically on input errors.

Return a NumPy score history, or a Python final score when return_history=False. Empty chunks are no-ops, including on new objects. The entire chunk must fit the remaining cluster capacity. This is an incremental scan, distinct from the one-time batch get_cvi operation.