Choosing an Index¶
Each cluster validity index emphasizes different properties of a partition. Values from different index families are not on a common scale, so compare candidate partitions with the same index rather than comparing one index’s number directly with another’s.
The following operation support describes the default NumPy backend. Every listed index supports batch evaluation.
Index |
Name |
Prefer |
Range |
Incremental |
Remove/merge/split |
|---|---|---|---|---|---|
|
Calinski–Harabasz |
Larger |
|
Yes |
Yes |
|
Connectivity |
Larger |
|
FuzzyART Backend Only |
No |
|
Centroid-based Silhouette |
Larger |
|
Yes |
Yes |
|
Davies–Bouldin |
Smaller |
|
Yes |
Yes |
|
Generalized Dunn 43 |
Larger |
|
Yes |
Yes |
|
Generalized Dunn 53 |
Larger |
|
Yes |
Yes |
|
Partition Separation |
Larger |
|
Yes |
Yes |
|
Representative Cross Information Potential |
Smaller |
|
Yes |
Yes |
|
Within/Between |
Smaller |
|
Yes |
Yes |
|
Xie–Beni |
Smaller |
|
Yes |
Yes |
Index |
Numba batch |
Numba sample updates |
Numba remove/merge |
JAX batch |
JAX streaming |
|---|---|---|---|---|---|
|
Yes |
Unchanged |
Unchanged |
Yes |
Capacity |
|
Unavailable |
Unavailable |
Unavailable |
Unavailable |
Unavailable |
|
Yes |
Unchanged |
Unchanged |
Unavailable |
Unavailable |
|
Yes |
Distances |
Distances |
Unavailable |
Unavailable |
|
Yes |
Distances |
Distances |
Unavailable |
Unavailable |
|
Yes |
Unchanged |
Unchanged |
Unavailable |
Unavailable |
|
Yes |
Distances |
Distances |
Unavailable |
Unavailable |
|
Unavailable |
Unavailable |
Unavailable |
Unavailable |
Unavailable |
|
Yes |
Unchanged |
Unchanged |
Yes |
Capacity |
|
Yes |
Distances |
Distances |
Yes |
Capacity |
Yes means compiled batch kernels are available; it does not mean that every
operation is compiled. Distances means centroid-distance kernels are
compiled while other update logic remains in Python/NumPy. Unchanged means
the operation is supported with backend="numba" but uses its existing NumPy
implementation. Unavailable means the index rejects that numerical backend.
Capacity means JAX sample additions and update_many chunks require a
positive capacity at construction. CH, WB, and XB also provide functional
JAX batch and streaming APIs. All JAX modes require x64, and none support remove
or merge. Numba retains remove/merge support for all eight supported indices,
including the operations marked Unchanged.
NumPy remains the default for every index. CONN’s model_type selects its
prototype learner separately from numerical backend selection; it does not
enable the optional Numba or JAX numerical backends.
See Backends for installation, accelerated kernels, precision rules, and compilation costs. Backend availability does not guarantee a speedup.
CH, DB, GD43, GD53, WB, and XB summarize variants of
within-cluster compactness and between-cluster separation. cSIL offers a
bounded, centroid-based silhouette measure. PS measures partition
separation, while rCIP uses distributional information. CONN is the
specialized choice when connectivity between learned prototypes is important;
see Using CONN before using it.
The value numpy.nan is used when an index is not yet defined, including
many one-cluster states. When monitoring a stream, consider the trajectory
only after enough clusters and samples have been observed. An undefined batch
evaluation also emits a RuntimeWarning; incremental startup and structural
operations remain silent. Use numpy.isnan to test whether a result is
undefined. A computed score of 0.0 remains a valid result.
The conditions for a defined value are:
Index |
The score is defined when |
|---|---|
|
There are at least two clusters and within-cluster sum of squares is positive. |
|
At least two prototypes or ART categories have been learned. |
|
There are at least two clusters. A local term with equal zero compactness and separation contributes |
|
There are at least two clusters and every pair of centroids has positive separation. |
|
There are at least two clusters and at least one cluster has positive dispersion. |
|
There are at least two clusters and the cluster centroids have positive dispersion. |
|
There are at least two clusters. |
|
There are at least two clusters and between-cluster sum of squares is positive. |
|
There are at least two clusters and minimum centroid separation is positive. |
The checks use exact zero comparisons. No small value is added to a denominator, so a very small nonzero denominator remains part of the metric’s result.
For example, exclude undefined values when consuming a streaming score:
import cvi
import numpy as np
index = cvi.CH()
for sample, label in zip(samples, labels):
value = index.get_cvi(sample, label)
if not np.isnan(value):
print(value)