Choosing an Index

Each cluster validity index emphasizes different properties of a partition. Values from different index families are not on a common scale, so compare candidate partitions with the same index rather than comparing one index’s number directly with another’s.

The following operation support describes the default NumPy backend. Every listed index supports batch evaluation.

Implemented indices

Index

Name

Prefer

Range

Incremental

Remove/merge/split

CH

Calinski–Harabasz

Larger

[0, inf)

Yes

Yes

CONN

Connectivity

Larger

[0, 1]

FuzzyART Backend Only

No

cSIL

Centroid-based Silhouette

Larger

[-1, 1]

Yes

Yes

DB

Davies–Bouldin

Smaller

[0, inf)

Yes

Yes

GD43

Generalized Dunn 43

Larger

[0, inf)

Yes

Yes

GD53

Generalized Dunn 53

Larger

[0, inf)

Yes

Yes

PS

Partition Separation

Larger

[0, 1]

Yes

Yes

rCIP

Representative Cross Information Potential

Smaller

[0, inf)

Yes

Yes

WB

Within/Between

Smaller

[0, inf)

Yes

Yes

XB

Xie–Beni

Smaller

[0, inf)

Yes

Yes

Optional optimization coverage

Index

Numba batch

Numba sample updates

Numba remove/merge

JAX batch

JAX streaming

CH

Yes

Unchanged

Unchanged

Yes

Capacity

CONN

Unavailable

Unavailable

Unavailable

Unavailable

Unavailable

cSIL

Yes

Unchanged

Unchanged

Unavailable

Unavailable

DB

Yes

Distances

Distances

Unavailable

Unavailable

GD43

Yes

Distances

Distances

Unavailable

Unavailable

GD53

Yes

Unchanged

Unchanged

Unavailable

Unavailable

PS

Yes

Distances

Distances

Unavailable

Unavailable

rCIP

Unavailable

Unavailable

Unavailable

Unavailable

Unavailable

WB

Yes

Unchanged

Unchanged

Yes

Capacity

XB

Yes

Distances

Distances

Yes

Capacity

Yes means compiled batch kernels are available; it does not mean that every operation is compiled. Distances means centroid-distance kernels are compiled while other update logic remains in Python/NumPy. Unchanged means the operation is supported with backend="numba" but uses its existing NumPy implementation. Unavailable means the index rejects that numerical backend.

Capacity means JAX sample additions and update_many chunks require a positive capacity at construction. CH, WB, and XB also provide functional JAX batch and streaming APIs. All JAX modes require x64, and none support remove or merge. Numba retains remove/merge support for all eight supported indices, including the operations marked Unchanged.

NumPy remains the default for every index. CONN’s model_type selects its prototype learner separately from numerical backend selection; it does not enable the optional Numba or JAX numerical backends.

See Backends for installation, accelerated kernels, precision rules, and compilation costs. Backend availability does not guarantee a speedup.

CH, DB, GD43, GD53, WB, and XB summarize variants of within-cluster compactness and between-cluster separation. cSIL offers a bounded, centroid-based silhouette measure. PS measures partition separation, while rCIP uses distributional information. CONN is the specialized choice when connectivity between learned prototypes is important; see Using CONN before using it.

The value numpy.nan is used when an index is not yet defined, including many one-cluster states. When monitoring a stream, consider the trajectory only after enough clusters and samples have been observed. An undefined batch evaluation also emits a RuntimeWarning; incremental startup and structural operations remain silent. Use numpy.isnan to test whether a result is undefined. A computed score of 0.0 remains a valid result.

The conditions for a defined value are:

Definition conditions

Index

The score is defined when

CH

There are at least two clusters and within-cluster sum of squares is positive.

CONN

At least two prototypes or ART categories have been learned.

cSIL

There are at least two clusters. A local term with equal zero compactness and separation contributes 0.0.

DB

There are at least two clusters and every pair of centroids has positive separation.

GD43 and GD53

There are at least two clusters and at least one cluster has positive dispersion.

PS

There are at least two clusters and the cluster centroids have positive dispersion.

rCIP

There are at least two clusters.

WB

There are at least two clusters and between-cluster sum of squares is positive.

XB

There are at least two clusters and minimum centroid separation is positive.

The checks use exact zero comparisons. No small value is added to a denominator, so a very small nonzero denominator remains part of the metric’s result.

For example, exclude undefined values when consuming a streaming score:

import cvi
import numpy as np

index = cvi.CH()
for sample, label in zip(samples, labels):
    value = index.get_cvi(sample, label)
    if not np.isnan(value):
        print(value)