FAQ¶
Which detector should I use, online or offline?¶
online_changepoint_detection (Adams & MacKay 2007) processes the series one
point at a time and, after each point, gives the posterior over how long the
current segment has lasted. Use it for streams, or when you want to know how
quickly a change would have been noticed. offline_changepoint_detection
(Fearnhead 2006) sees the whole series and returns the posterior probability
of a changepoint at each position, using data on both sides of it. Use it for
retrospective analysis; it is usually sharper. Both cost O(T²).
The two detectors report the same change at indices one apart. Why?¶
Different conventions, both documented in the docstrings:
- Online (
get_map_changepoints,changepoint_probabilities,viterbi_changepoints): the index of the first point of the new segment. A series whose first 80 points come from one regime reports 80. - Offline (
Pcp[j, t], andtorch.exp(Pcp).sum(0)[t]): the probability that a segment ends att, i.e. the last point of the old regime. The same series reports 79.
So offline index + 1 == online index.
Does the scale of my data matter? (issue #34)¶
Yes. The priors are on the mean and variance of the data, so their
hyperparameters have units, and rescaling the data without rescaling them
changes the model. For the univariate Normal-Gamma model (online StudentT
with alpha, beta, kappa, mu; offline StudentT with alpha0, beta0,
kappa0, mu0):
| parameter | meaning | units |
|---|---|---|
mu |
prior mean of a segment | data units |
kappa |
how many observations the prior mean is worth | none |
alpha |
half the number of observations the variance prior is worth | none |
beta |
alpha times the prior guess of the variance |
data units² |
Multiplying the data by c is equivalent to using mu * c and beta * c²
with kappa and alpha unchanged. With beta / alpha far from the actual
within-segment variance, or mu far from the data, the first points of
every segment look surprising and the detector over- or under-reacts.
Practical choices: standardize the data (subtract a typical level, divide by
a typical within-segment standard deviation, ideally estimated on a
calibration window rather than on the whole series), or set mu to the
expected level and beta = alpha * expected_variance. The values in the
examples (alpha=0.1, beta=0.01, kappa=1, mu=0) encode "around zero,
variance about 0.1, but I am not sure": with df = 2 * alpha = 0.2 the
predictive is extremely heavy-tailed, which is why they still work on
roughly unit-scale data.
The multivariate classes work the same way but parametrize the prior on
the covariance differently. Online MultivariateT takes scale, the
Wishart scale W on the precision: to encode a prior covariance C pass
scale = inv(C) / dof (default I / dof, unit prior covariance). Offline
MultivariateT takes Psi0, the inverse-Wishart scale on the covariance
side (Psi0 = inv(W)): the same prior covariance C is Psi0 = dof0 * C,
and the default dof0 * I is the same unit prior covariance as online.
mu/mu0 are in data units in both.
Data far from zero (timestamps, prices, counters) is handled without
loss of precision: the offline likelihoods compute their statistics on data
centered on its mean, and the online ones keep their state relative to the
first observation, so an offset of 1e8 gives the same result as the same
series around 0. Pass such data as float64 (a float64 tensor or NumPy
array): a float32 value near 1e8 is only resolved to steps of 8, before
the detector ever sees it.
How do I make the detector more or less sensitive? (issue #31)¶
In order of importance:
- The hazard, i.e. the expected segment length.
constant_hazard(lam)puts prior probability1 / lamon a change at every step. Largerlammeans fewer detections, more confidence needed, slightly longer delay; smallerlammeans more, earlier, and more false alarms. This is the main knob and it is about the data, not the model: set it near the segment length you expect. If segments shorter than some minimum are implausible,negative_binomial_hazard(k, p)(mean lengthk / p) puts little probability on changes soon after the last one. - How much you trust the prior versus the first points of a new segment.
kappa(for the mean) andalpha(for the variance) act as pseudo-counts. Small values let a few points establish a new regime quickly; larger values make the detector wait for more evidence.betaandmushould describe the data (previous question) rather than be used as sensitivity knobs. - How you read the output.
changepoint_probabilities(R, lag)trades delay for confidence: a largerlaggives a more decisive probability,lagobservations later.get_map_changepoints(R, min_separation=k)drops starts closer thankpoints to an earlier one, for when the posterior hesitates between neighboring points.
Offline, the equivalent of the hazard is the segment-length prior:
const_prior(p=1/(T+1)) is the flat default; geometric_prior(p=1/L)
encodes an expected segment length L; negative_binomial_prior allows a
peaked length distribution, and negative_binomial_hazard with the same
k and p is its online counterpart. Leave truncate at its default: the sum is exact
and the legacy truncation can drop the dominant term.
My data are not normally distributed. Can I still use this? (issue #36)¶
Every likelihood here assumes that within a segment the observations are
independent draws from one distribution. The Gaussian ones detect changes in
the mean and/or the (co)variance; Poisson detects changes in the rate of
counts:
| likelihood | within-segment model |
|---|---|
online StudentT, offline StudentT |
i.i.d. Normal, unknown mean and variance (Normal-Gamma prior) |
online MultivariateT |
i.i.d. multivariate Normal, unknown mean and covariance (Normal-Wishart) |
offline IndependentFeaturesLikelihood |
one Normal-Gamma model per dimension, independent |
offline MultivariateT |
i.i.d. multivariate Normal, unknown mean and covariance (Normal-Wishart) |
online NormalKnownVariance, offline NormalKnownVariance |
i.i.d. Normal with a known variance, unknown mean (Normal prior): mean changes only; multivariate offline input is independent dimensions |
online Poisson, offline Poisson |
i.i.d. Poisson counts, unknown rate (Gamma prior); multivariate offline input is independent Poisson dimensions |
offline FullCovarianceLikelihood |
multivariate Normal with unknown covariance and no mean parameter (mean zero, Xuan & Murphy 2007): it detects covariance changes; segments that differ in mean are misread as scale changes, so use MultivariateT when means move |
When the data are not Gaussian the detector still runs, and the question is what the misspecification does to it:
- Heavy tails or outliers: single extreme points look like the start of
a new segment. The Student-t predictive already tolerates some of this;
a larger
lamorkappahelps, and so does a transform (log for positive, right-skewed quantities such as latencies or prices). - Counts: use
online_likelihoods.Poisson/offline_likelihoods.Poisson(non-negative integers only). Counts that vary more than a Poisson allows (overdispersion) will show extra changepoints; a square-root or Anscombe transform withStudentTis the alternative. - Bounded data: a transform (logit for proportions) usually gets you close enough.
- Autocorrelation or slow drift: the model has no notion of dynamics within a segment, so a drift is reported as a sequence of small changes. Differencing, or modeling residuals from a trend, is the usual fix.
- Changes in something other than mean or variance (e.g. in autocorrelation) are not detected.
In short: use it when "piecewise stationary with Gaussian-ish noise" is a reasonable description after a transform, and check on a segment you trust that the residuals look plausible.
Why is it slow on my laptop with a GPU?¶
Since 1.2.0 the default device is the CPU. Earlier versions selected CUDA or Apple MPS automatically when present, but the online recursion is a sequential loop over small tensors, and each step on an accelerator pays a launch cost. Measured on an Apple M-series laptop (PyTorch 2.14), CPU against MPS:
| workload | CPU | MPS |
|---|---|---|
online StudentT, 1 000 points |
0.16 s | 2.5 s |
online StudentT, 5 000 points |
1.7 s | 11 s |
online MultivariateT, 10-D, 1 000 points |
0.56 s | 17 s |
The offline detector needs float64 and always runs on the CPU when MPS is
selected. On 1.1.0 or earlier, pass device="cpu" to both the likelihood
and the detector. Opt into an accelerator only after measuring on your
hardware; CUDA has not been benchmarked (issue #43).