Your model is being retired. What breaks?

Providers publish retirement dates. They do not tell you what changes in your prompts. We run yours against both models and measure the difference against your endpoint's own normal variation. Free, no signup, results in a few minutes.

What you get. For each prompt: whether the replacement model behaves differently in a way that exceeds ordinary run-to-run noise — reported as genuine drift, expected variation, or insufficient evidence. We never call disagreement a defect without a baseline.
Only pairs published in the provider's own deprecation table are offered.
Held in memory for this run only. Never written to our database, never logged, never shown in the report. Use a scoped or throwaway key if you prefer.
Each prompt is sent 8 times to each model — that is how we measure your endpoint's normal variation before comparing anything.

How the verdict is decided

We compute Cliff's delta between the model-to-model differences and the model-to-itself differences. A high-temperature endpoint that disagrees with itself constantly needs a large gap to register; a deterministic one registers a small one. Thresholds were fixed in writing before any data was collected.

Limitation, stated up front: this measures surface form — wording, structure, format, refusal, JSON validity. A change that keeps the wording and reverses the meaning is not reliably caught in this version.