AI 中转降智检测
测测你正在用的中转 API 会不会在简单题上明显答错 —— 被降配、砍上下文或换成很差的模型,通常会在这里露馅。填你自己的中转信息,依次跑 6 道答案固定的题。
How it works
Six fixed questions with known answers — instruction following, Base64 decoding, a needle in about 3.5k characters of filler, a reasoning trap, arithmetic and code evaluation — checked by string matching, no judge model. A capable model normally passes all six; a badly downgraded model or one with its context cut tends to fail some. It cannot tell which model answered: cheap small models can pass too. You get a 0-100 score plus a per-question breakdown; each question waits at most 60 seconds.
FAQ
Some relays quietly route your request to a cheaper, smaller model than the one you asked for — you pay flagship prices and get budget output. Others trim context windows or strip capabilities. This tool catches the cases where that makes simple answers go wrong; it cannot tell which model actually answered.
It sends six fixed questions whose correct answers are known in advance (instruction following, Base64 decoding, a needle in about 3.5k characters, a reasoning trap, arithmetic, code evaluation) and checks the replies by string matching — no judge model. A capable model normally passes all six; one that is badly downgraded or has its context cut tends to fail some. Cheap small models can pass too, so a pass is not proof of which model answered. You get a 0-100 score and a per-question breakdown; if any request fails, the run is marked inconclusive.
The key is used for that one test run and is never stored, logged or reused. The questions and pass rules are open source (probe_spec.json in the repo), and you can run the same checks locally with the CLI instead.
Yes — that is the point. It works against any OpenAI-compatible endpoint. Test your current provider, test us, compare. We publish our own score rather than asking you to trust us.