Available models
Models this org's plan may select, with measured characteristics.
- Authentication
- Bearer token
- Retries
- Safe to repeat
- Body
- None
- Version
- 2026-09-03
Shaped for a model switcher rather than for a developer. The owner sees the lab that made it and one word for what it is good at - name is "Alibaba - Long memory", not qwen/qwen3.7-flash - because a model id is meaningless to the person choosing and mildly alarming to the rest.
ttft_ms is the number worth putting next to the choice, and it is ours: OpenRouter publishes no latency at all (its latency_last_30m returns null on every endpoint), so every figure here came from scripts/model_probe.py running one real support turn against our own host. measured_on ships with it, because a latency claim with no date is marketing rather than measurement.
Responses#
Errors#
Failures use one envelope on every endpoint, described in Retries, versioning and limits. The codes you are most likely to meet here:
unauthenticated· 401 — The request carried no API key, or one the API could not verify.forbidden· 403 — The key is valid, but it is not allowed to do this — either the scope is missing or the resource belongs to another workspace.rate_limited· 429 — Too many requests in the current window. The limit is per workspace, and some endpoints add a per-agent limit on top.
More Agents endpoints#
Something unclear or missing? Tell us and we’ll fix it.