An OpenAI-compatible API accepts OpenAI-style requests and returns responses that OpenAI clients can parse, most often for Chat Completions. Reuse the OpenAI SDK by changing the base URL, API key and model ID. Tools, JSON schemas, streaming metadata and Responses support vary by endpoint and model. Compatibility describes the interface, as the Modular handbook explains.
The useful question is: “Does this endpoint support the requests my application sends?”
What OpenAI compatibility means
For a basic chat request, your client sends a model identifier and a messages array to a Chat Completions route. It expects an assistant reply under choices[].message. Streaming returns Server-Sent Events (SSE) containing choices[].delta. These are parts of the OpenAI Chat Completions contract.
A compatible provider implements that contract for the features it supports. Embeddings, model discovery, images, audio and Responses are separate API capabilities.

Base URL means the prefix, not the full operation URL. With Requesty, set base_url="https://router.requesty.ai/v1"; the SDK appends /chat/completions. Gemini's prefix ends in /v1beta/openai/. DeepSeek uses https://api.deepseek.com. Follow the Requesty quickstart, Gemini setup or DeepSeek setup rather than adding /v1 everywhere.
The destination determines credentials, billing and data policies. Compatibility does not transfer an OpenAI API balance or ChatGPT subscription to another provider. A shared payload format does not imply the same answer quality. Even OpenAI documents different parameter support between its own models.
Which providers support OpenAI-compatible APIs?
Search the matrix by provider, filter by category or require the features your application uses. Expand a row for its base URL, restrictions and sources. Tools means developer-defined function calling; provider-hosted search and code execution need separate checks.
| Status | Meaning |
|---|---|
| Documented | Official docs describe implemented support |
| Model-dependent | Requires a suitable model, deployment or server configuration |
| Limited | Implemented with documented restrictions |
| Unsupported | The documented contract excludes the feature |
| Ignored | Accepted without the requested behavior |
| Not documented | Reviewed docs do not establish support |
OpenAI-compatible endpoint reference
Documentation reviewed September 28, 2026.
Model-dependent results need model or configuration checks. Limited results need restriction checks.
- Hosted
Google's compatibility layer is in beta.
Base URL, caveats and sources
OpenAI SDK base URLhttps://generativelanguage.googleapis.com/v1beta/openai/Google's compatibility layer is in beta. Schema output and base64 image input are documented. reasoning_effort cannot be combined with Google's thinking_level or thinking_budget.
- Chat stream
- Documented
- Function tools
- Documented
- JSON object
- Not documented
- JSON Schema
- Model-dependent
- Vision
- Model-dependent
- Responses
- Not documented
- Anthropic direct layerDocsHosted
Anthropic recommends this layer primarily for testing, not most production use.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.anthropic.com/v1/Native Claude schema support does not change this layer's ignored response_format behavior.
Anthropic recommends this layer primarily for testing, not most production use. response_format, function strict, logprobs and reasoning_effort are ignored; n must be 1. Images use image_url.url; detail is ignored. System and developer messages are concatenated and hoisted to the start. Applicable multi-workspace personal/service-account keys require anthropic-workspace-id via default_headers in Python or defaultHeaders in JavaScript.
- Chat stream
- Documented
- Function tools
- Limited
- JSON object
- Ignored
- JSON Schema
- Ignored
- Vision
- Limited
- Responses
- Not documented
- MistralDocsHosted
Chat documents both JSON modes, named tool selection and parallel_tool_calls.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.mistral.ai/v1Chat documents both JSON modes, named tool selection and parallel_tool_calls. In JSON object mode, instruct the model to output JSON. Vision examples use a string image_url; test your client's content shape.
- Chat stream
- Documented
- Function tools
- Documented
- JSON object
- Documented
- JSON Schema
- Documented
- Vision
- Model-dependent
- Responses
- Not documented
- GroqDocsHosted
Parallel tools and strict schema output vary by model.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.groq.com/openai/v1Check previous_response_id behavior separately before relying on Responses state.
Parallel tools and strict schema output vary by model. Supplied logprobs, logit_bias, top_logprobs and messages[].name cause HTTP 400; n must be 1. Responses is in beta and its documentation conflicts on stored continuation.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Documented
- JSON Schema
- Model-dependent
- Vision
- Model-dependent
- Responses
- Limited
- TogetherDocsHosted
service_tier, store, metadata and prediction are accepted but ignored.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.together.ai/v1service_tier, store, metadata and prediction are accepted but ignored. n support varies and logprobs uses Together's shape. Vision models accept remote URLs and data URIs; detail is ignored. Cached-token and reasoning fields have provider-specific locations. Match errors by HTTP status and documented codes.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Model-dependent
- JSON Schema
- Model-dependent
- Vision
- Model-dependent
- Responses
- Not documented
- FireworksDocsHosted
Chat usage arrives on the final finish chunk.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.fireworks.ai/inference/v1Chat usage arrives on the final finish chunk. Default context-overflow behavior reduces the output budget; select error behavior if required. Schema output excludes external references and has regex restrictions. Responses supports stored continuation and store=false; background=true cannot combine with stream=true.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Documented
- JSON Schema
- Limited
- Vision
- Model-dependent
- Responses
- Documented
- DeepSeekDocsHosted
Chat response_format allows only text and json_object.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.deepseek.comStrict tool arguments use the separate beta base https://api.deepseek.com/beta. They do not implement Chat assistant JSON Schema output.
Chat response_format allows only text and json_object. Thinking mode rejects required and named tool choices. Vision is documented for deepseek-flash. Chat usage arrives on the final finish chunk. Responses is stateless; previous_response_id and other unsupported parameters are ignored.
- Chat stream
- Documented
- Function tools
- Limited
- JSON object
- Documented
- JSON Schema
- Unsupported
- Vision
- Model-dependent
- Responses
- Limited
- xAIDocsHosted
Schema output has best-effort keywords; oneOf behaves like anyOf.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.x.ai/v1Schema output has best-effort keywords; oneOf behaves like anyOf. Function calling supports named and parallel selection. Responses supports stored continuation; background is a compatibility field without implemented background execution.
- Chat stream
- Documented
- Function tools
- Documented
- JSON object
- Documented
- JSON Schema
- Limited
- Vision
- Model-dependent
- Responses
- Documented
- Perplexity Agent APIDocsHosted · Responses interface
Agent documents Responses streaming and /v1/responses as an alias for /v1/agent.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.perplexity.ai/v1Agent features do not establish Chat Completions support or OpenAI Responses text.format mapping. Router is a separate product.
Agent documents Responses streaming and /v1/responses as an alias for /v1/agent. Agent custom functions, image input and schema output belong to that surface. Agent schema output uses provider response_format. The September 27, 2026 Sonar transition preserves synchronous and streaming requests through reformulation as Agent requests; asynchronous Sonar requests ended.
- Chat stream
- Not documented
- Function tools
- Not documented
- JSON object
- Not documented
- JSON Schema
- Not documented
- Vision
- Not documented
- Responses
- Documented
- Hosted
Router is separate from Agent and is in private preview.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.perplexity.ai/router/v1Router is separate from Agent and is in private preview. Chat schemas are always strict. Several compatibility fields support only defaults; unknown fields cause HTTP 400. The schema includes image_url; select an image-capable model. Router Responses is stateless.
- Chat stream
- Documented
- Function tools
- Documented
- JSON object
- Documented
- JSON Schema
- Documented
- Vision
- Model-dependent
- Responses
- Limited
- CerebrasDocsHosted
JSON object mode cannot stream; n must be 1.
Base URL, caveats and sources
OpenAI SDK base URLhttps://api.cerebras.ai/v1JSON object mode cannot stream; n must be 1. Selected vision models accept base64 PNG/JPEG, not external HTTPS URLs or image_url.detail. gpt-oss-120b rejects tools plus response_format together.
- Chat stream
- Documented
- Function tools
- Documented
- JSON object
- Limited
- JSON Schema
- Model-dependent
- Vision
- Limited
- Responses
- Not documented
- Azure OpenAI v1DocsCloud
Azure v1 uses OpenAI(), without a required dated api-version parameter.
Base URL, caveats and sources
OpenAI SDK base URLhttps://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/Azure v1 uses OpenAI(), without a required dated api-version parameter. Replace the resource placeholder; model is your deployment name. Authentication supports a key or Microsoft Entra token provider. Disable parallel tools for strict structured function use.
- Chat stream
- Model-dependent
- Function tools
- Model-dependent
- JSON object
- Documented
- JSON Schema
- Model-dependent
- Vision
- Model-dependent
- Responses
- Documented
- Amazon Bedrock runtimeDocsCloud
AWS recommends runtime for new applications.
Base URL, caveats and sources
OpenAI SDK base URLhttps://bedrock-runtime.REGION.amazonaws.com/openai/v1Native Converse schema or vision support does not establish OpenAI-compatible response_format or image coverage.
AWS recommends runtime for new applications. Replace REGION and check the selected model's API coverage. Ordinary OpenAI SDK setup uses a Bedrock API key. Runtime does not implement compatible GET /models; use AWS discovery operations and model cards. GPT OSS Responses requires mantle; verify the model-specific endpoint path.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Not documented
- JSON Schema
- Not documented
- Vision
- Not documented
- Responses
- Model-dependent
- Ollama localDocsLocal / self-hosted
Local Ollama ignores the key value required by the client.
Base URL, caveats and sources
OpenAI SDK base URLhttp://localhost:11434/v1/Local Ollama ignores the key value required by the client. Load a suitable model for tools and vision. tool_choice and remote image URLs are unchecked in the published compatibility checklist. Responses is stateless. Ollama Cloud authentication differs and its structured-output support differs from local.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Documented
- JSON Schema
- Documented
- Vision
- Model-dependent
- Responses
- Limited
- vLLMDocsLocal / self-hosted
The address is illustrative; confirm your server host and port.
Base URL, caveats and sources
OpenAI SDK base URLhttp://localhost:8000/v1The address is illustrative; confirm your server host and port. Chat needs a chat template. Automatic tool choice needs a matching parser/template and server flags; named and required selection use structured decoding, enabled by default. Schema decoding is backend-dependent; image detail is unsupported. Match docs to the installed version. API-key settings do not protect every HTTP route; review network-hardening guidance before exposure.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Documented
- JSON Schema
- Model-dependent
- Vision
- Model-dependent
- Responses
- Documented
- LM StudioDocsLocal / self-hosted
Enable the server and load a suitable model.
Base URL, caveats and sources
OpenAI SDK base URLhttp://localhost:1234/v1Enable the server and load a suitable model. Tools use native or default prompting modes; malformed tool output can appear in message.content instead of message.tool_calls. Responses supports prior-response state.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Not documented
- JSON Schema
- Model-dependent
- Vision
- Model-dependent
- Responses
- Documented
- RequestyDocsGateway
One gateway interface reaches multiple providers.
Base URL, caveats and sources
OpenAI SDK base URLhttps://router.requesty.ai/v1One gateway interface reaches multiple providers. Select models using live catalog capability flags. Translated and native Responses routes differ in state, file, tool and event behavior.
- Chat stream
- Documented
- Function tools
- Model-dependent
- JSON object
- Model-dependent
- JSON Schema
- Model-dependent
- Vision
- Model-dependent
- Responses
- Limited
Perplexity Agent's custom functions, schema output and image inputs use its Agent surface. Its Sonar migration notice explains the September 27 transition. Keep that surface separate from Router when configuring a client.
Use an OpenAI-compatible endpoint
These examples use Requesty and openai/gpt-6-luna. Replace the URL, key and model to use another provider. Check current identifiers in the Models API reference.
Export these values; the examples read the LLM_* variables explicitly:
export LLM_BASE_URL="https://router.requesty.ai/v1"
export LLM_API_KEY="your-requesty-key"
export LLM_MODEL="openai/gpt-6-luna"The OpenAI Python SDK requires Python 3.10 or newer. Save this as example.py and run python example.py.
python -m pip install openaiimport os
from openai import OpenAI
client = OpenAI(
base_url=os.environ["LLM_BASE_URL"],
api_key=os.environ["LLM_API_KEY"],
)
response = client.chat.completions.create(
model=os.environ["LLM_MODEL"],
messages=[{"role": "user", "content": "Explain an API in one sentence."}],
)
print(response.choices[0].message.content)The OpenAI Node SDK requires Node 22 or newer. Save this as openai-example.ts.
npm install openaiimport OpenAI from 'openai';
const baseURL = process.env.LLM_BASE_URL;
const apiKey = process.env.LLM_API_KEY;
const model = process.env.LLM_MODEL;
if (!baseURL || !apiKey || !model) {
throw new Error('Set LLM_BASE_URL, LLM_API_KEY, and LLM_MODEL');
}
const client = new OpenAI({ baseURL, apiKey });
const response = await client.chat.completions.create({
model,
messages: [{ role: 'user', content: 'Explain an API in one sentence.' }],
});
console.log(response.choices[0].message.content);curl needs the full operation URL, as shown in the Requesty quickstart:
curl https://router.requesty.ai/v1/chat/completions \
-H "Authorization: Bearer $LLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-6-luna",
"messages": [{"role": "user", "content": "Explain an API in one sentence."}]
}'Run the TypeScript files with npx tsx; they use top-level await, so the project needs "type": "module" in package.json.
Only use endpoints you trust, keep keys server-side, and test with synthetic inputs. OpenAI's key-security guidance warns against exposing API keys in browsers or apps.
Use @ai-sdk/openai-compatible
Vercel's OpenAI Compatible provider wraps an OpenAI-shaped API for the AI SDK. Supply name, baseURL and credentials. Save this as ai-sdk-example.ts and run npx tsx ./ai-sdk-example.ts:
npm install ai @ai-sdk/openai-compatibleimport { createOpenAICompatible } from '@ai-sdk/openai-compatible';
import { generateText } from 'ai';
const baseURL = process.env.LLM_BASE_URL;
const apiKey = process.env.LLM_API_KEY;
const model = process.env.LLM_MODEL;
if (!baseURL || !apiKey || !model) {
throw new Error('Set LLM_BASE_URL, LLM_API_KEY, and LLM_MODEL');
}
const provider = createOpenAICompatible({ name: 'custom', baseURL, apiKey });
const { text } = await generateText({
model: provider(model),
prompt: 'Explain an API in one sentence.',
});
console.log(text);The provider interface requires name and baseURL. Choosing the adapter also chooses an API surface:
| Integration | Language-model route | Use it for |
|---|---|---|
@ai-sdk/openai-compatible | Chat Completions | Generic compatible chat endpoints |
@ai-sdk/openai | Responses by default; .chat() selects Chat | OpenAI provider features or an explicit Chat route |
@ai-sdk/open-responses | Full Responses endpoint URL | The required Open Responses contract |
See the generic adapter source, OpenAI provider guide and Open Responses provider guide. The generic adapter has no .responses() language-model factory.
Set supportsStructuredOutputs: true for a schema-capable model. It tells the adapter what to request; server support must already exist. For streaming, includeUsage: true requests stream_options.include_usage. Both settings are covered in Vercel's provider reference.
Where compatibility breaks
Tools: test the return trip
Your application validates the function's JSON arguments, executes an authorized operation and returns the result with the matching tool-call ID. The provider then needs to accept that conversation on the next request. OpenAI documents the tool-message shape and streamed argument fields.
Test auto, required, named selection, strict arguments and parallel calls. Groq separates parallel support by model. DeepSeek rejects forced choices in thinking mode. Anthropic's direct layer ignores strictness.
For self-hosting, vLLM automatic selection needs a matching parser/template; named and required selection use structured decoding. LM Studio describes malformed calls appearing in message content. Rich tool results need another check: Vercel's compatible-message converter stringifies content-style tool results, even when user image inputs are supported.
JSON mode, JSON Schema and strict tools are different
| Feature | Chat Completions setting | What it constrains |
|---|---|---|
| JSON object mode | response_format.type="json_object" | JSON syntax |
| JSON Schema output | response_format.type="json_schema" | Assistant output shape within a supported schema subset |
| Strict function arguments | tools[].function.strict=true | Tool arguments |
These distinctions follow the OpenAI Chat reference and Azure structured-output guide.
HTTP 200 does not establish enforcement. Anthropic ignores response_format. DeepSeek's Chat enum permits only text and json_object; its strict tool schemas use a beta base.
Also test combinations and keywords. Cerebras disallows JSON object mode with streaming. Fireworks documents schema restrictions. Validate output locally and handle refusal or truncation before using it. Our structured-output testing article covers the gap between valid shape and correct content.
Streaming: guard empty choices and missing usage
OpenAI permits a final usage chunk with an empty choices array. Interrupted streams might never deliver final usage. Both behaviors appear in the streaming reference.
Reuse the Python client above for this text-only loop, following our streaming guide:
stream = client.chat.completions.create(
model=os.environ["LLM_MODEL"],
messages=[{"role": "user", "content": "Write two sentences about HTTP."}],
stream=True,
stream_options={"include_usage": True},
)
usage = None
for chunk in stream:
if chunk.choices:
text = chunk.choices[0].delta.content
if text:
print(text, end="", flush=True)
if chunk.usage is not None:
usage = chunk.usage
print()
print("Usage:", usage)Keep missing usage separate from zero tokens. DeepSeek and Fireworks report usage on a final finish chunk. Cached-token and reasoning fields also differ, as Together documents.
For streamed tools, accumulate argument fragments by index and ID before parsing. The text loop above does not collect tool arguments.
Vision and reasoning: test the exact format
Cerebras accepts base64 PNG/JPEG on selected models, but rejects external HTTPS image URLs. Ollama's local checklist checks base64 support and leaves remote URLs unchecked. Test your intended transport, MIME type and content shape.
Reasoning settings also differ. Gemini prohibits combining reasoning_effort with its thinking controls. Anthropic ignores reasoning_effort in its direct layer. Preserve provider-required reasoning or signature fields when replaying tool conversations.
Choose a native SDK when you need native content types, caching, reasoning or hosted tools outside the compatible surface.
Chat Completions does not imply Responses support
Responses changes the request and response contract. OpenAI recommends trying Responses for new projects; Chat Completions remains supported in its Node SDK.
| Contract | Chat Completions | Responses |
|---|---|---|
| Relative route | /chat/completions | /responses |
| Input | messages | input and optional instructions |
| Output | choices[].message | Typed output[] items |
| Function definition | Nested under tools[].function | Name and parameters under tools[] |
| Output schema | response_format.json_schema | text.format |
| Stream | Choice deltas | Semantic event types |
See the OpenAI Chat contract, Requesty Responses reference and Azure schema examples.
With the Python client above, a basic Requesty Responses call is:
response = client.responses.create(
model=os.environ["LLM_MODEL"],
input="Explain an API in one sentence.",
)
print(response.output_text)Stored continuation needs its own test. DeepSeek and Ollama implement stateless Responses. Fireworks supports previous_response_id. Check storage semantics and hosted tools on the selected route as well.
Open Responses defines an open specification based on Responses and provides compliance tests. A generic compatibility label does not establish compliance with that specification.
Test compatibility before switching providers
Start with the requests your application sends. A drafting tool needs text. An agent needs tool replay. A document extractor needs validated output and image handling. Their acceptance tests should reflect those jobs.

| Figure label | Example and primary source |
|---|---|
| A: Wrong route | The SDK builds the operation URL from a prefix: OpenAI route generation and Requesty setup. |
| B: Empty choices crash | Valid final usage chunks have empty choices; unguarded client code crashes. |
| C: Lost tool-call ID | Tool results need the matching tool_call_id. |
| D: Ignored format | Anthropic's direct layer ignores response_format. |
| E: Unsupported image URL | Cerebras requires base64 PNG/JPEG instead of external HTTPS URLs. |
| F: No stored continuation | DeepSeek Responses ignores unsupported previous_response_id. |
Run a smoke test
Use the exported variables and Python installation above. Create a synthetic PNG, such as a red circle on white, save it as probe.png, and export its path:
export LLM_IMAGE_PATH="./probe.png"Save the script as compatibility_probe.py and run python compatibility_probe.py. It makes billable requests for chat, streaming, JSON object output, JSON Schema, a tool round trip and vision. Run without Python's optimization flag so assertions remain enabled.
Copy the Python compatibility smoke test
import base64
import json
import os
import secrets
from pathlib import Path
from openai import OpenAI
client = OpenAI(base_url=os.environ["LLM_BASE_URL"],
api_key=os.environ["LLM_API_KEY"], max_retries=0, timeout=60)
model = os.environ["LLM_MODEL"]
def request(messages, **options):
return client.chat.completions.create(model=model, messages=messages, **options)
def prompt(text, **options):
return request([{"role": "user", "content": text}], **options)
def chat():
r = prompt("Say hello.")
assert r.choices and r.choices[0].message.role == "assistant"
assert r.choices[0].message.content
def streaming():
chunks = prompt("Say hello.", stream=True,
stream_options={"include_usage": True})
text, usage, finished = [], None, False
for chunk in chunks:
for choice in chunk.choices:
if choice.delta.content:
text.append(choice.delta.content)
finished |= choice.finish_reason is not None
if chunk.usage is not None:
usage = chunk.usage
assert text and finished and usage is not None
def json_object():
r = prompt("Return a JSON object with value 7.",
response_format={"type": "json_object"})
assert r.choices[0].finish_reason != "length"
assert isinstance(json.loads(r.choices[0].message.content or ""), dict)
def schema_output():
schema = {"type": "object", "properties": {"value": {"type": "integer"}},
"required": ["value"], "additionalProperties": False}
r = prompt("Return the integer 7 as JSON.", response_format={
"type": "json_schema", "json_schema": {
"name": "probe", "strict": True, "schema": schema}})
assert r.choices[0].finish_reason != "length"
value = json.loads(r.choices[0].message.content or "")
assert isinstance(value, dict) and set(value) == {"value"}
assert type(value["value"]) is int
def tools():
definitions = [{"type": "function", "function": {
"name": "get_marker", "description": "Get the marker needed to answer.",
"parameters": {"type": "object", "properties": {}, "required": [],
"additionalProperties": False}}}]
messages = [{"role": "user", "content": "Use get_marker; return its exact marker."}]
assistant = request(messages, tools=definitions).choices[0].message
assert assistant.tool_calls, "No tool call under automatic selection"
messages.append(assistant.model_dump(exclude_none=True))
marker = "probe-" + secrets.token_hex(8)
for call in assistant.tool_calls:
assert call.type == "function" and call.id
assert call.function.name == "get_marker"
assert json.loads(call.function.arguments) == {}
messages.append({"role": "tool", "tool_call_id": call.id,
"content": json.dumps({"marker": marker})})
assert marker in (request(messages).choices[0].message.content or "")
def vision():
image = base64.b64encode(Path(os.environ["LLM_IMAGE_PATH"]).read_bytes()).decode()
r = request([{"role": "user", "content": [
{"type": "text", "text": "Describe the image in one sentence."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64," + image}}
]}])
assert r.choices[0].message.content
print("Image description:", r.choices[0].message.content)
for test in [chat, streaming, json_object, schema_output, tools, vision]:
try:
test()
print("PASS:", test.__name__)
except Exception as error:
print("FAIL:", test.__name__, type(error).__name__, str(error))The request and stream fields follow the OpenAI Chat contract and stream schema. Client settings follow the Python constructor; the data URI follows our image-input guide.
We ran this probe through Requesty on September 28, 2026 against openai/gpt-6-luna, google/gemini-2.5-flash and mistral/mistral-small-latest, and all six checks passed. In the same session, openai/gpt-4.1-mini named the wrong color for a solid-color test image in two of three tries, which is why you should read the image description rather than trust a non-empty answer.
Read PASS as the result of that probe. One valid object does not prove strict schema enforcement. Automatic tool selection also tests model behavior; add required and named-selection probes where supported. Inspect the image description against the synthetic image, then use a JSON Schema validator and representative task evaluations. A failed advanced-feature check leaves basic Chat support as a separate result.
Check the production contract
OpenAI compatibility acceptance checklist
0 of 9 done
Troubleshoot by symptom
| Symptom | First check |
|---|---|
| 401 or 403 | Destination key, model permissions, workspace or deployment access |
| 404 | Final URL, duplicated route suffix, wrong API surface or model ID |
| 400 | Unsupported fields, schema keywords or parameter combinations |
| 429 | Provider quota, concurrency and retry policy |
| HTTP 200, setting has no effect | Accepted-but-ignored parameters; validate output |
| Stream crashes | Empty choices, optional delta fields or incomplete tool arguments |
| Missing usage | Opt-in setting, final-chunk format or interruption |
| Tool call appears as prose | Model capability, parser/template or prompting mode |
| Valid response, worse answers | Prompt and task-quality evaluation |
Match errors by HTTP status and documented codes, as Together recommends. For quota failures, use our 429 causes and retry guide. Protocol checks and provider-switching evaluations answer different questions.
When a gateway helps
We provide one OpenAI-compatible endpoint and an Anthropic-compatible endpoint for 600+ models. The same client code reaches providers through one Requesty key. You also get fallbacks when a provider fails, monthly spend caps and per-request logs alongside unified usage analytics.
Use https://router.requesty.ai/v1 for the OpenAI SDK or https://router.requesty.ai for the Anthropic SDK. For the coding-agent setup, see Claude Code with an API key.
Inspect the catalog capability fields before selecting models: supports_tool_calling, supports_vision, supports_reasoning, supports_output_json_object and supports_output_json_schema. This read-only query lists approved chat models advertising schema output and requires jq:
curl https://router.requesty.ai/v1/models \
-H "Authorization: Bearer $LLM_API_KEY" \
| jq '.data[] | select(.supports_output_json_schema == true) | .id'Authenticated discovery reflects organization and key approvals; /v1/models covers chat models. Choose a fallback chain and its order, then call policy/your-policy-name. Our routing-policy guide covers that setup.
Provider-specific features still vary by model and route. Test the features you depend on across every model in your fallback chain.
Follow our quickstart to configure the endpoint and key, then run your acceptance tests. Create a Requesty account to add models, fallbacks and spend controls.
Sources
Provider sources are linked in the matrix and beside each limitation.
Frequently asked questions
- What is an OpenAI-compatible API?
- An OpenAI-compatible API accepts OpenAI-style requests and returns responses that OpenAI clients can parse, most often for Chat Completions. Compatibility applies to specific endpoints and features, not necessarily the whole OpenAI platform.
- Is an OpenAI-compatible API the same as OpenAI's API?
- Another provider exposes a similar interface while using its own models, billing, data policies and parameter behavior. Your credentials must belong to the endpoint you call.
- Which providers offer OpenAI-compatible endpoints?
- Documented options include Gemini, Anthropic, Mistral, Groq, Together, Fireworks, DeepSeek, xAI, Perplexity, Cerebras, Azure OpenAI and Amazon Bedrock. Ollama, vLLM and LM Studio provide local or self-hosted options, while Requesty provides a multi-provider gateway. Feature coverage varies by endpoint and model.
- How do I use an OpenAI-compatible endpoint with the OpenAI SDK?
- Set the Python client's base_url or the JavaScript client's baseURL, supply the destination provider's API key, and select its model or deployment ID. Use the documented URL prefix, not the full Chat Completions route. Some authentication configurations also require headers or token providers.
- Are Anthropic and Gemini OpenAI-compatible?
- Both document OpenAI SDK compatibility layers. Gemini's layer is in beta; Anthropic recommends its layer primarily for testing and ignores response_format and function strictness. Native APIs provide features these layers do not expose.
- Does OpenAI compatibility include the Responses API?
- Not automatically. Check Responses support separately, including streaming events, stored continuation, schema fields and built-in tools. A provider can implement Responses without supporting previous_response_id.
- What is @ai-sdk/openai-compatible?
- It is Vercel's adapter for OpenAI-shaped APIs. Configure createOpenAICompatible with a name, baseURL and any required credentials; its language-model factory calls Chat Completions, not Responses.
- SEP '26
429 Too Many Requests: Causes and Fixes for LLM APIs
Fix 429 too many requests errors with provider-specific headers, billing checks, Retry-After parsing, bounded Python and TypeScript retries, and fallbacks.
- JAN '25
Switching LLM Providers: Why It’s Harder Than It Seems
- MAY '26
Structured Outputs Across LLM Providers: 244 Models Tested (2026)
We tested structured outputs (json_schema) across 244 models, 23 providers, 3 endpoints, and 10 popular SDKs. Over 2,400 tests. Here is the full compatibility matrix, what works, what breaks, and how to get consistent JSON schema enforcement without rewriting your integration for every model.
