Requesty
Back|SEP '26AI GATEWAY / INTEGRATIONS
17 MIN READ|

OpenAI-Compatible APIs: Meaning, Providers and Examples

Last updated

An OpenAI-compatible API accepts OpenAI-style requests and returns responses that OpenAI clients can parse, most often for Chat Completions. Reuse the OpenAI SDK by changing the base URL, API key and model ID. Tools, JSON schemas, streaming metadata and Responses support vary by endpoint and model. Compatibility describes the interface, as the Modular handbook explains.

The useful question is: “Does this endpoint support the requests my application sends?”

What OpenAI compatibility means

For a basic chat request, your client sends a model identifier and a messages array to a Chat Completions route. It expects an assistant reply under choices[].message. Streaming returns Server-Sent Events (SSE) containing choices[].delta. These are parts of the OpenAI Chat Completions contract.

A compatible provider implements that contract for the features it supports. Embeddings, model discovery, images, audio and Responses are separate API capabilities.

The same OpenAI Python client uses different base URLs, API keys and model IDs for OpenAI and Requesty. The messages and basic Chat Completions response shape stay the same.
Two constructor settings change, plus the model identifier in the request. Check advanced features separately. Sources: OpenAI SDK and Requesty quickstart.

Base URL means the prefix, not the full operation URL. With Requesty, set base_url="https://router.requesty.ai/v1"; the SDK appends /chat/completions. Gemini's prefix ends in /v1beta/openai/. DeepSeek uses https://api.deepseek.com. Follow the Requesty quickstart, Gemini setup or DeepSeek setup rather than adding /v1 everywhere.

The destination determines credentials, billing and data policies. Compatibility does not transfer an OpenAI API balance or ChatGPT subscription to another provider. A shared payload format does not imply the same answer quality. Even OpenAI documents different parameter support between its own models.

Which providers support OpenAI-compatible APIs?

Search the matrix by provider, filter by category or require the features your application uses. Expand a row for its base URL, restrictions and sources. Tools means developer-defined function calling; provider-hosted search and code execution need separate checks.

StatusMeaning
DocumentedOfficial docs describe implemented support
Model-dependentRequires a suitable model, deployment or server configuration
LimitedImplemented with documented restrictions
UnsupportedThe documented contract excludes the feature
IgnoredAccepted without the requested behavior
Not documentedReviewed docs do not establish support

OpenAI-compatible endpoint reference

Documentation reviewed September 28, 2026.

Needs:
17 of 17 documentation rows

Model-dependent results need model or configuration checks. Limited results need restriction checks.

  • Gemini Developer APIDocsBeta
    Hosted

    Google's compatibility layer is in beta.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://generativelanguage.googleapis.com/v1beta/openai/

    Google's compatibility layer is in beta. Schema output and base64 image input are documented. reasoning_effort cannot be combined with Google's thinking_level or thinking_budget.

    Chat stream
    Documented
    Function tools
    Documented
    JSON object
    Not documented
    JSON Schema
    Model-dependent
    Vision
    Model-dependent
    Responses
    Not documented
  • Anthropic direct layerDocs
    Hosted

    Anthropic recommends this layer primarily for testing, not most production use.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.anthropic.com/v1/

    Native Claude schema support does not change this layer's ignored response_format behavior.

    Anthropic recommends this layer primarily for testing, not most production use. response_format, function strict, logprobs and reasoning_effort are ignored; n must be 1. Images use image_url.url; detail is ignored. System and developer messages are concatenated and hoisted to the start. Applicable multi-workspace personal/service-account keys require anthropic-workspace-id via default_headers in Python or defaultHeaders in JavaScript.

    Chat stream
    Documented
    Function tools
    Limited
    JSON object
    Ignored
    JSON Schema
    Ignored
    Vision
    Limited
    Responses
    Not documented
  • MistralDocs
    Hosted

    Chat documents both JSON modes, named tool selection and parallel_tool_calls.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.mistral.ai/v1

    Chat documents both JSON modes, named tool selection and parallel_tool_calls. In JSON object mode, instruct the model to output JSON. Vision examples use a string image_url; test your client's content shape.

    Chat stream
    Documented
    Function tools
    Documented
    JSON object
    Documented
    JSON Schema
    Documented
    Vision
    Model-dependent
    Responses
    Not documented
  • GroqDocs
    Hosted

    Parallel tools and strict schema output vary by model.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.groq.com/openai/v1

    Check previous_response_id behavior separately before relying on Responses state.

    Parallel tools and strict schema output vary by model. Supplied logprobs, logit_bias, top_logprobs and messages[].name cause HTTP 400; n must be 1. Responses is in beta and its documentation conflicts on stored continuation.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Documented
    JSON Schema
    Model-dependent
    Vision
    Model-dependent
    Responses
    Limited
  • TogetherDocs
    Hosted

    service_tier, store, metadata and prediction are accepted but ignored.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.together.ai/v1

    service_tier, store, metadata and prediction are accepted but ignored. n support varies and logprobs uses Together's shape. Vision models accept remote URLs and data URIs; detail is ignored. Cached-token and reasoning fields have provider-specific locations. Match errors by HTTP status and documented codes.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Model-dependent
    JSON Schema
    Model-dependent
    Vision
    Model-dependent
    Responses
    Not documented
  • FireworksDocs
    Hosted

    Chat usage arrives on the final finish chunk.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.fireworks.ai/inference/v1

    Chat usage arrives on the final finish chunk. Default context-overflow behavior reduces the output budget; select error behavior if required. Schema output excludes external references and has regex restrictions. Responses supports stored continuation and store=false; background=true cannot combine with stream=true.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Documented
    JSON Schema
    Limited
    Vision
    Model-dependent
    Responses
    Documented
  • DeepSeekDocs
    Hosted

    Chat response_format allows only text and json_object.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.deepseek.com

    Strict tool arguments use the separate beta base https://api.deepseek.com/beta. They do not implement Chat assistant JSON Schema output.

    Chat response_format allows only text and json_object. Thinking mode rejects required and named tool choices. Vision is documented for deepseek-flash. Chat usage arrives on the final finish chunk. Responses is stateless; previous_response_id and other unsupported parameters are ignored.

    Chat stream
    Documented
    Function tools
    Limited
    JSON object
    Documented
    JSON Schema
    Unsupported
    Vision
    Model-dependent
    Responses
    Limited
  • xAIDocs
    Hosted

    Schema output has best-effort keywords; oneOf behaves like anyOf.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.x.ai/v1

    Schema output has best-effort keywords; oneOf behaves like anyOf. Function calling supports named and parallel selection. Responses supports stored continuation; background is a compatibility field without implemented background execution.

    Chat stream
    Documented
    Function tools
    Documented
    JSON object
    Documented
    JSON Schema
    Limited
    Vision
    Model-dependent
    Responses
    Documented
  • Perplexity Agent APIDocs
    Hosted · Responses interface

    Agent documents Responses streaming and /v1/responses as an alias for /v1/agent.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.perplexity.ai/v1

    Agent features do not establish Chat Completions support or OpenAI Responses text.format mapping. Router is a separate product.

    Agent documents Responses streaming and /v1/responses as an alias for /v1/agent. Agent custom functions, image input and schema output belong to that surface. Agent schema output uses provider response_format. The September 27, 2026 Sonar transition preserves synchronous and streaming requests through reformulation as Agent requests; asynchronous Sonar requests ended.

    Chat stream
    Not documented
    Function tools
    Not documented
    JSON object
    Not documented
    JSON Schema
    Not documented
    Vision
    Not documented
    Responses
    Documented
  • Perplexity Router APIDocsPrivate preview
    Hosted

    Router is separate from Agent and is in private preview.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.perplexity.ai/router/v1

    Router is separate from Agent and is in private preview. Chat schemas are always strict. Several compatibility fields support only defaults; unknown fields cause HTTP 400. The schema includes image_url; select an image-capable model. Router Responses is stateless.

    Chat stream
    Documented
    Function tools
    Documented
    JSON object
    Documented
    JSON Schema
    Documented
    Vision
    Model-dependent
    Responses
    Limited
  • CerebrasDocs
    Hosted

    JSON object mode cannot stream; n must be 1.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://api.cerebras.ai/v1

    JSON object mode cannot stream; n must be 1. Selected vision models accept base64 PNG/JPEG, not external HTTPS URLs or image_url.detail. gpt-oss-120b rejects tools plus response_format together.

    Chat stream
    Documented
    Function tools
    Documented
    JSON object
    Limited
    JSON Schema
    Model-dependent
    Vision
    Limited
    Responses
    Not documented
  • Azure OpenAI v1Docs
    Cloud

    Azure v1 uses OpenAI(), without a required dated api-version parameter.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://YOUR-RESOURCE-NAME.openai.azure.com/openai/v1/

    Azure v1 uses OpenAI(), without a required dated api-version parameter. Replace the resource placeholder; model is your deployment name. Authentication supports a key or Microsoft Entra token provider. Disable parallel tools for strict structured function use.

    Chat stream
    Model-dependent
    Function tools
    Model-dependent
    JSON object
    Documented
    JSON Schema
    Model-dependent
    Vision
    Model-dependent
    Responses
    Documented
  • Amazon Bedrock runtimeDocs
    Cloud

    AWS recommends runtime for new applications.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://bedrock-runtime.REGION.amazonaws.com/openai/v1

    Native Converse schema or vision support does not establish OpenAI-compatible response_format or image coverage.

    AWS recommends runtime for new applications. Replace REGION and check the selected model's API coverage. Ordinary OpenAI SDK setup uses a Bedrock API key. Runtime does not implement compatible GET /models; use AWS discovery operations and model cards. GPT OSS Responses requires mantle; verify the model-specific endpoint path.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Not documented
    JSON Schema
    Not documented
    Vision
    Not documented
    Responses
    Model-dependent
  • Ollama localDocs
    Local / self-hosted

    Local Ollama ignores the key value required by the client.

    Base URL, caveats and sources
    OpenAI SDK base URL
    http://localhost:11434/v1/

    Local Ollama ignores the key value required by the client. Load a suitable model for tools and vision. tool_choice and remote image URLs are unchecked in the published compatibility checklist. Responses is stateless. Ollama Cloud authentication differs and its structured-output support differs from local.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Documented
    JSON Schema
    Documented
    Vision
    Model-dependent
    Responses
    Limited
  • vLLMDocs
    Local / self-hosted

    The address is illustrative; confirm your server host and port.

    Base URL, caveats and sources
    OpenAI SDK base URL
    http://localhost:8000/v1

    The address is illustrative; confirm your server host and port. Chat needs a chat template. Automatic tool choice needs a matching parser/template and server flags; named and required selection use structured decoding, enabled by default. Schema decoding is backend-dependent; image detail is unsupported. Match docs to the installed version. API-key settings do not protect every HTTP route; review network-hardening guidance before exposure.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Documented
    JSON Schema
    Model-dependent
    Vision
    Model-dependent
    Responses
    Documented
  • LM StudioDocs
    Local / self-hosted

    Enable the server and load a suitable model.

    Base URL, caveats and sources
    OpenAI SDK base URL
    http://localhost:1234/v1

    Enable the server and load a suitable model. Tools use native or default prompting modes; malformed tool output can appear in message.content instead of message.tool_calls. Responses supports prior-response state.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Not documented
    JSON Schema
    Model-dependent
    Vision
    Model-dependent
    Responses
    Documented
  • RequestyDocs
    Gateway

    One gateway interface reaches multiple providers.

    Base URL, caveats and sources
    OpenAI SDK base URL
    https://router.requesty.ai/v1

    One gateway interface reaches multiple providers. Select models using live catalog capability flags. Translated and native Responses routes differ in state, file, tool and event behavior.

    Chat stream
    Documented
    Function tools
    Model-dependent
    JSON object
    Model-dependent
    JSON Schema
    Model-dependent
    Vision
    Model-dependent
    Responses
    Limited
Streaming, tools, JSON object, JSON Schema and vision columns describe Chat Completions. Responses is a separate endpoint check; model and parameter combinations still matter.

Perplexity Agent's custom functions, schema output and image inputs use its Agent surface. Its Sonar migration notice explains the September 27 transition. Keep that surface separate from Router when configuring a client.

Use an OpenAI-compatible endpoint

These examples use Requesty and openai/gpt-6-luna. Replace the URL, key and model to use another provider. Check current identifiers in the Models API reference.

Export these values; the examples read the LLM_* variables explicitly:

Shell
export LLM_BASE_URL="https://router.requesty.ai/v1"
export LLM_API_KEY="your-requesty-key"
export LLM_MODEL="openai/gpt-6-luna"

The OpenAI Python SDK requires Python 3.10 or newer. Save this as example.py and run python example.py.

Shell
python -m pip install openai
example.py
import os
from openai import OpenAI
 
client = OpenAI(
    base_url=os.environ["LLM_BASE_URL"],
    api_key=os.environ["LLM_API_KEY"],
)
response = client.chat.completions.create(
    model=os.environ["LLM_MODEL"],
    messages=[{"role": "user", "content": "Explain an API in one sentence."}],
)
print(response.choices[0].message.content)

Run the TypeScript files with npx tsx; they use top-level await, so the project needs "type": "module" in package.json.

Use @ai-sdk/openai-compatible

Vercel's OpenAI Compatible provider wraps an OpenAI-shaped API for the AI SDK. Supply name, baseURL and credentials. Save this as ai-sdk-example.ts and run npx tsx ./ai-sdk-example.ts:

Shell
npm install ai @ai-sdk/openai-compatible
ai-sdk-example.ts
import { createOpenAICompatible } from '@ai-sdk/openai-compatible';
import { generateText } from 'ai';
 
const baseURL = process.env.LLM_BASE_URL;
const apiKey = process.env.LLM_API_KEY;
const model = process.env.LLM_MODEL;
if (!baseURL || !apiKey || !model) {
  throw new Error('Set LLM_BASE_URL, LLM_API_KEY, and LLM_MODEL');
}
 
const provider = createOpenAICompatible({ name: 'custom', baseURL, apiKey });
const { text } = await generateText({
  model: provider(model),
  prompt: 'Explain an API in one sentence.',
});
console.log(text);

The provider interface requires name and baseURL. Choosing the adapter also chooses an API surface:

IntegrationLanguage-model routeUse it for
@ai-sdk/openai-compatibleChat CompletionsGeneric compatible chat endpoints
@ai-sdk/openaiResponses by default; .chat() selects ChatOpenAI provider features or an explicit Chat route
@ai-sdk/open-responsesFull Responses endpoint URLThe required Open Responses contract

See the generic adapter source, OpenAI provider guide and Open Responses provider guide. The generic adapter has no .responses() language-model factory.

Set supportsStructuredOutputs: true for a schema-capable model. It tells the adapter what to request; server support must already exist. For streaming, includeUsage: true requests stream_options.include_usage. Both settings are covered in Vercel's provider reference.

Where compatibility breaks

Tools: test the return trip

Your application validates the function's JSON arguments, executes an authorized operation and returns the result with the matching tool-call ID. The provider then needs to accept that conversation on the next request. OpenAI documents the tool-message shape and streamed argument fields.

Test auto, required, named selection, strict arguments and parallel calls. Groq separates parallel support by model. DeepSeek rejects forced choices in thinking mode. Anthropic's direct layer ignores strictness.

For self-hosting, vLLM automatic selection needs a matching parser/template; named and required selection use structured decoding. LM Studio describes malformed calls appearing in message content. Rich tool results need another check: Vercel's compatible-message converter stringifies content-style tool results, even when user image inputs are supported.

JSON mode, JSON Schema and strict tools are different

FeatureChat Completions settingWhat it constrains
JSON object moderesponse_format.type="json_object"JSON syntax
JSON Schema outputresponse_format.type="json_schema"Assistant output shape within a supported schema subset
Strict function argumentstools[].function.strict=trueTool arguments

These distinctions follow the OpenAI Chat reference and Azure structured-output guide.

HTTP 200 does not establish enforcement. Anthropic ignores response_format. DeepSeek's Chat enum permits only text and json_object; its strict tool schemas use a beta base.

Also test combinations and keywords. Cerebras disallows JSON object mode with streaming. Fireworks documents schema restrictions. Validate output locally and handle refusal or truncation before using it. Our structured-output testing article covers the gap between valid shape and correct content.

Streaming: guard empty choices and missing usage

OpenAI permits a final usage chunk with an empty choices array. Interrupted streams might never deliver final usage. Both behaviors appear in the streaming reference.

Reuse the Python client above for this text-only loop, following our streaming guide:

Python
stream = client.chat.completions.create(
    model=os.environ["LLM_MODEL"],
    messages=[{"role": "user", "content": "Write two sentences about HTTP."}],
    stream=True,
    stream_options={"include_usage": True},
)
usage = None
for chunk in stream:
    if chunk.choices:
        text = chunk.choices[0].delta.content
        if text:
            print(text, end="", flush=True)
    if chunk.usage is not None:
        usage = chunk.usage
print()
print("Usage:", usage)

Keep missing usage separate from zero tokens. DeepSeek and Fireworks report usage on a final finish chunk. Cached-token and reasoning fields also differ, as Together documents.

For streamed tools, accumulate argument fragments by index and ID before parsing. The text loop above does not collect tool arguments.

Vision and reasoning: test the exact format

Cerebras accepts base64 PNG/JPEG on selected models, but rejects external HTTPS image URLs. Ollama's local checklist checks base64 support and leaves remote URLs unchecked. Test your intended transport, MIME type and content shape.

Reasoning settings also differ. Gemini prohibits combining reasoning_effort with its thinking controls. Anthropic ignores reasoning_effort in its direct layer. Preserve provider-required reasoning or signature fields when replaying tool conversations.

Choose a native SDK when you need native content types, caching, reasoning or hosted tools outside the compatible surface.

Chat Completions does not imply Responses support

Responses changes the request and response contract. OpenAI recommends trying Responses for new projects; Chat Completions remains supported in its Node SDK.

ContractChat CompletionsResponses
Relative route/chat/completions/responses
Inputmessagesinput and optional instructions
Outputchoices[].messageTyped output[] items
Function definitionNested under tools[].functionName and parameters under tools[]
Output schemaresponse_format.json_schematext.format
StreamChoice deltasSemantic event types

See the OpenAI Chat contract, Requesty Responses reference and Azure schema examples.

With the Python client above, a basic Requesty Responses call is:

Python
response = client.responses.create(
    model=os.environ["LLM_MODEL"],
    input="Explain an API in one sentence.",
)
print(response.output_text)

Stored continuation needs its own test. DeepSeek and Ollama implement stateless Responses. Fireworks supports previous_response_id. Check storage semantics and hosted tools on the selected route as well.

Open Responses defines an open specification based on Responses and provides compliance tests. A generic compatibility label does not establish compliance with that specification.

Test compatibility before switching providers

Start with the requests your application sends. A drafting tool needs text. An agent needs tool replay. A document extractor needs validated output and image handling. Their acceptance tests should reflect those jobs.

Six protocol checks labeled A to F: text, stream and usage, tool round trip, JSON constraints, image inputs, and Responses behavior. A separate task-quality track checks accepted results rather than protocol compatibility.
Choose the checks your application needs. Example failures A to F have primary-source links below.
Figure labelExample and primary source
A: Wrong routeThe SDK builds the operation URL from a prefix: OpenAI route generation and Requesty setup.
B: Empty choices crashValid final usage chunks have empty choices; unguarded client code crashes.
C: Lost tool-call IDTool results need the matching tool_call_id.
D: Ignored formatAnthropic's direct layer ignores response_format.
E: Unsupported image URLCerebras requires base64 PNG/JPEG instead of external HTTPS URLs.
F: No stored continuationDeepSeek Responses ignores unsupported previous_response_id.

Run a smoke test

Use the exported variables and Python installation above. Create a synthetic PNG, such as a red circle on white, save it as probe.png, and export its path:

Shell
export LLM_IMAGE_PATH="./probe.png"

Save the script as compatibility_probe.py and run python compatibility_probe.py. It makes billable requests for chat, streaming, JSON object output, JSON Schema, a tool round trip and vision. Run without Python's optimization flag so assertions remain enabled.

Copy the Python compatibility smoke test
compatibility_probe.py
import base64
import json
import os
import secrets
from pathlib import Path
from openai import OpenAI
 
client = OpenAI(base_url=os.environ["LLM_BASE_URL"],
                api_key=os.environ["LLM_API_KEY"], max_retries=0, timeout=60)
model = os.environ["LLM_MODEL"]
 
 
def request(messages, **options):
    return client.chat.completions.create(model=model, messages=messages, **options)
 
 
def prompt(text, **options):
    return request([{"role": "user", "content": text}], **options)
 
 
def chat():
    r = prompt("Say hello.")
    assert r.choices and r.choices[0].message.role == "assistant"
    assert r.choices[0].message.content
 
 
def streaming():
    chunks = prompt("Say hello.", stream=True,
                    stream_options={"include_usage": True})
    text, usage, finished = [], None, False
    for chunk in chunks:
        for choice in chunk.choices:
            if choice.delta.content:
                text.append(choice.delta.content)
            finished |= choice.finish_reason is not None
        if chunk.usage is not None:
            usage = chunk.usage
    assert text and finished and usage is not None
 
 
def json_object():
    r = prompt("Return a JSON object with value 7.",
               response_format={"type": "json_object"})
    assert r.choices[0].finish_reason != "length"
    assert isinstance(json.loads(r.choices[0].message.content or ""), dict)
 
 
def schema_output():
    schema = {"type": "object", "properties": {"value": {"type": "integer"}},
              "required": ["value"], "additionalProperties": False}
    r = prompt("Return the integer 7 as JSON.", response_format={
        "type": "json_schema", "json_schema": {
            "name": "probe", "strict": True, "schema": schema}})
    assert r.choices[0].finish_reason != "length"
    value = json.loads(r.choices[0].message.content or "")
    assert isinstance(value, dict) and set(value) == {"value"}
    assert type(value["value"]) is int
 
 
def tools():
    definitions = [{"type": "function", "function": {
        "name": "get_marker", "description": "Get the marker needed to answer.",
        "parameters": {"type": "object", "properties": {}, "required": [],
                       "additionalProperties": False}}}]
    messages = [{"role": "user", "content": "Use get_marker; return its exact marker."}]
    assistant = request(messages, tools=definitions).choices[0].message
    assert assistant.tool_calls, "No tool call under automatic selection"
    messages.append(assistant.model_dump(exclude_none=True))
    marker = "probe-" + secrets.token_hex(8)
    for call in assistant.tool_calls:
        assert call.type == "function" and call.id
        assert call.function.name == "get_marker"
        assert json.loads(call.function.arguments) == {}
        messages.append({"role": "tool", "tool_call_id": call.id,
                         "content": json.dumps({"marker": marker})})
    assert marker in (request(messages).choices[0].message.content or "")
 
 
def vision():
    image = base64.b64encode(Path(os.environ["LLM_IMAGE_PATH"]).read_bytes()).decode()
    r = request([{"role": "user", "content": [
        {"type": "text", "text": "Describe the image in one sentence."},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64," + image}}
    ]}])
    assert r.choices[0].message.content
    print("Image description:", r.choices[0].message.content)
 
 
for test in [chat, streaming, json_object, schema_output, tools, vision]:
    try:
        test()
        print("PASS:", test.__name__)
    except Exception as error:
        print("FAIL:", test.__name__, type(error).__name__, str(error))

The request and stream fields follow the OpenAI Chat contract and stream schema. Client settings follow the Python constructor; the data URI follows our image-input guide.

We ran this probe through Requesty on September 28, 2026 against openai/gpt-6-luna, google/gemini-2.5-flash and mistral/mistral-small-latest, and all six checks passed. In the same session, openai/gpt-4.1-mini named the wrong color for a solid-color test image in two of three tries, which is why you should read the image description rather than trust a non-empty answer.

Read PASS as the result of that probe. One valid object does not prove strict schema enforcement. Automatic tool selection also tests model behavior; add required and named-selection probes where supported. Inspect the image description against the synthetic image, then use a JSON Schema validator and representative task evaluations. A failed advanced-feature check leaves basic Chat support as a separate result.

Check the production contract

OpenAI compatibility acceptance checklist

0 of 9 done

Endpoint
Streaming
Tools
Output
Multimodal
Reasoning
Operations
Responses
Quality

Troubleshoot by symptom

SymptomFirst check
401 or 403Destination key, model permissions, workspace or deployment access
404Final URL, duplicated route suffix, wrong API surface or model ID
400Unsupported fields, schema keywords or parameter combinations
429Provider quota, concurrency and retry policy
HTTP 200, setting has no effectAccepted-but-ignored parameters; validate output
Stream crashesEmpty choices, optional delta fields or incomplete tool arguments
Missing usageOpt-in setting, final-chunk format or interruption
Tool call appears as proseModel capability, parser/template or prompting mode
Valid response, worse answersPrompt and task-quality evaluation

Match errors by HTTP status and documented codes, as Together recommends. For quota failures, use our 429 causes and retry guide. Protocol checks and provider-switching evaluations answer different questions.

When a gateway helps

We provide one OpenAI-compatible endpoint and an Anthropic-compatible endpoint for 600+ models. The same client code reaches providers through one Requesty key. You also get fallbacks when a provider fails, monthly spend caps and per-request logs alongside unified usage analytics.

Use https://router.requesty.ai/v1 for the OpenAI SDK or https://router.requesty.ai for the Anthropic SDK. For the coding-agent setup, see Claude Code with an API key.

Inspect the catalog capability fields before selecting models: supports_tool_calling, supports_vision, supports_reasoning, supports_output_json_object and supports_output_json_schema. This read-only query lists approved chat models advertising schema output and requires jq:

Shell
curl https://router.requesty.ai/v1/models \
  -H "Authorization: Bearer $LLM_API_KEY" \
  | jq '.data[] | select(.supports_output_json_schema == true) | .id'

Authenticated discovery reflects organization and key approvals; /v1/models covers chat models. Choose a fallback chain and its order, then call policy/your-policy-name. Our routing-policy guide covers that setup.

Provider-specific features still vary by model and route. Test the features you depend on across every model in your fallback chain.

Requesty
Keep one integration across providers

Follow our quickstart to configure the endpoint and key, then run your acceptance tests. Create a Requesty account to add models, fallbacks and spend controls.

Sources

Provider sources are linked in the matrix and beside each limitation.

Frequently asked questions
What is an OpenAI-compatible API?
An OpenAI-compatible API accepts OpenAI-style requests and returns responses that OpenAI clients can parse, most often for Chat Completions. Compatibility applies to specific endpoints and features, not necessarily the whole OpenAI platform.
Is an OpenAI-compatible API the same as OpenAI's API?
Another provider exposes a similar interface while using its own models, billing, data policies and parameter behavior. Your credentials must belong to the endpoint you call.
Which providers offer OpenAI-compatible endpoints?
Documented options include Gemini, Anthropic, Mistral, Groq, Together, Fireworks, DeepSeek, xAI, Perplexity, Cerebras, Azure OpenAI and Amazon Bedrock. Ollama, vLLM and LM Studio provide local or self-hosted options, while Requesty provides a multi-provider gateway. Feature coverage varies by endpoint and model.
How do I use an OpenAI-compatible endpoint with the OpenAI SDK?
Set the Python client's base_url or the JavaScript client's baseURL, supply the destination provider's API key, and select its model or deployment ID. Use the documented URL prefix, not the full Chat Completions route. Some authentication configurations also require headers or token providers.
Are Anthropic and Gemini OpenAI-compatible?
Both document OpenAI SDK compatibility layers. Gemini's layer is in beta; Anthropic recommends its layer primarily for testing and ignores response_format and function strictness. Native APIs provide features these layers do not expose.
Does OpenAI compatibility include the Responses API?
Not automatically. Check Responses support separately, including streaming events, stored continuation, schema fields and built-in tools. A provider can implement Responses without supporting previous_response_id.
What is @ai-sdk/openai-compatible?
It is Vercel's adapter for OpenAI-shaped APIs. Configure createOpenAICompatible with a name, baseURL and any required credentials; its language-model factory calls Chat Completions, not Responses.
Related reading

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders