The two APIs look similar and the models are comparable. The differences that break your code when switching are in request structure, system-prompt handling, streaming, and how you count tokens.

On paper, swapping one LLM provider for another looks like changing a base URL and an API key. The request bodies are both JSON with a messages array; the models are broadly comparable. Then you deploy, and something silently misbehaves — a system prompt that stops taking effect, a streaming parser that returns nothing, a cost estimate that's off by 15%. The APIs are similar enough to lull you and different enough to bite. Here's where the switching bugs actually live.

1. Request structure

Both APIs take {"model": ..., "messages": [...]}, and that's where the comfortable similarity ends. On the Anthropic API, the system prompt is a top-level system field, not an entry in the messages array. And max_tokens is required on Anthropic while it's optional on OpenAI. Port code that omits max_tokens and the Anthropic request fails outright — an easy first stumble, at least a loud one.

2. The system prompt that vanishes

This is the quiet one. On the OpenAI API you put a system instruction as a message with role: "system" inside the array. Do that on the Anthropic API and it doesn't error — it just isn't treated as the system prompt, because Anthropic expects a top-level system field. Your instructions become an ordinary message or get dropped, the model's behaviour shifts, and nothing tells you why. When you switch, the system prompt is the first thing to move to its new home and the first thing to check when behaviour drifts.

3. Streaming formats differ

Both stream over Server-Sent Events, which sounds like a shared standard until you parse the payloads:

"Server-Sent Events (SSE) is a server push technology enabling a client to receive automatic updates from a server via an HTTP connection, and describes how servers can initiate data transmission towards clients once an initial client connection has been established."

— Wikipedia, "Server-sent events" (CC BY-SA 4.0)

The transport is shared; the event shapes are not. OpenAI emits deltas under a choices[].delta.content structure; Anthropic emits typed events like content_block_delta with the text under delta.text. A client-side streaming parser written for one provider will not read the other's stream — it needs a rewrite, not a config change. This catches teams who assumed "both use SSE" meant "same parser."

4. Tool use / function calling

Both support tool use, and both describe tool parameters with JSON Schema, which is the good news. The field names and response shapes differ — the concepts line up, the wire format doesn't. Expect to remap how tools are declared and how tool-call results are read back. Budget for it; don't assume the tool-use layer ports for free just because the schemas look familiar.

5. Tokens and cost aren't equivalent

Here's the one that quietly breaks your cost model. Both providers tokenize with BPE-style tokenizers, but with different vocabularies — so the same paragraph of text is a different number of tokens on each. That means a prompt costing a known amount on one provider does not cost the proportional amount on the other, even at identical per-token prices; the token count itself shifts, often by low double-digit percentages. Never assume equivalent cost from equivalent text. Measure the actual token count on each side before you trust a migration's budget.

6. Error responses and rate limits

Error handling is another place where the similarity is superficial. Both APIs return HTTP error codes, but the error body shapes differ — field names, nesting, and the level of detail in error messages are all provider-specific. Rate-limit responses follow the same pattern: both use HTTP 429, but the headers that tell you when you can retry (retry-after, rate-limit remaining counts) use different names and conventions. Client code that parses rate-limit headers to implement exponential backoff needs to be rewritten, not just reconfigured. And Anthropic's overloaded-model response (HTTP 529) has no direct equivalent in OpenAI's error taxonomy. If your retry logic assumes a fixed set of error codes, the switch will introduce unhandled cases.

7. Model naming and capability mapping

Model names don't map one-to-one across providers, and assuming they do leads to mismatched expectations. OpenAI's gpt-4o and Anthropic's claude-sonnet-4-6 are broadly comparable in capability class but differ in context window, output style, instruction-following behaviour, and pricing. Picking the "equivalent" model on the new provider is a judgement call, not a lookup table. Test with representative prompts before committing — model behaviour differences are often larger than API-plumbing differences, and they're harder to fix because there's no code change that makes one model behave identically to another. Budget time for prompt tuning as part of the migration, not just the API integration.

8. Thinking and extended reasoning

Anthropic's API supports an extended thinking feature where the model can reason step-by-step before producing a final answer, with the thinking process visible in the response. OpenAI's reasoning models (the o-series) handle this differently — reasoning happens internally and the response structure reflects that architecture. If your application uses extended thinking or chain-of-thought features, the migration isn't just a field rename; it's a different interaction pattern. Applications that display thinking steps to users, or that parse intermediate reasoning for monitoring, will need their UI and parsing logic rebuilt for the target provider's approach.

Translate, measure, inspect

Make the switch mechanical instead of error-prone. Run request payloads through an OpenAI ↔ Anthropic converter to move the system prompt to the right place and reshape the body correctly, rather than hand-editing and hoping. Before committing to a cost estimate, put representative prompts through a token counter so the budget reflects real token counts, not an assumption carried over from the old provider. And when a request or a streamed response isn't behaving, a JSON formatter makes the actual structure legible so you can see exactly where your payload diverges from what the API expects. The models are comparable; the plumbing isn't — treat the plumbing with respect and the switch is routine.

← All articles