GLM-5.3 Flash reads images and scores higher on reasoning and agent benchmarks. DeepSeek V4 Flash 0731 is text-only but starts answering about three times sooner on traffic routed through us. Both are open-weight, million-token flash models in our catalog, released within a month of each other, and both cost pennies per session after caching.
The short answer
Pick GLM-5.3 Flash for agent pipelines that need vision or stronger multi-step reasoning: it leads Artificial Analysis's Intelligence and Agentic indexes and blended to $0.06 per million tokens on our traffic. Pick DeepSeek V4 Flash 0731 for fast-start, text-only chat, where its 541 ms median first token wins. Plain code generation is close either way.
At a glance
Both are open-weight, million-token flash models in our catalog, released within a month of each other. The practical split: GLM-5.3 Flash is natively multimodal, so it accepts images on most of our endpoints, while DeepSeek V4 Flash 0731 is text-only but ships as a 304B MoE with speculative decoding aimed at fast, cheap inference and agentic tool use. Both support reasoning, tool calling and prompt caching across our endpoints.
At a glance
| GLM-5.3 Flash | DeepSeek V4 Flash 0731 | |
|---|---|---|
| Trained by | Z.AI | DeepSeek |
| Released | 2026-08-26 | 2026-07-31 |
| Context window | 1.049M tokens (1M on some endpoints) | 1.049M tokens (256K on some endpoints) |
| Max output | 1.049M tokens | 1.049M tokens |
| Providers on Requesty | zai, fireworks, novita, deepinfra, tensorx, runware, sference | fireworks, novita, deepinfra, tensorx, runware, scaleway, sail, sference |
| Image input | Yes | No |
| Reasoning controls | Yes | Yes |
| Prompt caching | Yes | Yes |
| Tool calling | Yes | Yes |
| Computer use | No | No |
| Web search | No | No |
| Structured output | Yes | Yes |
| Open weights | Yes | Yes |
| Data retention (first-party endpoint) | None | None |
| Trains on your data | No | No |
| Share of Requesty tokens | 2.8% | 9.1% |
Pricing and caching
Across traffic routed through us from July 6 to September 7, GLM-5.3 Flash blended to $0.06 per million tokens and DeepSeek V4 Flash 0731 to $0.07, with caching cutting bills by roughly 60% and 57% respectively (our model economics). List prices sit at $0.15 input and $0.50 output for GLM on zai/glm-5.3-flash versus $0.22 and $0.66 for DeepSeek on fireworks/deepseek-v4-flash-0731, though DeepSeek's $0.007 cached reads are far cheaper than GLM's $0.03. On DeepSeek's own API, peak-hour rates are double off-peak, with peak at 01:00-04:00 and 06:00-10:00 UTC since August 16, 2026 (DeepSeek changelog).
Pricing and caching
Our routed customer traffic · 2026-07-06 to 2026-09-07

▶Price breakdown and benchmark costs
▶Endpoint list prices
| Model | Endpoint | Input | Output | Cached input | Context |
|---|---|---|---|---|---|
| GLM-5.3 Flash | runware | $0.075 | $0.250 | $0.015 | 1.0M |
| GLM-5.3 Flash | zai | $0.150 | $0.500 | $0.030 | 1M |
| GLM-5.3 Flash | fireworks | $0.150 | $0.500 | $0.030 | 1.0M |
| GLM-5.3 Flash | novita | $0.150 | $0.500 | $0.030 | 1.0M |
| GLM-5.3 Flash | deepinfra | $0.150 | $0.500 | $0.030 | 1.0M |
| GLM-5.3 Flash | tensorx | $0.200 | $0.500 | $0.050 | 1.0M |
| GLM-5.3 Flash | sference | $0.200 | $0.600 | $0.070 | 1M |
| DeepSeek V4 Flash 0731 | deepinfra | $0.072 | $0.144 | $0.014 | 1.0M |
| DeepSeek V4 Flash 0731 | runware | $0.076 | $0.153 | $0.014 | 1.0M |
| DeepSeek V4 Flash 0731 | deepinfra | $0.090 | $0.180 | $0.018 | 1.0M |
| DeepSeek V4 Flash 0731 | sail | $0.090 | $0.180 | $0.020 | 1.0M |
| DeepSeek V4 Flash 0731 | novita | $0.140 | $0.280 | $0.028 | 1.0M |
| DeepSeek V4 Flash 0731 | tensorx | $0.250 | $0.300 | $0.060 | 1.0M |
| DeepSeek V4 Flash 0731 | sference | $0.280 | $0.560 | $0.070 | 1.0M |
| DeepSeek V4 Flash 0731 | fireworks | $0.220 | $0.660 | $0.0070 | 1M |
| DeepSeek V4 Flash 0731 | scaleway | $0.460 | $0.930 | $0.090 | 256K |
Benchmarks and use cases
On Artificial Analysis's version 4.3 indexes, GLM-5.3 Flash leads DeepSeek V4 Flash 0731 on the Intelligence Index (41.9 vs 34.5) and the Agentic Index (51.2 vs 41.7), with the DeepSeek figures measured in reasoning mode at max effort (GLM-5.3 Flash, DeepSeek V4 Flash). Coding is closer, 71.5 against 69.1, so for straightforward code generation either model is a reasonable pick. The wider gaps sit in general reasoning and multi-step agent work.
Benchmark scores
Artificial Analysis · source score units

Latency and throughput
Across traffic routed through us from July 6 to September 7, DeepSeek V4 Flash 0731 started responding sooner, with a median time to first token of 541 ms against 1,684 ms for GLM-5.3 Flash (model economics). Once streaming, GLM-5.3 Flash delivered tokens far faster at the median, and its 5xx error rate was lower, 0.4% versus 1.9%. Endpoint choice matters for both: our latency rankings show first-token times ranging from roughly 600 ms to well over 10 s depending on provider.
Latency and throughput
Our routed customer traffic · 2026-07-06 to 2026-09-07

▶Endpoint results
| Endpoint | First token | Output speed | Total |
|---|---|---|---|
| sference | 1,219 ms | n/a | 4.1 s |
| runware | 2,909 ms | n/a | 4.7 s |
| fireworks | 3,171 ms | n/a | 7.9 s |
| tensorx | 5,290 ms | n/a | 10.1 s |
| zai | 7,383 ms | n/a | 9.8 s |
| deepinfra | 8,607 ms | n/a | 10.5 s |
| novita | 101,992 ms | n/a | 102.1 s |
| Endpoint | First token | Output speed | Total |
|---|---|---|---|
| deepseek | 618 ms | n/a | 4.5 s |
| sference | 1,432 ms | n/a | 5.4 s |
| fireworks | 1,829 ms | n/a | 4.2 s |
| tensorx | 1,891 ms | n/a | 2.1 s |
| runware | 2,184 ms | n/a | 2.3 s |
| novita | 2,889 ms | n/a | 9.3 s |
| scaleway | 3,697 ms | n/a | 14.0 s |
| deepinfra | 12,617 ms | n/a | 14.3 s |
| sail | 18,416 ms | n/a | 30.5 s |
Posts on X
Here are some posts on X comparing both models.
Independent reviews
- A category-wide review of flash-tier models, covering DeepSeek V4, GLM-5.3 and Qwen3.8 Flash-Next. It names no overall winner.Local AI Zone
How each model feels to use
GLM-5.3 Flash
- Default OpenAI-compatible requests to GLM-5.3 Flash returned HTTP 200 but spent most completion tokens on reasoning with little visible output, becoming usable only after switching to GLM-style thinking settings and a larger output ceiling.glbgpt.com
- Gateway support for GLM-5.3 Flash varies, with some providers marking structured output unsupported and handling system prompts differently.datacamp.com
- Once configured for chat, GLM-5.3 Flash followed strict schemas, wrote executable code, ignored simple prompt injections and retrieved the right records from noisy input on the first attempt.kingy.ai
DeepSeek V4 Flash 0731
- DeepSeek V4 Flash 0731 does not accept image input, and users cite GLM's vision support as a practical differentiator, with DeepSeek needing a separate vision-capable model as a sidecar.Artificial Analysis@gilfredstuart@imost
Across both models
- Both models carry the maximum 4/4 verbosity rating, with GLM generating 180M and DeepSeek 240M output tokens on the Intelligence Index against a 120M median, so reviewers recommend constraining response formats to control cost and context use in agent traces.Artificial AnalysisArtificial Analysisbuildfastwithai.com
Recent changes
- Aug 16, 2026: DeepSeek moved API pricing to peak and off-peak rates, with output at $1.32 per million peak and $0.66 off-peak, up from a flat $0.28.deepseek_aiapi-docs.deepseek.com
- Under the same change, cache-miss input rose to $0.44 per million at peak and $0.22 off-peak, with cache-hit input at $0.014 peak and $0.007 off-peak; peak hours are 01:00-04:00 and 06:00-10:00 UTC.deepseek_aiapi-docs.deepseek.com
- Context window, max output and concurrency limits were unchanged by the pricing update.api-docs.deepseek.com
Which should you choose?
GLM-5.3 Flash is the stronger default for agent pipelines and anything touching images: it leads on the Intelligence and Agentic indexes, takes image input on most of our endpoints, and blended to $0.06 per million tokens on our routed traffic. Budget setup time for its thinking settings, since default OpenAI-compatible requests can burn tokens on hidden reasoning. DeepSeek V4 Flash 0731 suits interactive text-only chat, where its 541 ms median first token beats GLM's 1,684 ms and its $0.007 cached reads reward heavy prompt reuse. Whichever you choose, constrain response formats: both models are rated maximally verbose.
Frequently asked questions
- Which model should I use for an agent that processes screenshots or documents with images?
- GLM-5.3 Flash. It accepts image input on most of our endpoints, while DeepSeek V4 Flash 0731 is text-only and would need a separate vision-capable model alongside it.
- Why does GLM-5.3 Flash return empty-looking responses on default settings?
- Reviewers found default OpenAI-compatible requests spent most completion tokens on hidden reasoning. Enable GLM-style thinking settings and raise the output ceiling to get usable answers.
- How do the costs compare in practice?
- Very close after caching: $0.06 per million tokens for GLM-5.3 Flash versus $0.07 for DeepSeek V4 Flash 0731 on our routed traffic. DeepSeek's $0.007 cached reads favor heavy prompt reuse, but its own API charges double at peak hours.
- AUG '26
Claude Opus 5 vs GPT-5.6 Sol: API speed, coding and cost
Sol starts responding sooner. Opus streams faster and recorded fewer server errors. Compare coding results, context limits and costs after caching.
- SEP '26
GPT-6 Astra is the best model you may not be allowed to call: gated access, the EU gap and safety stops
Astra launched 3 September at $10 and $50 per million tokens with the first Critical cybersecurity rating in OpenAI's framework. Three days later, most developers cannot call it, Foundry customers in the EU cannot keep it in region, and the ones who can call it are learning that a safety monitor can end a job mid-run. The benchmarks are one story. Getting it into production is another.
- SEP '26
The list price is now a range: peak hours, promo windows and host floors
In August the frontier rate card stopped being a number. DeepSeek bills V4 Pro at $1.32 input during seven weekday hours and $0.66 the rest of the time. OpenAI cut GPT-5.6 Sol to $4 and $20 but only through 21 November. Gemini 3.8 Flash doubles on 1 January. And the cheapest GPT-4-class price in the market is set by a reseller, not by the lab. Your cost model needs a clock, a calendar and a host column.


