Requesty
Back|SEP '26AI MODELS / BENCHMARKS
6 MIN READ|

GLM-5.3 Flash vs DeepSeek V4 Flash 0731: which flash model fits your API workload?

Last updated

GLM-5.3 Flash reads images and scores higher on reasoning and agent benchmarks. DeepSeek V4 Flash 0731 is text-only but starts answering about three times sooner on traffic routed through us. Both are open-weight, million-token flash models in our catalog, released within a month of each other, and both cost pennies per session after caching.

The short answer

Verdict

Pick GLM-5.3 Flash for agent pipelines that need vision or stronger multi-step reasoning: it leads Artificial Analysis's Intelligence and Agentic indexes and blended to $0.06 per million tokens on our traffic. Pick DeepSeek V4 Flash 0731 for fast-start, text-only chat, where its 541 ms median first token wins. Plain code generation is close either way.

At a glance

Both are open-weight, million-token flash models in our catalog, released within a month of each other. The practical split: GLM-5.3 Flash is natively multimodal, so it accepts images on most of our endpoints, while DeepSeek V4 Flash 0731 is text-only but ships as a 304B MoE with speculative decoding aimed at fast, cheap inference and agentic tool use. Both support reasoning, tool calling and prompt caching across our endpoints.

Intelligence
GLM-5.3 Flash41.9
DeepSeek V4 Flash 073134.5
Artificial Analysis Intelligence Index
Coding
GLM-5.3 Flash71.5
DeepSeek V4 Flash 073169.1
Artificial Analysis Coding Index
Context
GLM-5.3 Flash1,048.576K
DeepSeek V4 Flash 07311,048.576K
Widest context window any provider serves
Time to first token
GLM-5.3 Flash1,684 ms
DeepSeek V4 Flash 0731541 ms
Median across all Requesty traffic
Output speed
GLM-5.3 Flash3,849 tok/s
DeepSeek V4 Flash 0731122 tok/s
Median across all Requesty traffic
Blended price
GLM-5.3 Flash$0.06/M
DeepSeek V4 Flash 0731$0.07/M
Actual cost per million tokens across our routed customer traffic, after caching

At a glance

GLM-5.3 FlashDeepSeek V4 Flash 0731
Trained byZ.AIDeepSeek
Released2026-08-262026-07-31
Context window1.049M tokens (1M on some endpoints)1.049M tokens (256K on some endpoints)
Max output1.049M tokens1.049M tokens
Providers on Requestyzai, fireworks, novita, deepinfra, tensorx, runware, sferencefireworks, novita, deepinfra, tensorx, runware, scaleway, sail, sference
Image inputYesNo
Reasoning controlsYesYes
Prompt cachingYesYes
Tool callingYesYes
Computer useNoNo
Web searchNoNo
Structured outputYesYes
Open weightsYesYes
Data retention (first-party endpoint)NoneNone
Trains on your dataNoNo
Share of Requesty tokens2.8%9.1%
Specifications and endpoint features from our catalog. Token share from our routed traffic.

Pricing and caching

Across traffic routed through us from July 6 to September 7, GLM-5.3 Flash blended to $0.06 per million tokens and DeepSeek V4 Flash 0731 to $0.07, with caching cutting bills by roughly 60% and 57% respectively (our model economics). List prices sit at $0.15 input and $0.50 output for GLM on zai/glm-5.3-flash versus $0.22 and $0.66 for DeepSeek on fireworks/deepseek-v4-flash-0731, though DeepSeek's $0.007 cached reads are far cheaper than GLM's $0.03. On DeepSeek's own API, peak-hour rates are double off-peak, with peak at 01:00-04:00 and 06:00-10:00 UTC since August 16, 2026 (DeepSeek changelog).

Pricing and caching

Our routed customer traffic · 2026-07-06 to 2026-09-07

Requesty
MetricGLM-5.3 FlashDeepSeek V4 Flash 0731
What customers paid
Blended price includes cache savings.
Blended price / 1M tokens$0.06$0.07
Cache hit rate91%82%
Bill saved by caching60.5%57.3%
Price breakdown and benchmark costs
MetricGLM-5.3 FlashDeepSeek V4 Flash 0731
Customer traffic costs
Cost per request$0.004718$0.002380
Median session cost$0.04$0.03
List prices
Our catalog · 2026-09-10. Canonical endpoints: zai/glm-5.3-flash / fireworks/deepseek-v4-flash-0731.
Input / 1M tokens$0.15$0.22
Cached input / 1M tokens$0.03$0.007
Output / 1M tokens$0.5$0.66
Artificial Analysis cost per task
Weighted evaluation cost · Intelligence Index v4.3, 2026-09-10. GLM-5.3-Flash / DeepSeek V4 Flash 0731 (Reasoning, Max Effort).
Cost per benchmark task$0.2533$0.2196
Endpoint list prices
All priced endpoints
USD per million tokens. Cached-input rates apply to eligible cache reads.
ModelEndpointInputOutputCached inputContext
GLM-5.3 Flashrunware$0.075$0.250$0.0151.0M
GLM-5.3 Flashzai$0.150$0.500$0.0301M
GLM-5.3 Flashfireworks$0.150$0.500$0.0301.0M
GLM-5.3 Flashnovita$0.150$0.500$0.0301.0M
GLM-5.3 Flashdeepinfra$0.150$0.500$0.0301.0M
GLM-5.3 Flashtensorx$0.200$0.500$0.0501.0M
GLM-5.3 Flashsference$0.200$0.600$0.0701M
DeepSeek V4 Flash 0731deepinfra$0.072$0.144$0.0141.0M
DeepSeek V4 Flash 0731runware$0.076$0.153$0.0141.0M
DeepSeek V4 Flash 0731deepinfra$0.090$0.180$0.0181.0M
DeepSeek V4 Flash 0731sail$0.090$0.180$0.0201.0M
DeepSeek V4 Flash 0731novita$0.140$0.280$0.0281.0M
DeepSeek V4 Flash 0731tensorx$0.250$0.300$0.0601.0M
DeepSeek V4 Flash 0731sference$0.280$0.560$0.0701.0M
DeepSeek V4 Flash 0731fireworks$0.220$0.660$0.00701M
DeepSeek V4 Flash 0731scaleway$0.460$0.930$0.090256K

Benchmarks and use cases

On Artificial Analysis's version 4.3 indexes, GLM-5.3 Flash leads DeepSeek V4 Flash 0731 on the Intelligence Index (41.9 vs 34.5) and the Agentic Index (51.2 vs 41.7), with the DeepSeek figures measured in reasoning mode at max effort (GLM-5.3 Flash, DeepSeek V4 Flash). Coding is closer, 71.5 against 69.1, so for straightforward code generation either model is a reasonable pick. The wider gaps sit in general reasoning and multi-step agent work.

Benchmark scores

Artificial Analysis · source score units

Requesty
GLM-5.3 FlashDeepSeek V4 Flash 0731
Intelligence IndexVersioned Artificial Analysis composite, not a percentage
41.9
34.5
Coding IndexArtificial Analysis coding composite for this snapshot
71.5
69.1
Agentic IndexArtificial Analysis agentic composite for this snapshot
51.2
41.7
Snapshot 2026-09-10. GLM-5.3-Flash / DeepSeek V4 Flash 0731 (Reasoning, Max Effort). Index versions: 4.3 / 4.3.

Latency and throughput

Across traffic routed through us from July 6 to September 7, DeepSeek V4 Flash 0731 started responding sooner, with a median time to first token of 541 ms against 1,684 ms for GLM-5.3 Flash (model economics). Once streaming, GLM-5.3 Flash delivered tokens far faster at the median, and its 5xx error rate was lower, 0.4% versus 1.9%. Endpoint choice matters for both: our latency rankings show first-token times ranging from roughly 600 ms to well over 10 s depending on provider.

Latency and throughput

Our routed customer traffic · 2026-07-06 to 2026-09-07

Requesty
MetricGLM-5.3 FlashDeepSeek V4 Flash 0731
Speed and errors
Median time to first token1,684 ms541 ms
p95 time to first token7,475 ms3,052 ms
Median output speed3,849 tok/s122 tok/s
Upstream 5xx error rate0.4%1.9%
Endpoint results
GLM-5.3 Flash
EndpointFirst tokenOutput speedTotal
sference1,219 msn/a4.1 s
runware2,909 msn/a4.7 s
fireworks3,171 msn/a7.9 s
tensorx5,290 msn/a10.1 s
zai7,383 msn/a9.8 s
deepinfra8,607 msn/a10.5 s
novita101,992 msn/a102.1 s
DeepSeek V4 Flash 0731
EndpointFirst tokenOutput speedTotal
deepseek618 msn/a4.5 s
sference1,432 msn/a5.4 s
fireworks1,829 msn/a4.2 s
tensorx1,891 msn/a2.1 s
runware2,184 msn/a2.3 s
novita2,889 msn/a9.3 s
scaleway3,697 msn/a14.0 s
deepinfra12,617 msn/a14.3 s
sail18,416 msn/a30.5 s

Posts on X

Here are some posts on X comparing both models.

AIHubmix
AIHubmix
@AiHubMix · 2026-08-28
𝕏

GLM 5.3 Flash vs GLM 5.3 vs DeepSeek V4 Flash 0731 Same Falcon 9 prompt🚀 One shot. One gateway 🔹GLM 5.3-$0.459 | 2,078s 🔹GLM 5.3 Flash-$0.034 | 1,912s 🔹V4 Flash-$0.002 | 14s V4 was explicitly told to reason carefully. It still delivered in 14s. Benchmark via AIHubMix https://t.co/12yE3L8XWU

61 likes · 4 reposts · 9 repliesView post ↗
Jackson Atkins
Jackson Atkins
@JacksonAtkinsX · 2026-08-26
𝕏

GLM-5.3-Flash vs. DeepSeek-v4-Flash Same prompt, same harness (pi), same effort. GLM used 114k tokens over 93 turns taking 73min. Total cost 16 cents. DeepSeek used 70k tokens over 9 turns taking 10min. Total cost 2 cents. GLM is a richer scene but at 8x the cost and 7x the time!

157 likes · 13 reposts · 16 repliesView post ↗
Fede(URU) 🇺🇾
Fede(URU) 🇺🇾
@RealFedeURU · 2026-08-28
𝕏

Glm 5.3 vs Qwen 3.8 vs Gemini 3.7 vs Deepseek v4 flash https://t.co/OK622iHa3g

6 likes · 0 reposts · 0 repliesView post ↗

Independent reviews

  • A category-wide review of flash-tier models, covering DeepSeek V4, GLM-5.3 and Qwen3.8 Flash-Next. It names no overall winner.Local AI Zone

How each model feels to use

GLM-5.3 Flash

  • Default OpenAI-compatible requests to GLM-5.3 Flash returned HTTP 200 but spent most completion tokens on reasoning with little visible output, becoming usable only after switching to GLM-style thinking settings and a larger output ceiling.glbgpt.com
  • Gateway support for GLM-5.3 Flash varies, with some providers marking structured output unsupported and handling system prompts differently.datacamp.com
  • Once configured for chat, GLM-5.3 Flash followed strict schemas, wrote executable code, ignored simple prompt injections and retrieved the right records from noisy input on the first attempt.kingy.ai

DeepSeek V4 Flash 0731

Across both models

Recent changes

  • Aug 16, 2026: DeepSeek moved API pricing to peak and off-peak rates, with output at $1.32 per million peak and $0.66 off-peak, up from a flat $0.28.deepseek_aiapi-docs.deepseek.com
  • Under the same change, cache-miss input rose to $0.44 per million at peak and $0.22 off-peak, with cache-hit input at $0.014 peak and $0.007 off-peak; peak hours are 01:00-04:00 and 06:00-10:00 UTC.deepseek_aiapi-docs.deepseek.com
  • Context window, max output and concurrency limits were unchanged by the pricing update.api-docs.deepseek.com

Which should you choose?

GLM-5.3 Flash is the stronger default for agent pipelines and anything touching images: it leads on the Intelligence and Agentic indexes, takes image input on most of our endpoints, and blended to $0.06 per million tokens on our routed traffic. Budget setup time for its thinking settings, since default OpenAI-compatible requests can burn tokens on hidden reasoning. DeepSeek V4 Flash 0731 suits interactive text-only chat, where its 541 ms median first token beats GLM's 1,684 ms and its $0.007 cached reads reward heavy prompt reuse. Whichever you choose, constrain response formats: both models are rated maximally verbose.

Frequently asked questions
Which model should I use for an agent that processes screenshots or documents with images?
GLM-5.3 Flash. It accepts image input on most of our endpoints, while DeepSeek V4 Flash 0731 is text-only and would need a separate vision-capable model alongside it.
Why does GLM-5.3 Flash return empty-looking responses on default settings?
Reviewers found default OpenAI-compatible requests spent most completion tokens on hidden reasoning. Enable GLM-style thinking settings and raise the output ceiling to get usable answers.
How do the costs compare in practice?
Very close after caching: $0.06 per million tokens for GLM-5.3 Flash versus $0.07 for DeepSeek V4 Flash 0731 on our routed traffic. DeepSeek's $0.007 cached reads favor heavy prompt reuse, but its own API charges double at peak hours.
Related reading

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders