All models
GLM 5.3 Flash
Best value1,000,000 tokens of context
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead
Provider
WellFlow Premium
Provider model
glm-5.3-flash
Pricing
per 1M tokens
Input
Wellflow price: $0.04provider price: $0.15
Output
provider price: $0.50Wellflow price: $0.14
Savings
72%
Prompt caching
The repeated part of a request (system prompt, documents, chat history) is stored in the cache. Reading it back costs less than regular input.
- Cache read
- Wellflow price: $0.01
- -75% vs input price
- Cache write
- Wellflow price: $0.04
Performance
Availability and speed of the model through Wellflow.
Uptime
OperationalPartial outageOutage