deepseek-v4-flash-0731
DeepSeek
- Input
- $0.14/M$0.036/M
- Output
- $0.28/M$0.071/M
One OpenAI-compatible API for leading AI models. We find the lowest healthy marketplace price, never charge above a published list price, and optimize each request before it reaches the model.
Search the models available through one API. Every comparison comes from the same catalog used to price requests.
Ranked by aggregate token usage across the current filter.
61 models
Live marketplace rates
| Model | Modality | Input | Output | Discount | Details |
|---|---|---|---|---|---|
| deepseek-v4-flash-0731DeepSeek | TextReasoning | $0.14/M$0.036/M | $0.28/M$0.071/M | 74.54% off | View model |
| gpt-5.6-lunaOpenAI | MultimodalReasoning | $0.20/M$0.08/M | $1.20/M$0.48/M | 60% off | View model |
| gpt-5.6-terraOpenAI | MultimodalReasoning | $2.00/M$0.80/M | $12.00/M$4.80/M | 60% off | View model |
| gpt-oss-120bOpenAI | TextReasoning | $0.088/M$0.035/M | $0.50/M$0.20/M | 60% off | View model |
| gpt-5.6-solOpenAI | MultimodalReasoning | $2.00/M$1.00/M | $10.00/M$5.00/M | 50% off | View model |
| glm-4.5Z.ai | TextReasoning | $0.60/M$0.33/M | $2.20/M$1.21/M | 45% off | View model |
| glm-4.5-airZ.ai | TextReasoning | $0.20/M$0.11/M | $1.10/M$0.605/M | 45% off | View model |
| glm-4.6Z.ai | Text | $0.43/M$0.237/M | $1.75/M$0.963/M | 45% off | View model |
DeepSeek
OpenAI
OpenAI
OpenAI
OpenAI
Z.ai
Z.ai
Z.ai
Prices follow available marketplace capacity and never exceed a published provider list price. A discount appears only when a verified comparison price exists.
System architecture / 01
First a marketplace price, never above list. Then we remove tokens before they reach the model.
Monthly savings
$0
Yearly savings
$0
Effective bill
$0
Based on the 34.4% average discount across the live catalog; your mix of models decides your own, and real optimization varies by workload.
The application record, optional caches, and the provider serving a request have different data boundaries. Here is where each one starts and ends.
Core metering records retain model, endpoint, token counts, charges, status, and sanitized failure details. They do not include prompt or response bodies; optional caching is a separate data boundary described below.
Field notes / FAQ
Keep your SDK and prompts. Point the base URL at us.
base_url="https://api.cheaperinference.com/v1"Sellers compete for your traffic in an open order book. You get whoever is cheapest right now.
We trim, cache, and route, then attribute every saving.
Account response caching can store request and response pairs within that account and can be switched off. Provider prompt caching may also create provider-side cached state.
Inputs are sent to the provider selected to serve each request. That provider may process or retain content under the contractual settings and retention practices applicable to its service.
We do not make a universal zero-data-retention claim. Confirm exact route and provider coverage with security before sending production data.