View model prices
1
Open the Model marketplace
Click Model marketplace in the top navigation.
2
Search for a model
Use the search box and enter a model keyword such as
gpt-5.6-sol, claude-sonnet-5, or gemini-3.7-flash to jump straight to it.3
Read the multipliers and unit price
Each model shows:
- Model ID: used in the
modelfield for API calls - Prompt token multiplier: pricing coefficient for input
- Completion token multiplier: pricing coefficient for output
- Cache token multiplier: pricing coefficient when the cache hits
- Available groups: which groups can call this model
- Performance: model success rate over the last 24 hours
The list of supported models and their prices changes as the market shifts. Trust the live pricing page as the source of truth.
Billing formula
What is quota, and how is it calculated?Billing formula
- Prompt tokens: content you send to the model, including the system prompt, history, and the current message.
- Completion tokens: content the model returns.
- Group multiplier: price differences across models are expressed through multipliers. Higher multipliers mean higher cost per token.
Quota conversion
Quota is the platform’s internal billing unit, like “balance points” in your account. On the top-up page, the amount selector shows how much quota one CNY or USD is worth. Refer to the top-up page for the live rate.Image, audio, and other endpoints
Beyond chat completions, endpoints such as image generation (/v1/images/generations), speech-to-text (/v1/audio/transcriptions), and TTS (/v1/audio/speech) are typically billed per call or per second / character. The pricing page lists the billing unit for each of these models separately.
When a model is unavailable
- If the pricing page has no result for a model, the platform is not currently offering it (or it has been retired).
- Some models are only available to specific user groups and may not be visible to standard users. Contact support to request a group change.
- When an upstream provider is under maintenance, calls may fail. See Call the API for the meaning of error codes.
Next steps
Quota and top-up
Top up quota to call paid models.
Usage logs
Break down actual consumption by model and token.