On TokenHub, different models may have different billing methods. The following introduces the billing methods for each type of model. You can also go to the model details page from Model Gallery to view the billing rules for each model. Language Models
|
Input tokens | Pay-as-you-go | USD / million tokens | Token consumption of user input text (including system prompt) |
Output tokens | Pay-as-you-go | USD / million tokens | Token consumption of model-generated text |
Cached Input | Pay-as-you-go | USD / million tokens | Cached input Token consumption |
Billing Notes:
Some models support tiered pricing (such as different unit prices for different input length ranges).
Some reasoning models' thinking process (Reasoning) and regular output are priced separately.
Prerequisites
Each time you claim a free resource package for a model or enable postpaid service for a model, USD 1 will be automatically frozen in your account. If your account balance or credit limit is insufficient, the claim or enablement will fail.
The frozen amount will be returned to your account after the resources are terminated. If you want to get your pre-frozen amount back, delete all model endpoints on the Online Inference page. To delete the default endpoint, contact customer service. References
For detailed pricing information of each model, see Model Pricing. For details on the impact of billing arrears and recovery mechanisms, see Overdue Payments.