Skip to content

Currency and Billing Information

Currency and Billing Information

This site uniformly uses US dollars for billing, with a USD to CNY exchange rate of 1:7.3.

Pricing

For specific pricing, please check the pricing page. Discounts vary for different channel sources, please check user benefits for details. The final price is: Model Price × Channel Discount × User Discount.

User Benefits

User benefits refer to the privileges users enjoy when using the service, including request rates, discounts, etc. Currently, user benefits are automatically upgraded based on cumulative recharge amount. For details, please check the benefits page. If you have higher requirements for rates or need enterprise-level services, please contact our customer service.

Model Billing

Model billing is divided into the following types:

Per-Request Billing

Per-request billing means charging once for each request.

  • Billing Formula: Model Input Price × Channel Discount × User Discount

Token-Based Billing

Token-based billing charges based on the number of tokens used in the request.

  • Billing Formula: Cost = (Input tokens / 1000 × Input Price) + (Output tokens / 1000 × Output Price) Final Charge = Cost × Channel Discount × User Discount

Per-Second Billing

Per-second billing generally appears in video generation models.

  • Billing Formula: Cost = Input Price × Seconds Final Charge = Cost × Channel Discount × User Discount

Per-Megapixel Billing

Per-megapixel billing generally appears in image generation models.

  • Billing Formula: Cost = (Width × Height / 1000000) × Input Price Final Charge = Cost × Channel Discount × User Discount

Per-Image Billing

Per-image billing generally appears in image generation models, charging based on the number of generated images.

  • Billing Formula: Cost = Number of Images × Input Price Final Charge = Cost × Channel Discount × User Discount

Other Billing (Cache/Audio/Inference Additional Charges)

If the model supports cache/audio/inference, etc., these requests may also be charged. The system's billing method is: charging based on multiples of input/output prices.

  • Core Calculation Formula: Input/Output tokens = Input/Output tokens + ((Cache/Audio/Inference tokens) × (Multiple - 1))
  • Subsequent Billing: Based on the adjusted actual input/output tokens, calculate according to the corresponding model's billing method (such as token-based billing).

Billing Items and Corresponding Relationships

Billing ItemCorresponding RelationshipDescription
Cache TokensInput tokens × RateThis value appears when automatic caching is triggered
Input Audio TokensInput tokens × RateIf applicable, this value appears when inputting audio
Output Audio TokensOutput tokens × RateIf applicable, this value appears when outputting audio
Input Image TokensInput tokens × RateIf applicable, this value appears when inputting images
Output Image TokensOutput tokens × RateIf applicable, this value appears when outputting images
Cache Write 5M TokensInput tokens × RateGenerally exists in Claude models
Cache Write 1H TokensInput tokens × RateGenerally exists in Claude models
Cache Read TokensInput tokens × RateIf applicable, this value appears when cache is triggered
Inference TokensOutput tokens × RateBilling for some Gemini models

Field Multipliers

Some models adjust prices based on request field values, generally existing in video models.

For example: The Veo model can use priority pay go. By adding a header when making requests, you can get priority queuing, with each request price ×1.8.

Model Pricing Reference

Note

All prices below are in USD, per 1 million tokens ($/1M tokens) unless otherwise noted. Prices are subject to change — please refer to the latest platform announcements.

OpenAI (GPT)

GPT-5.6 Price Adjustment (Effective 2026-08-01)

  • gpt-5.6-luna: 80% price reduction
  • gpt-5.6-terra: 20% price reduction
  • Dedicated channels: effective immediately | Standard channels: effective August 1, 2026

Language Models

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)Cache Creation ($/1M tokens)
gpt-6-astra NewOfficial≤272K: $10 / ∞: $20 / Priority ≤272K: $20 / Priority ∞: $40≤272K: $50 / ∞: $75 / Priority ≤272K: $100 / Priority ∞: $150≤272K: $1 / ∞: $2 / Priority ≤272K: $2 / Priority ∞: $4≤272K: $12.5 / ∞: $25 / Priority ≤272K: $25 / Priority ∞: $50
gpt-5.5Official$5$30$0.5-
gpt-5.6-luna NewOfficial≤272K: $0.2 / ∞: $0.4 / Priority: $0.4≤272K: $1.2 / ∞: $1.8 / Priority: $2.4≤272K: $0.02 / ∞: $0.04 / Priority: $0.04≤272K: $0.25 / ∞: $0.5 / Priority: $0.5
gpt-5.6-terra NewOfficial≤272K: $2 / ∞: $4 / Priority: $4≤272K: $12 / ∞: $18 / Priority: $24≤272K: $0.2 / ∞: $0.4 / Priority: $0.4≤272K: $2.5 / ∞: $5 / Priority: $5
gpt-5.6-sol NewOfficial≤272K: $5 / ∞: $10 / Priority: $10≤272K: $30 / ∞: $45 / Priority: $60≤272K: $0.5 / ∞: $1 / Priority: $1≤272K: $6.25 / ∞: $12.5 / Priority: $12.5
gpt-5.4-proOfficial≤272K: $30 / >272K: $60≤272K: $180 / >272K: $270-
gpt-5.4Official≤272K: $2.5 / >272K: $5≤272K: $15 / >272K: $22.5≤272K: $0.25 / >272K: $0.5
gpt-5.4-miniOfficial / Cloud$0.75$4.5$0.075
gpt-5.4-nanoOfficial / Cloud$0.2$1.25$0.02
gpt-5.3-codexOfficial$1.75$14$0.175
gpt-5.3-codexCloud$1.75$14$0.18
gpt-5.2-proOfficial$21$168-
gpt-5.2-codexOfficial$1.75$14$0.175
gpt-5.2-codexCloud$1.75$14$0.18
gpt-5.2Official$1.75$14$0.175
gpt-5.2Cloud$1.75$14$0.18
gpt-5.1-chat-latestOfficial$1.25$10$0.125
gpt-5.1-codex-maxOfficial$1.25$10$0.125
gpt-5.1-codexOfficial$1.25$10$0.125
gpt-5.1Official$1.25$10$0.125
gpt-5.1Cloud$1.25$10$0.13
gpt-5-proOfficial$15$120-
gpt-5-chat-latestOfficial$1.25$10$0.125
gpt-5-codexOfficial$1.25$10$0.125
gpt-5Official$1.25$10$0.125
gpt-5Cloud$1.25$10$0.13
gpt-5-miniOfficial / Cloud$0.25$2$0.025
gpt-5-nanoOfficial$0.05$0.4$0.005
gpt-5-nanoCloud$0.05$0.4$0.01
gpt-4.1Official$2$8$0.5
gpt-4.1Cloud$0.4$1.6-
gpt-4.1-miniOfficial / Cloud$0.4$1.6$0.1
gpt-4.1-nanoOfficial$0.1$0.4$0.025
gpt-4oOfficial / Cloud$2.5$10$1.25
gpt-4o-2024-05-13Official$5$15-
gpt-4o-miniOfficial / Cloud$0.15$0.6$0.075
gpt-4-turboOfficial$10$30-
gpt-4Official$30$60-
gpt-3.5-turboOfficial$0.5$1.5-
gpt-3.5Official$1.5$2-
gpt-image-1.5 DeprecatedOfficial / Cloud$8$32$2
gpt-image-2Official / CloudImage: $8 / Text: $5Image: $30Image: $2 / Text: $1.25
o1-proOfficial$150$600-
o1Official$15$60$7.5
o1-miniOfficial$1.1$4.4$0.55
o3-proOfficial$20$80-
o3Official$2$8$0.5
o3Cloud$2.2$8.8$0.55
o3-deep-researchOfficial$10$40$2.5
o3-miniOfficial$1.1$4.4$0.55
o4-miniOfficial$1.1$4.4$0.275
o4-mini-deep-researchOfficial$2$8$0.5

Voice and Audio

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)Reference Price
gpt-realtime-2.1 NewOfficialText: $4Text: $24Text: $0.4Multimodal real-time model
gpt-realtime-2.1-mini NewOfficialText: $0.6Text: $2.4Text: $0.3Multimodal real-time model
gpt-realtime-2OfficialAudio: $32 / Text: $4 / Image: $5Audio: $64 / Text: $24Audio: $0.4 / Text: $0.4 / Image: $0.5Multimodal real-time model
gpt-realtime-miniOfficialText: $0.6Text: $2.4Text: $0.3Multimodal real-time model
whisper-1Official---$0.006 / minute
gpt-4o-transcribeOfficial$2.50$10-~$0.006 / minute
gpt-4o-transcribe-diarizeOfficial$2.50$10-~$0.006 / minute
gpt-4o-mini-transcribeOfficial$1.25$5-~$0.003 / minute

Video Generation (Sora)

ModelChannelPrice ($/sec)
sora-2 DeprecatedOfficial / Cloud$0.1
sora-2-pro DeprecatedOfficial / Cloud720p: $0.3 / 1024p: $0.5

Claude (Anthropic)

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)Cache Write 5m ($/1M tokens)Cache Write 1h ($/1M tokens)
claude-fable-5-1 NewOfficial / Cloud$10$50$0.25$12.5
claude-opus-5 NewOfficial / Cloud$5$25$0.5$6.25$10
claude-opus-4-6Official / Cloud≤200K: $5 / >200K: $10≤200K: $25 / >200K: $37.5≤200K: $0.5 / >200K: $1≤200K: $6.25 / >200K: $12.5
claude-opus-4-7Official / Cloud≤200K: $5 / >200K: $10≤200K: $25 / >200K: $37.5≤200K: $0.5 / >200K: $1≤200K: $6.25 / >200K: $12.5
claude-opus-4-8Official / Cloud≤200K: $5 / >200K: $10≤200K: $25 / >200K: $37.5≤200K: $0.5 / >200K: $1≤200K: $6.25 / >200K: $12.5
claude-sonnet-4-6Official / Cloud≤200K: $3 / >200K: $6≤200K: $15 / >200K: $22.5≤200K: $0.3 / >200K: $0.6≤200K: $3.75 / >200K: $7.5
claude-sonnet-4-5-20250929Official / Cloud≤200K: $3 / >200K: $6≤200K: $15 / >200K: $22.5≤200K: $0.3 / >200K: $0.6≤200K: $3.75 / >200K: $7.5
claude-sonnet-5Official / Cloud≤200K: $2 / >200K: $4≤200K: $10 / >200K: $15≤200K: $0.2 / >200K: $0.4≤200K: $2.5 / >200K: $5
claude-opus-4-5-20251101Official / Cloud$5$25$0.5$6.25
claude-opus-4-1-20250805 RetiredOfficial$15$75$1.5$18.5
claude-opus-4-20250514Official$15$75$1.5$18.5
claude-sonnet-4-20250514Official / Cloud$3$15$0.3$3.75
claude-fable-5Official / Cloud$10$50$0.25$12.5
claude-mythos-5Official / Cloud$10$50$1$12.5
claude-haiku-4-5-20251001Official$1$5$0.1$1.25
claude-3-7-sonnet-20250219 RetiredOfficial / Cloud$3$15$0.3$3.75
claude-3-5-sonnet-20241022 RetiredOfficial / Cloud$3$15--
claude-3-5-sonnet-20240620 RetiredOfficial / Cloud$3$15--
claude-3-5-haiku-20241022 RetiredOfficial$0.8$4$0.08$1
claude-3-haiku-20240307 RetiredOfficial / Cloud$0.25$1.25$0.03$0.3
claude-3-sonnet-20240229 RetiredOfficial / Cloud$3$15--
claude-3-opus-20240229 RetiredOfficial$15$75$1.5$18.5

Gemini (Google)

Language Models

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)
gemini-3.8-flash NewOfficial / Cloud$0.75$3.75$0.075
gemini-3.7-flash NewOfficial / Cloud$0.75$3.75$0.075
gemini-3.6-flash NewOfficial / Cloud$1.5$7.5$0.15
gemini-3.5-flash-lite NewOfficial / Cloud$1.5$9$0.15
gemini-3.5-flashOfficial / Cloud$1.5$9$0.15
gemini-3.1-pro-previewOfficial / Cloud≤200K: $2 / >200K: $4≤200K: $12 / >200K: $18≤200K: $0.2 / >200K: $0.4
gemini-3-pro-imageOfficial / Cloud$2Text $12 / Image $120-
gemini-3-pro-image-preview DeprecatedOfficial / Cloud$2Text $12 / Image $120-
gemini-3.1-flash-imageOfficial$0.5$3-
gemini-3.1-flash-image-preview DeprecatedOfficial$0.5Text $3 / Image $60-
gemini-3-flash-previewOfficial / Cloud$0.5$3$0.05
gemini-3.1-flash-liteOfficial$0.25$1.5$0.025
gemini-3.1-flash-lite-preview DeprecatedOfficial$0.25$1.5$0.025
gemini-3-pro-preview RetiredOfficial / Cloud$2$12$0.2
gemini-2.5-pro DeprecatedOfficial / Cloud≤200K: $1.25 / >200K: $2.5≤200K: $10 / >200K: $15≤200K: $0.125 / >200K: $0.25
gemini-2.5-flash DeprecatedOfficial / Cloud$0.3$2.5$0.03
gemini-2.5-flash-lite DeprecatedOfficial / Cloud$0.1$0.4$0.01
gemini-2.5-flash-image-previewOfficial$0.3$30-
gemini-2.5-flash-image DeprecatedOfficial$0.3$30-
gemini-2.0-flash RetiredOfficial / Cloud$0.15$0.6-
gemini-2.0-flash-001 RetiredOfficial$0.1$0.4$0.025
gemini-2.0-flash-lite-001 RetiredOfficial / Cloud$0.075$0.3-

Video Generation (Veo)

ModelChannelPrice ($/sec)
veo-3.1-generate-previewOfficial / Cloud$0.4
veo-3.1-fast-generate-previewOfficial / Cloud$0.15
veo-3.0-generate-001 RetiredOfficial / Cloud$0.4
veo-3.0-fast-generate-001 RetiredOfficial / Cloud$0.15

Grok (xAI)

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)
grok-4.6 NewOfficial≤200K: $2 / ∞: $4≤200K: $6 / ∞: $12≤200K: $0.5 / ∞: $1
grok-4.5 NewOfficial$2.00$6.00$0.50
grok-4-1-fast-reasoningOfficial$0.2$0.5$0.05
grok-4-1-fast-non-reasoningOfficial$0.2$0.5$0.05

Kimi (Moonshot)

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)
kimi-k3 NewOfficial$3$15$0.30
kimi-k2.7-codeOfficial$0.95$4$0.19
kimi-k2.7-code-highspeedOfficial$1.9$8$0.38
kimi-k2.6Official$0.6$3$0.1
kimi-k2.5Official$0.6$3$0.1
kimi-latestOfficial≤8K: $0.2 / ≤32K: $1 / >32K: $2≤8K: $2 / ≤32K: $3 / >32K: $5$0.15
kimi-k2-turbo-preview RetiredOfficial$1.15$8$0.15
kimi-k2-thinking-turbo RetiredOfficial$1.15$8$0.15
kimi-k2-thinking RetiredOfficial$0.6$2.5$0.15
kimi-k2-0905-preview RetiredOfficial$0.6$2.5$0.15
kimi-k2-0711-preview RetiredOfficial$0.6$2.5$0.15
moonshot-v1-128k-vision-preview DeprecatedOfficial$2$5-
moonshot-v1-32k-vision-preview DeprecatedOfficial$1$3-
moonshot-v1-8k-vision-preview DeprecatedOfficial$0.2$2-

MiniMax

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)Cache Write ($/1M tokens)
MiniMax-H3 (H3-Context-IR)Official$0.9$3.6--
MiniMax-M3Official≤512K: $0.3 / ∞: $1.2≤512K: $1.2 / ∞: $4.8≤512K: $0.06 / ∞: $0.24-
MiniMax-M2.7-highspeedOfficial$0.6$2.4$0.06$0.375
MiniMax-M2.7Official$0.3$1.2$0.06$0.375
MiniMax-M2.5-highspeedOfficial$0.6$2.4$0.03$0.375
MiniMax-M2.5Official$0.3$1.2$0.03$0.375
MiniMax-M2.1Official$0.3$1.2$0.03$0.375
M2-herOfficial$0.3$1.2--

MiniMax-H3 token pricing above applies to the /v2/h3_context_ir prompt-enhancement endpoint only (billed by usage.prompt_tokens / completion_tokens). Video generation/regeneration is billed by seconds and input image count below.

Video Generation (MiniMax-H3)

MetricModePrice
Video Seconds ($/second)generation_2k$0.13
Video Seconds ($/second)generation_768p$0.08
Video Seconds ($/second)regeneration_2k$0.05
Video Seconds ($/second)default (mode unspecified)$0.13
Generated Images ($/image)generation_2k$0.04
Generated Images ($/image)generation_768p$0.04
Generated Images ($/image)regeneration_2k$0.025
Generated Images ($/image)default (mode unspecified)$0.04

"Video Seconds" corresponds to usage.total_seconds in the Query Task response; "Generated Images" corresponds to usage.input_image_count (billed per input image, e.g. first/last-frame or reference images).

GLM (Zhipu)

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)
glm-5.3-flash NewOfficial$0.075$0.25$0.015
glm-ocrOfficial$0.03--
glm-5.3 NewOfficial$1.4$4.4$0.26
glm-5.2Official$1.4$4.4$0.26
glm-5.1Official$1.4$4.4$0.26
glm-5-turboOfficial$1.2$4$0.24
glm-5v-turboOfficial$1.2$4$0.24
glm-5-codeOfficial$1.2$5$0.3
glm-5 DeprecatedOfficial$1$3.2$0.2
glm-4.7Official$0.6$2.2$0.11
glm-4.7-flashxOfficial$0.07$0.4$0.01
glm-4.6v NewOfficial$0.30$0.90$0.05
glm-4.6 DeprecatedOfficial$0.6$2.2$0.11
autoglm-phone-multilingualOfficial$0.1$0.1-

Tongyi Wanxiang (Wan)

Language Model

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)Cache Write ($/1M tokens)
qwen3.8-max NewOfficial$2$6$0.25$2.5
qwen3.8-flash NewOfficial$0.15$0.47$0.016$0.2
qwen-plusOfficial≤256K: $0.4 / ≤1M: $1.2Default: ≤256K: $1.2 / ≤1M: $3.6≤256K: $0.08 / ≤1M: $0.24≤256K: $0.5 / ≤1M: $1.5

qwen-plus thinking mode output price: ≤256K: $4 / ≤1M: $12 ($/1M tokens)

Embedding Models

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)
text-embedding-v4 NewOfficial$0.07--

Video Generation

ModelChannel480p ($/sec)720p ($/sec)1080p ($/sec)
wan2.5-t2v-previewOfficial$0.05$0.1$0.15
wan2.5-i2v-previewOfficial$0.05$0.1$0.15
wan2.2-t2v-plusOfficial$0.02-$0.1
wan2.2-i2v-plusOfficial$0.02-$0.1
wan2.2-i2v-flashOfficial$0.015$0.036-

ByteDance (Seed)

Language Models

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)
deepseek-r1-250528Official$1.35$5.4$0.27
deepseek-v3-1-250821Official$0.56$1.68$0.112
deepseek-v3-2-251201Official$0.28$0.42$0.056

Image Generation

ModelChannelPrice ($/image)
seedream-4-5-251128Official$0.04
seedream-4-0-250828Official$0.03
seededit-3-0-i2i-250628Official$0.03

Video Generation (Seedance)

ModelChannelPrice ($/generation)
seedance-1-5-pro-251215OfficialStandard+Audio: $2.4 / Standard: $1.2 / Draft+Audio: $1.44 / Draft: $0.84
seedance-1-0-pro-250528Official$2.5
seedance-1-0-pro-fast-251015Official$1
seedance-1-0-lite-i2v-250428Official$1.8
seedance-1-0-lite-t2v-250428Official$1.8

Video Generation (Seedance 2.0)

ModelChannelPrice ($/sec)
dreamina-seedance-2-0-260128Official1080p: $0.05–0.10
dreamina-seedance-2-0-fast-260128Official720p: $0.01–0.02

DeepSeek

Language Models

ModelChannelInput ($/1M tokens)Output ($/1M tokens)Cache Read ($/1M tokens)
deepseek-v4-proOfficial$1.74$3.48$0.015
deepseek-v4-flashOfficial$0.14$0.28$0.003

deepseek-flash (DeepSeek-V4.1-Flash)

deepseek-flash is the platform model ID for DeepSeek-V4.1-Flash and supports a 1M-token context window.

MeterConditionPrice ($/1M tokens)
Input Tokens--$0.15
Input TokensPeak: true$0.30
Cache Read Input Tokens--$0.003
Cache Read Input TokensPeak: true$0.006
Output Tokens--$0.60
Output TokensPeak: true$1.20

Happyhorse

Video Generation

ModelChannel720P ($/sec)1080P ($/sec)
happyhorse-1.0-t2vOfficial$0.12$0.22
happyhorse-1.0-i2vOfficial$0.12$0.22
happyhorse-1.0-r2vOfficial$0.12$0.22
happyhorse-1.0-video-editOfficial$0.12$0.22

The above prices are for reference only. Actual charges are subject to the system's billing.

This documentation is licensed under CC BY-SA 4.0.