Model pricing

Provider prices in USD per 1,000,000 tokens, Standard global service. These are reference rates, not a Pumpkin invoice or a quote for shared-account usage.

Model capabilities · Editor seat pricing · Catalog JSON

With your own keys, your provider bills you directly. Shared Fireworks requests pass through Pumpkin using its server key; these provider rates do not define shared-account billing terms. This catalog does not add credit billing or automatic reloads. Tool charges, other service tiers, regions and taxes can differ. Missing prices mean Unknown, never free.

Reviewed .

Catalog revisionebfc771294f78902ba21009cf9f4792859e5bf9c8cc0bcd56db795618ef348d3

Fireworks live discovery: available; last successful fetch 2026-09-23T09:56:25Z.

Active Standard token rates; cache writes replace ordinary input charges, not an extra fee on the same tokens.
Provider / modelInputOutputCache readCache writeQualifications and sources
Claude Fable 5.1anthropic · defaultclaude-fable-5-1 $10$50$0.25$12.5
Details and sources

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

cacheWrite is the 5-minute TTL rate. One-hour cache write: $20 per million tokens. Cache writes replace ordinary input charges, not an additive fee.

Standard global routing; Batch, fast modes and regional processing differ.

Standard Messages output limit; thinking shares output budget.

Adaptive thinking is always on; forced tool choice is unsupported.

https://platform.claude.com/docs/en/models/fable-5-1/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing
Claude Opus 5.5anthropic · defaultclaude-opus-5-5 $4$20$0.2$5
Details and sources

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

cacheWrite is the 5-minute TTL rate. One-hour cache write: $8 per million tokens. Cache writes replace ordinary input charges, not an additive fee.

Cache read is $0.20 per million tokens (0.05x base input).

Standard global routing; Batch, fast modes and regional processing differ.

Standard Messages output limit; thinking shares output budget.

Adaptive thinking is always on; forced tool choice is unsupported. Default effort is medium.

Cache read is 0.05x base input, not the usual 0.1x.

https://platform.claude.com/docs/en/models/opus-5-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing
Claude Opus 5anthropic · defaultclaude-opus-5 $5$25$0.5$6.25
Details and sources

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

cacheWrite is the 5-minute TTL rate. One-hour cache write: $10 per million tokens. Cache writes replace ordinary input charges, not an additive fee.

Standard global routing; Batch, fast modes and regional processing differ.

Standard Messages output limit; thinking shares output budget.

https://platform.claude.com/docs/en/models/opus-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing
Claude Sonnet 5anthropic · defaultclaude-sonnet-5 $2$10$0.2$2.5
Details and sources

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

cacheWrite is the 5-minute TTL rate. One-hour cache write: $4 per million tokens. Cache writes replace ordinary input charges, not an additive fee.

Standard global routing; Batch, fast modes and regional processing differ.

Standard Messages output limit; thinking shares output budget.

https://platform.claude.com/docs/en/models/sonnet-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing
Claude Haiku 4.5anthropic · defaultclaude-haiku-4-5 $1$5$0.1$1.25
Details and sources

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

cacheWrite is the 5-minute TTL rate. One-hour cache write: $2 per million tokens. Cache writes replace ordinary input charges, not an additive fee.

Standard global routing; Batch, fast modes and regional processing differ.

Standard Messages output limit; thinking shares output budget.

https://platform.claude.com/docs/en/models/haiku-4-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing
Grok 4.7xai · defaultgrok-4.7 $2$6$0.5Unknown Higher full-request rates at 200000+ prompt/input tokens.
Details and sources

Long-context input $4, output $12, cache read $1, cache write Unknown.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when prompt tokens >=200000.

Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free.

Provider documents no separate text output limit. maxOutput=128000 is Pumpkin request policy, not a provider maximum.

Reasoning cannot be disabled.

https://docs.x.ai/developers/models/grok-4.7https://docs.x.ai/developers/model-capabilities/images/understandinghttps://docs.x.ai/developers/pricing
Grok 4.6xai · defaultgrok-4.6 $2$6$0.5Unknown Higher full-request rates at 200000+ prompt/input tokens.
Details and sources

Long-context input $4, output $12, cache read $1, cache write Unknown.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when prompt tokens >=200000.

Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free.

Provider documents no separate text output limit. maxOutput=128000 is Pumpkin request policy, not a provider maximum.

Reasoning cannot be disabled.

https://docs.x.ai/developers/models/grok-4.6https://docs.x.ai/developers/model-capabilities/images/understandinghttps://docs.x.ai/developers/pricing
GPT-5.6 Solopenai · defaultgpt-5.6-sol $4$20$0.4$5 Higher full-request rates at 272001+ prompt/input tokens.
Details and sources

Long-context input $8, output $30, cache read $0.8, cache write $10.

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001).

Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free.

Maximum input 922000 tokens; reasoning and visible output share the output budget.

Promotional prices at least through 2026-11-21; subsequent rates unknown.

https://developers.openai.com/api/docs/models/gpt-5.6-solhttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing
GPT-5.6 Terraopenai · defaultgpt-5.6-terra $2$12$0.2$2.5 Higher full-request rates at 272001+ prompt/input tokens.
Details and sources

Long-context input $4, output $18, cache read $0.4, cache write $5.

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001).

Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free.

Maximum input 922000 tokens; reasoning and visible output share the output budget.

https://developers.openai.com/api/docs/models/gpt-5.6-terrahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing
GPT-5.6 Lunaopenai · defaultgpt-5.6-luna $0.2$1.2$0.02$0.25 Higher full-request rates at 272001+ prompt/input tokens.
Details and sources

Long-context input $0.4, output $1.8, cache read $0.04, cache write $0.5.

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001).

Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free.

Maximum input 922000 tokens; reasoning and visible output share the output budget.

https://developers.openai.com/api/docs/models/gpt-5.6-lunahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing
GPT-6 Astraopenai · defaultgpt-6-astra $10$50$1$12.5 Higher full-request rates at 272001+ prompt/input tokens.
Details and sources

Long-context input $20, output $75, cache read $2, cache write $25.

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001).

Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free.

Maximum input 922000 tokens; reasoning and visible output share the output budget.

Tool calling requires Responses; Chat Completions supports text only.

https://developers.openai.com/api/docs/models/gpt-6-astrahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing
Muse Spark 1.3meta · defaultmuse-spark-1.3 $1.25$4.25$0.15Unknown
Details and sources

Web search $0.0025 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

Standard tier (prompts not used for training). The -contributor tier is cheaper but trains on prompts and is deliberately not cataloged.

No long-context premium: the same rate applies at any context fill. Cached input is $0.15/M; Meta lists no separate cache-write price (omission means Unknown, not free).

Reasoning cannot be turned off (effort "none" returns HTTP 400); reasoning shares the output budget with visible text.

Meta documents no per-request output ceiling; 65536 is a safe request cap below the 1,048,576 context window (input and output share one budget).

https://dev.meta.ai/docs/modelshttps://dev.meta.ai/docs/pricing-rate-limitshttps://dev.meta.ai/docs/reasoninghttps://dev.meta.ai/docs/image-understanding
DeepSeek V4 Pro 0813fireworks · opt-inaccounts/fireworks/models/deepseek-v4-pro-0813Availability: availableServerless retirement: Sep 25, 2026 $1.32$3.96$0.044Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded.

https://app.fireworks.ai/models/fireworks/deepseek-v4-pro-0813https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Kimi K3fireworks · opt-inaccounts/fireworks/models/kimi-k3Availability: available $3$15$0.3Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

https://fireworks.ai/models/fireworks/kimi-k3https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
GLM 5.3fireworks · opt-inaccounts/fireworks/models/glm-5p3Availability: available $1.4$4.4$0.26Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

https://fireworks.ai/models/fireworks/glm-5p3https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
GLM 5.2fireworks · opt-inaccounts/fireworks/models/glm-5p2Availability: availableServerless retirement: Sep 25, 2026 $1.4$4.4$0.14Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded.

https://fireworks.ai/models/fireworks/glm-5p2https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Kimi K2.7 Codefireworks · opt-inaccounts/fireworks/models/kimi-k2p7-codeAvailability: availableServerless retirement: Sep 25, 2026 $0.95$4$0.19Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded.

https://fireworks.ai/models/fireworks/kimi-k2p7-codehttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Kimi K2.6fireworks · opt-inaccounts/fireworks/models/kimi-k2p6Availability: availableServerless retirement: Sep 25, 2026 $0.95$4$0.16Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded.

https://fireworks.ai/models/fireworks/kimi-k2p6https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Qwen 3.8 2.4T A95Bfireworks · opt-inaccounts/fireworks/models/qwen3p8-2p4t-a95bAvailability: unavailable UnknownUnknownUnknownUnknown
Details and sources

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Public model page says serverless not supported. No substitution of qwen3p8-max pricing.

https://fireworks.ai/models/fireworks/qwen3p8-2p4t-a95bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
MiniMax M3fireworks · opt-inaccounts/fireworks/models/minimax-m3Availability: available $0.3$1.2$0.06Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Exact serving-model feature metadata says no images despite generic multimodal architecture prose.

https://fireworks.ai/models/fireworks/minimax-m3https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
DeepSeek V4.1 Flashfireworks · opt-inaccounts/fireworks/models/deepseek-v4p1-flashAvailability: available $0.22$0.66$0.007Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Scheduled 2026-10-01T00:00:00Z: input $0.3, output $1.2, cache read $0.006, cache write Unknown (revision 2026-10-01).

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

https://app.fireworks.ai/models/fireworks/deepseek-v4p1-flashhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Nemotron 3 Ultra NVFP4fireworks · opt-inaccounts/fireworks/models/nemotron-3-ultra-nvfp4Availability: available $0.6$2.4$0.12Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

https://fireworks.ai/models/fireworks/nemotron-3-ultra-nvfp4https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
DeepSeek V4 Flash 0731fireworks · opt-inaccounts/fireworks/models/deepseek-v4-flash-0731Availability: availableServerless retirement: Sep 25, 2026 $0.22$0.66$0.007Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded.

https://app.fireworks.ai/models/fireworks/deepseek-v4-flash-0731https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
DeepSeek V4 Flash Vision Experimentalfireworks · opt-inaccounts/fireworks/models/deepseek-v4-flash-vision-expAvailability: availableServerless retirement: Sep 25, 2026 $0.22$0.66$0.007Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded.

https://app.fireworks.ai/models/fireworks/deepseek-v4-flash-vision-exphttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
GPT OSS 120Bfireworks · opt-inaccounts/fireworks/models/gpt-oss-120bAvailability: available $0.15$0.6$0.015Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

https://fireworks.ai/models/fireworks/gpt-oss-120bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
GLM 5.3 Flashfireworks · opt-inaccounts/fireworks/models/glm-5p3-flashAvailability: available $0.15$0.5$0.03Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

https://fireworks.ai/models/fireworks/glm-5p3-flashhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Muse Glimmer 30Bfireworks · opt-inaccounts/fireworks/models/muse-glimmer-30bAvailability: availableServerless retirement: Sep 25, 2026 $0.35$1.5$0.04Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded.

https://fireworks.ai/models/fireworks/muse-glimmer-30bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Nemotron Lightning 3.5 30B A3Bfireworks · opt-inaccounts/fireworks/models/nemotron-lightning-3p5-30b-a3bAvailability: available $0.05$0.2$0.01Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

https://fireworks.ai/models/fireworks/nemotron-lightning-3p5-30b-a3bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
Inklingfireworks · opt-inaccounts/fireworks/models/inklingAvailability: available $1$4.05$0.17Unknown
Details and sources

Price revision 2026-09-22; verified Sep 22, 2026.

Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate.

Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning.

Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits.

https://fireworks.ai/models/fireworks/inklinghttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models
GPT-6 Solopenai · defaultgpt-6-sol $2$10$0.2$2.5 Higher full-request rates at 272001+ prompt/input tokens.
Details and sources

Long-context input $4, output $15, cache read $0.4, cache write $5.

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001).

Cache writes are 1.25x uncached input; cache reads are 0.1x. Batch/Flex are half Standard; Fast is twice the applicable rates. Regional processing adds 10% where available. EU data residency requires Standard processing.

Use Responses for built-in tools and function calling. Chat Completions function calling requires reasoning_effort=none.

Exact model-specific image resize limits are not stated in the reviewed image guide; Pumpkin uses its conservative client policy, not an inferred provider maximum.

https://developers.openai.com/api/docs/models/gpt-6-solhttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing
GPT-6 Lunaopenai · defaultgpt-6-luna $0.1$0.5$0.01$0.125 Higher full-request rates at 272001+ prompt/input tokens.
Details and sources

Long-context input $0.2, output $0.75, cache read $0.02, cache write $0.25.

Web search $0.01 per call, plus applicable tokens.

Price revision 2026-09-22; verified Sep 22, 2026.

Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001).

Cache writes are 1.25x uncached input; cache reads are 0.1x. Batch/Flex are half Standard; Fast is twice the applicable rates. Regional processing adds 10% where available. EU data residency requires Standard processing.

Use Responses for built-in tools and function calling. Chat Completions function calling requires reasoning_effort=none.

Exact model-specific image resize limits are not stated in the reviewed image guide; Pumpkin uses its conservative client policy, not an inferred provider maximum.

https://developers.openai.com/api/docs/models/gpt-6-lunahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing