Model pricing
Provider prices in USD per 1,000,000 tokens, Standard global service. These are reference rates, not a Pumpkin invoice or a quote for shared-account usage.
Model capabilities · Editor seat pricing · Catalog JSON
With your own keys, your provider bills you directly. Shared Fireworks requests pass through Pumpkin using its server key; these provider rates do not define shared-account billing terms. This catalog does not add credit billing or automatic reloads. Tool charges, other service tiers, regions and taxes can differ. Missing prices mean Unknown, never free.
Reviewed .
Catalog revision
ebfc771294f78902ba21009cf9f4792859e5bf9c8cc0bcd56db795618ef348d3Fireworks live discovery: available; last successful fetch 2026-09-23T09:56:25Z.
| Provider / model | Input | Output | Cache read | Cache write | Qualifications and sources |
|---|---|---|---|---|---|
Claude Fable 5.1anthropic · defaultclaude-fable-5-1 |
$10 | $50 | $0.25 | $12.5 |
Details and sourcesWeb search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. cacheWrite is the 5-minute TTL rate. One-hour cache write: $20 per million tokens. Cache writes replace ordinary input charges, not an additive fee. Standard global routing; Batch, fast modes and regional processing differ. Standard Messages output limit; thinking shares output budget. Adaptive thinking is always on; forced tool choice is unsupported. https://platform.claude.com/docs/en/models/fable-5-1/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing |
Claude Opus 5.5anthropic · defaultclaude-opus-5-5 |
$4 | $20 | $0.2 | $5 |
Details and sourcesWeb search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. cacheWrite is the 5-minute TTL rate. One-hour cache write: $8 per million tokens. Cache writes replace ordinary input charges, not an additive fee. Cache read is $0.20 per million tokens (0.05x base input). Standard global routing; Batch, fast modes and regional processing differ. Standard Messages output limit; thinking shares output budget. Adaptive thinking is always on; forced tool choice is unsupported. Default effort is medium. Cache read is 0.05x base input, not the usual 0.1x. https://platform.claude.com/docs/en/models/opus-5-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing |
Claude Opus 5anthropic · defaultclaude-opus-5 |
$5 | $25 | $0.5 | $6.25 |
Details and sourcesWeb search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. cacheWrite is the 5-minute TTL rate. One-hour cache write: $10 per million tokens. Cache writes replace ordinary input charges, not an additive fee. Standard global routing; Batch, fast modes and regional processing differ. Standard Messages output limit; thinking shares output budget. https://platform.claude.com/docs/en/models/opus-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing |
Claude Sonnet 5anthropic · defaultclaude-sonnet-5 |
$2 | $10 | $0.2 | $2.5 |
Details and sourcesWeb search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. cacheWrite is the 5-minute TTL rate. One-hour cache write: $4 per million tokens. Cache writes replace ordinary input charges, not an additive fee. Standard global routing; Batch, fast modes and regional processing differ. Standard Messages output limit; thinking shares output budget. https://platform.claude.com/docs/en/models/sonnet-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing |
Claude Haiku 4.5anthropic · defaultclaude-haiku-4-5 |
$1 | $5 | $0.1 | $1.25 |
Details and sourcesWeb search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. cacheWrite is the 5-minute TTL rate. One-hour cache write: $2 per million tokens. Cache writes replace ordinary input charges, not an additive fee. Standard global routing; Batch, fast modes and regional processing differ. Standard Messages output limit; thinking shares output budget. https://platform.claude.com/docs/en/models/haiku-4-5/overviewhttps://platform.claude.com/docs/en/build-with-claude/visionhttps://platform.claude.com/docs/en/about-claude/pricing |
Grok 4.7xai · defaultgrok-4.7 |
$2 | $6 | $0.5 | Unknown | Higher full-request rates at 200000+ prompt/input tokens.
Details and sourcesLong-context input $4, output $12, cache read $1, cache write Unknown. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when prompt tokens >=200000. Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free. Provider documents no separate text output limit. maxOutput=128000 is Pumpkin request policy, not a provider maximum. Reasoning cannot be disabled. https://docs.x.ai/developers/models/grok-4.7https://docs.x.ai/developers/model-capabilities/images/understandinghttps://docs.x.ai/developers/pricing |
Grok 4.6xai · defaultgrok-4.6 |
$2 | $6 | $0.5 | Unknown | Higher full-request rates at 200000+ prompt/input tokens.
Details and sourcesLong-context input $4, output $12, cache read $1, cache write Unknown. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when prompt tokens >=200000. Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free. Provider documents no separate text output limit. maxOutput=128000 is Pumpkin request policy, not a provider maximum. Reasoning cannot be disabled. https://docs.x.ai/developers/models/grok-4.6https://docs.x.ai/developers/model-capabilities/images/understandinghttps://docs.x.ai/developers/pricing |
GPT-5.6 Solopenai · defaultgpt-5.6-sol |
$4 | $20 | $0.4 | $5 | Higher full-request rates at 272001+ prompt/input tokens.
Details and sourcesLong-context input $8, output $30, cache read $0.8, cache write $10. Web search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001). Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free. Maximum input 922000 tokens; reasoning and visible output share the output budget. Promotional prices at least through 2026-11-21; subsequent rates unknown. https://developers.openai.com/api/docs/models/gpt-5.6-solhttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing |
GPT-5.6 Terraopenai · defaultgpt-5.6-terra |
$2 | $12 | $0.2 | $2.5 | Higher full-request rates at 272001+ prompt/input tokens.
Details and sourcesLong-context input $4, output $18, cache read $0.4, cache write $5. Web search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001). Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free. Maximum input 922000 tokens; reasoning and visible output share the output budget. https://developers.openai.com/api/docs/models/gpt-5.6-terrahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing |
GPT-5.6 Lunaopenai · defaultgpt-5.6-luna |
$0.2 | $1.2 | $0.02 | $0.25 | Higher full-request rates at 272001+ prompt/input tokens.
Details and sourcesLong-context input $0.4, output $1.8, cache read $0.04, cache write $0.5. Web search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001). Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free. Maximum input 922000 tokens; reasoning and visible output share the output budget. https://developers.openai.com/api/docs/models/gpt-5.6-lunahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing |
GPT-6 Astraopenai · defaultgpt-6-astra |
$10 | $50 | $1 | $12.5 | Higher full-request rates at 272001+ prompt/input tokens.
Details and sourcesLong-context input $20, output $75, cache read $2, cache write $25. Web search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001). Standard global routing; other service tiers and regions differ. Cache write/read omissions mean Unknown, not free. Maximum input 922000 tokens; reasoning and visible output share the output budget. Tool calling requires Responses; Chat Completions supports text only. https://developers.openai.com/api/docs/models/gpt-6-astrahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing |
Muse Spark 1.3meta · defaultmuse-spark-1.3 |
$1.25 | $4.25 | $0.15 | Unknown |
Details and sourcesWeb search $0.0025 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. Standard tier (prompts not used for training). The -contributor tier is cheaper but trains on prompts and is deliberately not cataloged. No long-context premium: the same rate applies at any context fill. Cached input is $0.15/M; Meta lists no separate cache-write price (omission means Unknown, not free). Reasoning cannot be turned off (effort "none" returns HTTP 400); reasoning shares the output budget with visible text. Meta documents no per-request output ceiling; 65536 is a safe request cap below the 1,048,576 context window (input and output share one budget). https://dev.meta.ai/docs/modelshttps://dev.meta.ai/docs/pricing-rate-limitshttps://dev.meta.ai/docs/reasoninghttps://dev.meta.ai/docs/image-understanding |
DeepSeek V4 Pro 0813fireworks · opt-inaccounts/fireworks/models/deepseek-v4-pro-0813Availability: availableServerless retirement: Sep 25, 2026 |
$1.32 | $3.96 | $0.044 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded. https://app.fireworks.ai/models/fireworks/deepseek-v4-pro-0813https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Kimi K3fireworks · opt-inaccounts/fireworks/models/kimi-k3Availability: available |
$3 | $15 | $0.3 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. https://fireworks.ai/models/fireworks/kimi-k3https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
GLM 5.3fireworks · opt-inaccounts/fireworks/models/glm-5p3Availability: available |
$1.4 | $4.4 | $0.26 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. https://fireworks.ai/models/fireworks/glm-5p3https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
GLM 5.2fireworks · opt-inaccounts/fireworks/models/glm-5p2Availability: availableServerless retirement: Sep 25, 2026 |
$1.4 | $4.4 | $0.14 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded. https://fireworks.ai/models/fireworks/glm-5p2https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Kimi K2.7 Codefireworks · opt-inaccounts/fireworks/models/kimi-k2p7-codeAvailability: availableServerless retirement: Sep 25, 2026 |
$0.95 | $4 | $0.19 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded. https://fireworks.ai/models/fireworks/kimi-k2p7-codehttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Kimi K2.6fireworks · opt-inaccounts/fireworks/models/kimi-k2p6Availability: availableServerless retirement: Sep 25, 2026 |
$0.95 | $4 | $0.16 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded. https://fireworks.ai/models/fireworks/kimi-k2p6https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Qwen 3.8 2.4T A95Bfireworks · opt-inaccounts/fireworks/models/qwen3p8-2p4t-a95bAvailability: unavailable |
Unknown | Unknown | Unknown | Unknown |
Details and sourcesProvider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Public model page says serverless not supported. No substitution of qwen3p8-max pricing. https://fireworks.ai/models/fireworks/qwen3p8-2p4t-a95bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
MiniMax M3fireworks · opt-inaccounts/fireworks/models/minimax-m3Availability: available |
$0.3 | $1.2 | $0.06 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Exact serving-model feature metadata says no images despite generic multimodal architecture prose. https://fireworks.ai/models/fireworks/minimax-m3https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
DeepSeek V4.1 Flashfireworks · opt-inaccounts/fireworks/models/deepseek-v4p1-flashAvailability: available |
$0.22 | $0.66 | $0.007 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Scheduled 2026-10-01T00:00:00Z: input $0.3, output $1.2, cache read $0.006, cache write Unknown (revision 2026-10-01). Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. https://app.fireworks.ai/models/fireworks/deepseek-v4p1-flashhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Nemotron 3 Ultra NVFP4fireworks · opt-inaccounts/fireworks/models/nemotron-3-ultra-nvfp4Availability: available |
$0.6 | $2.4 | $0.12 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. https://fireworks.ai/models/fireworks/nemotron-3-ultra-nvfp4https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
DeepSeek V4 Flash 0731fireworks · opt-inaccounts/fireworks/models/deepseek-v4-flash-0731Availability: availableServerless retirement: Sep 25, 2026 |
$0.22 | $0.66 | $0.007 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded. https://app.fireworks.ai/models/fireworks/deepseek-v4-flash-0731https://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
DeepSeek V4 Flash Vision Experimentalfireworks · opt-inaccounts/fireworks/models/deepseek-v4-flash-vision-expAvailability: availableServerless retirement: Sep 25, 2026 |
$0.22 | $0.66 | $0.007 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded. https://app.fireworks.ai/models/fireworks/deepseek-v4-flash-vision-exphttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
GPT OSS 120Bfireworks · opt-inaccounts/fireworks/models/gpt-oss-120bAvailability: available |
$0.15 | $0.6 | $0.015 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. https://fireworks.ai/models/fireworks/gpt-oss-120bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
GLM 5.3 Flashfireworks · opt-inaccounts/fireworks/models/glm-5p3-flashAvailability: available |
$0.15 | $0.5 | $0.03 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. https://fireworks.ai/models/fireworks/glm-5p3-flashhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Muse Glimmer 30Bfireworks · opt-inaccounts/fireworks/models/muse-glimmer-30bAvailability: availableServerless retirement: Sep 25, 2026 |
$0.35 | $1.5 | $0.04 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. Serverless retirement announced for 2026-09-25; exact cutoff time unknown. Dedicated deployments are excluded. https://fireworks.ai/models/fireworks/muse-glimmer-30bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Nemotron Lightning 3.5 30B A3Bfireworks · opt-inaccounts/fireworks/models/nemotron-lightning-3p5-30b-a3bAvailability: available |
$0.05 | $0.2 | $0.01 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. https://fireworks.ai/models/fireworks/nemotron-lightning-3p5-30b-a3bhttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
Inklingfireworks · opt-inaccounts/fireworks/models/inklingAvailability: available |
$1 | $4.05 | $0.17 | Unknown |
Details and sourcesPrice revision 2026-09-22; verified Sep 22, 2026. Standard global serverless rates. Priority, Fast and regional prices differ. No separately verified cache-write rate. Provider output limit Unknown. thinkingMode=none and the 16384 output cap describe Pumpkin adapter policy, not absence of provider reasoning. Exact context Unknown until metadata discovery; rounded public context labels are not treated as exact limits. https://fireworks.ai/models/fireworks/inklinghttps://docs.fireworks.ai/guides/querying-vision-language-modelshttps://docs.fireworks.ai/serverless/pricinghttps://docs.fireworks.ai/updates/changeloghttps://api.fireworks.ai/v1/accounts/fireworks/models |
GPT-6 Solopenai · defaultgpt-6-sol |
$2 | $10 | $0.2 | $2.5 | Higher full-request rates at 272001+ prompt/input tokens.
Details and sourcesLong-context input $4, output $15, cache read $0.4, cache write $5. Web search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001). Cache writes are 1.25x uncached input; cache reads are 0.1x. Batch/Flex are half Standard; Fast is twice the applicable rates. Regional processing adds 10% where available. EU data residency requires Standard processing. Use Responses for built-in tools and function calling. Chat Completions function calling requires reasoning_effort=none. Exact model-specific image resize limits are not stated in the reviewed image guide; Pumpkin uses its conservative client policy, not an inferred provider maximum. https://developers.openai.com/api/docs/models/gpt-6-solhttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing |
GPT-6 Lunaopenai · defaultgpt-6-luna |
$0.1 | $0.5 | $0.01 | $0.125 | Higher full-request rates at 272001+ prompt/input tokens.
Details and sourcesLong-context input $0.2, output $0.75, cache read $0.02, cache write $0.25. Web search $0.01 per call, plus applicable tokens. Price revision 2026-09-22; verified Sep 22, 2026. Long-context rates apply to all tokens in the request when input tokens >272000 (inclusive lower bound 272001). Cache writes are 1.25x uncached input; cache reads are 0.1x. Batch/Flex are half Standard; Fast is twice the applicable rates. Regional processing adds 10% where available. EU data residency requires Standard processing. Use Responses for built-in tools and function calling. Chat Completions function calling requires reasoning_effort=none. Exact model-specific image resize limits are not stated in the reviewed image guide; Pumpkin uses its conservative client policy, not an inferred provider maximum. https://developers.openai.com/api/docs/models/gpt-6-lunahttps://developers.openai.com/api/docs/guides/images-visionhttps://developers.openai.com/api/docs/pricing |