

Fireworks AI
Fireworks AI pricing runs two meters: per token on serverless, and per GPU-second on on-demand deployments.
Updated on:
Fireworks AI pricing: the token rate held while the GPU rate doubled
Fireworks AI pricing runs two meters: per token on serverless, and per GPU-second on on-demand deployments. Both draw down a prepaid balance rather than billing in arrears. The token rates sit where the market sits. The GPU rates don't. An H100 hour costs $8.00 today against $4.00 in January 2026, so the meter has moved further than any model you could pick.
Key takeaways
GPT-OSS 120B costs $0.15 per million input tokens and $0.60 output, matching Groq, Together AI, Bedrock and Nebius, but cached input runs $0.015, a 90% cut against Groq's 50%.
An H100 hour costs $8.00 on-demand, up from $4.00 in January 2026. That's 46% above Together AI's $5.49 dedicated rate and twice its $3.99 cluster rate.
Pinning work to a geography costs a flat 1.5x on both meters, on region-restricted deployments and US-only serverless variants alike.
Billing went prepaid on 1 July 2026, so there's no overage. A zero balance pauses serverless, deployments and training at once.
Fireworks AI pricing in 2026
Mode | Unit | Rate | Notes |
|---|---|---|---|
Serverless | Per 1M tokens | GPT-OSS 120B $0.15 / $0.015 cached / $0.60 | Three dimensions |
Serverless priority | Per 1M tokens | 1.2x to 1.5x standard, by model | US-only variants 1.5x |
Batch | Per 1M tokens | 50% of serverless, input and output | |
On-demand | Per GPU-hour, by the second | H100 $8.00, B200 $13.00, GB300 $20.00 | 1.5x if pinned |
Managed training | Per 1M training tokens | $0.50 to $40.00 by size and method | |
Reserved | Custom | Lower GPU-hour prices, about a 1 year term | Sales only |
What Fireworks AI actually meters
Fireworks meters tokens on one product and GPU time on the other, and the two never share a line. Serverless splits every request three ways: input, cached input from the prompt cache, and output. Embeddings bill on input alone, $0.008 to $0.10 per million. Models without a named row fall back to a size-based rate covering input and output alike: $0.10 under 4B, $0.20 to 16B, $0.90 above.
On-demand deployments switch the meter to hardware. Charges start when a replica begins accepting requests and run whether or not calls arrive, and each extra GPU adds its full rate. Deployments scale to zero by default, so an idle one costs nothing, though Fireworks says plainly that turning that off means continuous charges.
Two things deserve care. The rate card prints an hourly and a per-minute figure, and the per-minute one rounds up at the third decimal: $0.134 times 60 is $8.04, not the $8.00 printed beside it, so budget from the hourly figure. And geography costs 1.5x on both meters, putting a region-restricted H100 at $12.00 an hour and GLM 5.3 (US) at $2.10 per million input against $1.40.
How credits work
Credits are the only way to pay, and they behave as a balance rather than a wallet with rules. Fireworks moved self-serve accounts to prepaid credits on 1 July 2026, and serverless, deployments and training all draw on the same balance.
Auto Reload buys more when the balance drops to a minimum you choose. With it off, a zero balance pauses usage until you top up. Fireworks publishes no expiry date, no rollover rule and no deduction order, which is defensible with one credit type but leaves a gap for anyone modelling a year out.
What happens when you hit the limit
Fireworks pauses rather than charging you, at three gates. A zero balance with Auto Reload off stops all usage. So does the monthly spend limit, and reaching 100% of it pauses serverless, deployments and training together. Adding credits doesn't raise that limit, and Fireworks devotes a whole FAQ to accounts suspended with credit still on them.
Rate limits bite earlier. Accounts with no payment method or no credits get 10 requests a minute. With both, the account-wide ceiling is a fixed 6,000 RPM that no spend tier raises, though your Tier 1 to Tier 4 spend tier does cap how far the serverless token ceilings climb. Those adaptive limits grow 25% when usage passes half the current limit and reset after 72 quiet hours. On-demand deployments sit outside those ceilings, so a 429 there means busy GPUs.
How Fireworks AI's pricing has changed
Date | Milestone | Source |
|---|---|---|
1 Oct 2026 | DeepSeek V4.1 Flash output goes $0.66 to $1.20 per 1M, input $0.22 to $0.30 | Vendor |
1 Sep 2026 | H100 and H200 go $7.00 to $8.00 an hour, B200 $10.00 to $13.00 | Vendor |
1 Jul 2026 | Self-serve accounts move from post-paid invoicing to prepaid credits | Vendor |
1 May 2026 | H100 and H200 go $6.00 to $7.00, B200 $9.00 to $10.00. A100 leaves the card | Vendor |
By 1 Jan 2026 | H100 cut $5.80 to $4.00, B200 $11.99 to $9.00. AMD MI300X drops off | Vendor |
8 Dec 2025 | Model pages start showing cached and uncached input rates separately | Vendor |
Customer
Sentiment Highlights
"I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig."
Pay-as-you-go Fireworks customer, Hacker News, August 2026
"They force you into making a top up (cause otherwise you are by default suspended) then tell you that almost no model is available on serverless and you literally have to rent a GPU (pay more) to use the simplest most basic models and obviously pay them a gazilion dollars."
Mahmoud Hosseinypour, Trustpilot review of Fireworks, July 2026
Explore other providers

Together AI
Infrastructure Platform
Together AI pricing charges three meters for the same model: per token on serverless, per GPU-minute on dedicated endpoints, and per token trained on fine-tuning.

Ideogram
Image / Video Generation
Ideogram pricing sells a monthly bucket of priority credits, and priority buys speed rather than access.

Baseten
Infrastructure Platform
Baseten pricing meters a minute of instance time, and almost every minute a container is alive counts: image builds, cold starts and idle warm replicas all bill.
How much does Fireworks AI cost?
Does Fireworks AI have a free tier?
Is Fireworks AI cheaper than Together AI?
Do Fireworks AI credits expire?
























