specification: FinOps Framework specificationVersion: '1.0' schema: https://www.finops.org/framework/ provider: Fireworks AI providerId: fireworks-ai created: '2026-05-08' # Provenance stamped 2026-08-11: this artifact was written by the API Evangelist # bulk sweep dated 2026-05-08, not harvested from the provider. See roadmap#35. method: generated modified: '2026-05-08' reconciled: true tags: - AI - LLM - Inference - Multimodal - Fine-tuning - GPU - FinOps - Cost Management - FOCUS description: >- FinOps view of Fireworks AI spend. Postpaid usage-based billing in USD, with serverless per-token rates per model, embedding rates by model size, per-1M-token training-token fine-tuning rates, per-GPU-hour on-demand deployments, and a 50% discount on Batch and Cached Input tokens. notes: >- $1 free credit at signup; postpaid invoicing thereafter. Fine-tuned model serving costs the same as the base model. sources: - https://fireworks.ai/pricing - https://docs.fireworks.ai/ - https://focus.finops.org/focus-specification/v1-3/ alignedWith: framework: FinOps Foundation Framework frameworkUrl: https://www.finops.org/framework/ dataSpec: FOCUS dataSpecVersion: '1.3' dataSpecUrl: https://focus.finops.org/focus-specification/v1-3/ publisherName: Fireworks AI serviceCategory: AI and Machine Learning billingModel: pricingCategory: Usage-Based billingFrequency: Monthly billingCurrency: USD chargeCategories: - Usage - Purchase - Adjustment focusColumns: ServiceName: Fireworks AI ServiceCategory: AI and Machine Learning ProviderName: Fireworks AI PublisherName: Fireworks AI InvoiceIssuerName: Fireworks AI BillingCurrency: USD ChargeCategory: Usage PricingCategory: Usage-Based meters: - name: input_tokens description: Tokens sent in chat / vision requests, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: cached_input_tokens description: Cached-input tokens at 50% of standard input rate. unit: tokens aggregation: sum dimensions: - account - model - name: output_tokens description: Tokens generated by the model, per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: embedding_tokens description: Tokens processed for embeddings, per 1M tokens per embedding-model size class. unit: tokens aggregation: sum dimensions: - account - model - name: rerank_documents description: Documents reranked, per model. unit: documents aggregation: sum dimensions: - account - model - name: image_generations description: Image generation requests per model and resolution. unit: images aggregation: sum dimensions: - account - model - name: audio_seconds description: Audio seconds for STT/TTS workloads per model. unit: seconds aggregation: sum dimensions: - account - model - name: batch_tokens description: Tokens consumed via the Batch API at 50% discount. unit: tokens aggregation: sum dimensions: - account - model - name: fine_tuning_tokens description: Training tokens for SFT (LoRA / full) jobs per base-model size class. unit: tokens aggregation: sum dimensions: - account - base_model - method - name: gpu_seconds description: On-demand dedicated GPU runtime per GPU class (H100, H200, B200, B300). unit: seconds aggregation: sum dimensions: - account - deployment - gpu_class principles: - name: Visibility description: Pull token / GPU-second usage from the Fireworks console; reconcile against postpaid invoices. - name: Allocation description: Tag API keys per workload/team and join usage to internal cost centers. - name: Optimization description: Use Batch (-50%) for non-realtime work; enable prompt caching (-50%); pick smaller embedding models when sufficient; right-size GPU deployments and autoscale to zero where possible. - name: Accountability description: Assign owners per deployment and per fine-tuning job; review monthly token/GPU burn vs. budget. maintainers: - FN: Kin Lane email: kin@apievangelist.com