specification: FinOps Framework specificationVersion: '1.0' schema: https://www.finops.org/framework/ provider: Groq providerId: groq created: '2026-05-08' # Provenance stamped 2026-08-11: this artifact was written by the API Evangelist # bulk sweep dated 2026-05-08, not harvested from the provider. See roadmap#35. method: generated modified: '2026-05-08' reconciled: true tags: - AI - LLM - Inference - LPU - Low Latency - FinOps - Cost Management - FOCUS description: >- FinOps view of GroqCloud spend. Groq bills usage-based per-token rates for chat / vision / reasoning per model, per-million-character rates for TTS, per-hour transcription rates for STT, per-call or per-hour rates for tools, and a 50% Batch discount. Prompt Caching gives 50% off cached input tokens. notes: >- Per-model rates are subject to change; verify against the Groq pricing page during reconciliation. sources: - https://groq.com/pricing - https://console.groq.com/docs - https://focus.finops.org/focus-specification/v1-3/ alignedWith: framework: FinOps Foundation Framework frameworkUrl: https://www.finops.org/framework/ dataSpec: FOCUS dataSpecVersion: '1.3' dataSpecUrl: https://focus.finops.org/focus-specification/v1-3/ publisherName: Groq serviceCategory: AI and Machine Learning billingModel: pricingCategory: Usage-Based billingFrequency: Monthly billingCurrency: USD chargeCategories: - Usage - Purchase - Adjustment focusColumns: ServiceName: GroqCloud ServiceCategory: AI and Machine Learning ProviderName: Groq PublisherName: Groq InvoiceIssuerName: Groq BillingCurrency: USD ChargeCategory: Usage PricingCategory: Usage-Based meters: - name: input_tokens description: Tokens sent in chat / vision / reasoning requests, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: cached_input_tokens description: Cached-input tokens billed at 50% of the standard input rate. unit: tokens aggregation: sum dimensions: - account - model - name: output_tokens description: Tokens generated, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: tts_characters description: TTS characters synthesized, billed per 1M characters per voice/model. unit: characters aggregation: sum dimensions: - account - model - name: stt_audio_hours description: Audio hours transcribed, billed per hour per Whisper variant. unit: hours aggregation: sum dimensions: - account - model - name: tool_invocations description: Tool calls (web search, Wolfram) priced per 1,000 invocations. unit: invocations aggregation: sum dimensions: - account - tool - name: tool_compute_hours description: Tool compute hours (e.g., Code Execution at $0.18/hr). unit: hours aggregation: sum dimensions: - account - tool - name: batch_tokens description: Tokens consumed via the Batch API at 50% discount. unit: tokens aggregation: sum dimensions: - account - model - name: flex_tokens description: Tokens consumed via Flex Processing tier at relaxed-latency discount. unit: tokens aggregation: sum dimensions: - account - model principles: - name: Visibility description: Pull GroqCloud usage and billing exports; inspect per-model token burn. - name: Allocation description: Tag API keys per workload/team; map to internal cost centers. - name: Optimization description: Use Batch (-50%) for non-realtime jobs; enable Prompt Caching; route lower-quality work to Flex; pick smaller open-source models when sufficient. - name: Accountability description: Assign owners per project/key; review token spend monthly against budget. maintainers: - FN: Kin Lane email: kin@apievangelist.com