specification: FinOps Framework specificationVersion: '1.0' schema: https://www.finops.org/framework/ provider: CentML providerId: centml created: '2026-06-21' modified: '2026-06-21' reconciled: false tags: - AI - LLM - Inference - Serverless - GPU - FinOps - Cost Management - FOCUS description: >- FinOps view of CentML spend. CentML bills on a credit-based model where 1 credit equals 1 USD. Serverless inference is metered by input and output tokens with per-token rates that vary by model. Dedicated deployments are metered by GPU hardware type and duration on a per-minute basis (per-GPU-hour equivalent). Exact per-model and per-GPU-hour rates are not reconciled here. notes: >- Per-model serverless token rates and per-GPU-hour dedicated rates change over time; verify against the CentML pricing page during reconciliation. sources: - https://centml.ai/pricing/ - https://docs.centml.ai/apps/serverless - https://focus.finops.org/focus-specification/v1-3/ alignedWith: framework: FinOps Foundation Framework frameworkUrl: https://www.finops.org/framework/ dataSpec: FOCUS dataSpecVersion: '1.3' dataSpecUrl: https://focus.finops.org/focus-specification/v1-3/ publisherName: CentML serviceCategory: AI and Machine Learning billingModel: pricingCategory: Usage-Based billingFrequency: Monthly billingCurrency: USD chargeCategories: - Usage - Purchase - Adjustment focusColumns: ServiceName: CentML Platform ServiceCategory: AI and Machine Learning ProviderName: CentML PublisherName: CentML InvoiceIssuerName: CentML BillingCurrency: USD ChargeCategory: Usage PricingCategory: Usage-Based meters: - name: serverless_input_tokens description: Input tokens sent to serverless chat / text completions, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: serverless_output_tokens description: Output tokens generated by serverless endpoints, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: dedicated_gpu_hours description: GPU hours consumed by dedicated inference / compute deployments, billed per minute by hardware instance type. unit: gpu_hours aggregation: sum dimensions: - account - deployment - hardware_instance - name: credits description: CentML credits consumed, where 1 credit equals 1 USD. unit: credits aggregation: sum dimensions: - account principles: - name: Visibility description: Track credit balance and burn; inspect per-model serverless token usage and per-deployment GPU-hour consumption. - name: Allocation description: Tag API keys and deployments per workload/team; map serverless token spend and dedicated GPU hours to cost centers. - name: Optimization description: Use serverless for spiky/low-volume workloads and dedicated deployments for sustained high throughput; right-size GPU hardware and autoscaling replicas; pick smaller models when sufficient. - name: Accountability description: Assign owners per deployment/key; review credit spend monthly against budget; pause idle dedicated deployments. maintainers: - FN: Kin Lane email: kin@apievangelist.com