specification: FinOps Framework specificationVersion: '1.0' schema: https://www.finops.org/framework/ provider: Predibase providerId: predibase created: '2026-06-20' modified: '2026-06-20' reconciled: true tags: - AI - LLM - Fine-Tuning - Inference - LoRA - FinOps - Cost Management - FOCUS description: >- FinOps view of Predibase spend. Predibase bills three usage meters: serverless inference per token (scaled by model size), batch inference at a flat per-million-token rate, fine-tuning (training) per token of training data scaled by base-model size, and dedicated deployments per GPU-hour by accelerator. LoRAX multi-LoRA serving packs many adapters onto one GPU, reducing dedicated serving cost versus one deployment per fine-tuned model. notes: >- Per-model and per-accelerator rates are subject to change; verify against the Predibase pricing page during reconciliation. Predibase was acquired by Rubrik in June 2025. sources: - https://predibase.com/pricing - https://docs.predibase.com - https://focus.finops.org/focus-specification/v1-3/ alignedWith: framework: FinOps Foundation Framework frameworkUrl: https://www.finops.org/framework/ dataSpec: FOCUS dataSpecVersion: '1.3' dataSpecUrl: https://focus.finops.org/focus-specification/v1-3/ publisherName: Predibase serviceCategory: AI and Machine Learning billingModel: pricingCategory: Usage-Based billingFrequency: Monthly billingCurrency: USD chargeCategories: - Usage - Purchase - Adjustment focusColumns: ServiceName: Predibase ServiceCategory: AI and Machine Learning ProviderName: Predibase PublisherName: Predibase InvoiceIssuerName: Predibase BillingCurrency: USD ChargeCategory: Usage PricingCategory: Usage-Based meters: - name: serverless_inference_tokens description: Tokens served via serverless / shared endpoints, billed per 1M tokens scaled by model size. unit: tokens aggregation: sum dimensions: - account - model - deployment - name: batch_inference_tokens description: Tokens served via batch inference jobs, billed at a flat per-1M-token rate. unit: tokens aggregation: sum dimensions: - account - model - name: finetuning_training_tokens description: Training tokens consumed by supervised / GRPO fine-tuning jobs, billed per 1M tokens scaled by base-model size. unit: tokens aggregation: sum dimensions: - account - base_model - task - name: dedicated_gpu_hours description: GPU-hours consumed by dedicated / private deployments, billed per accelerator type (A10, A100, H100). unit: hours aggregation: sum dimensions: - account - deployment - accelerator principles: - name: Visibility description: Pull Predibase usage exports; inspect token burn per deployment and GPU-hours per accelerator. - name: Allocation description: Tag API tokens and deployments per workload/team; map adapters and repos to internal cost centers. - name: Optimization description: Prefer serverless / batch for spiky or non-realtime traffic; consolidate many fine-tuned adapters onto one LoRAX-enabled deployment instead of one deployment per model; right-size the GPU accelerator; pick smaller base models when sufficient. - name: Accountability description: Assign owners per project/deployment; review token and GPU-hour spend monthly against budget. maintainers: - FN: Kin Lane email: kin@apievangelist.com