specification: FinOps Framework specificationVersion: '1.0' schema: https://www.finops.org/framework/ provider: Nscale providerId: nscale created: '2026-06-21' modified: '2026-06-21' reconciled: false tags: - AI - GPU - Inference - Serverless - Cloud Compute - FinOps - Cost Management - FOCUS description: >- FinOps view of Nscale spend. Nscale bills usage-based per-token rates for serverless text inference (input and output), per-megapixel rates for image generation, and per-GPU-hour rates for compute instances, nodes, and clusters. Reserved capacity and enterprise commitments reduce effective GPU-hour rates. notes: >- Per-model and per-GPU-hour rates change frequently and were not all individually confirmed (reconciled false); verify against the Nscale pricing page and console during reconciliation. sources: - https://www.nscale.com/pricing - https://www.nscale.com/product/serverless - https://www.nscale.com/product/gpu-nodes - https://focus.finops.org/focus-specification/v1-3/ alignedWith: framework: FinOps Foundation Framework frameworkUrl: https://www.finops.org/framework/ dataSpec: FOCUS dataSpecVersion: '1.3' dataSpecUrl: https://focus.finops.org/focus-specification/v1-3/ publisherName: Nscale serviceCategory: AI and Machine Learning billingModel: pricingCategory: Usage-Based billingFrequency: Monthly billingCurrency: USD chargeCategories: - Usage - Purchase - Adjustment focusColumns: ServiceName: Nscale Serverless Inference and GPU Compute ServiceCategory: AI and Machine Learning ProviderName: Nscale PublisherName: Nscale InvoiceIssuerName: Nscale BillingCurrency: USD ChargeCategory: Usage PricingCategory: Usage-Based meters: - name: input_tokens description: Tokens sent in serverless chat / completions requests, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: output_tokens description: Tokens generated by serverless inference, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - api - name: embedding_tokens description: Tokens embedded via the embeddings endpoint, billed per 1M tokens per model. unit: tokens aggregation: sum dimensions: - account - model - name: image_megapixels description: Megapixels of generated imagery, billed per megapixel and scaled by diffusion steps. unit: megapixels aggregation: sum dimensions: - account - model - name: gpu_hours description: GPU-hours consumed by compute instances, nodes, and clusters, billed per accelerator type. unit: gpu_hours aggregation: sum dimensions: - account - accelerator - region principles: - name: Visibility description: Pull Nscale usage and billing exports; inspect per-model token burn and per-accelerator GPU-hours. - name: Allocation description: Tag API keys and projects per workload/team; map serverless and compute spend to internal cost centers. - name: Optimization description: Pick smaller open-source serverless models when sufficient; cap diffusion steps and image resolution; use reserved GPU capacity for steady compute workloads. - name: Accountability description: Assign owners per project/key; review token and GPU-hour spend monthly against budget. maintainers: - FN: Kin Lane email: kin@apievangelist.com