specification: FinOps Framework specificationVersion: '1.0' schema: https://www.finops.org/framework/ provider: glhf providerId: glhf-chat created: '2026-06-21' modified: '2026-06-21' reconciled: false tags: - AI - LLM - Inference - Open Source Models - Hugging Face - FinOps - Cost Management - FOCUS description: >- FinOps view of glhf (glhf.chat) spend. glhf bills on a usage basis for GPU resources consumed while running open-source models through its OpenAI-compatible API, with no required subscription and no charge for boot time. Specific per-token or per-GPU-hour rates are not publicly documented, so meters below are modeled from the usage-based description and are not reconciled. notes: >- Per-token and per-GPU-hour rates are not published; verify against the glhf pricing page and account settings during reconciliation. sources: - https://glhf.chat - https://glhf.chat/pricing - https://focus.finops.org/focus-specification/v1-3/ alignedWith: framework: FinOps Foundation Framework frameworkUrl: https://www.finops.org/framework/ dataSpec: FOCUS dataSpecVersion: '1.3' dataSpecUrl: https://focus.finops.org/focus-specification/v1-3/ publisherName: glhf serviceCategory: AI and Machine Learning billingModel: pricingCategory: Usage-Based billingFrequency: Monthly billingCurrency: USD chargeCategories: - Usage - Purchase - Adjustment focusColumns: ServiceName: glhf ServiceCategory: AI and Machine Learning ProviderName: glhf PublisherName: glhf InvoiceIssuerName: glhf BillingCurrency: USD ChargeCategory: Usage PricingCategory: Usage-Based meters: - name: gpu_hours description: GPU hours consumed while running models on the auto-scaling scheduler, billed per usage (rate not publicly documented). unit: hours aggregation: sum dimensions: - account - model - name: input_tokens description: Tokens sent in chat completion requests (rate not publicly documented). unit: tokens aggregation: sum dimensions: - account - model - name: output_tokens description: Tokens generated in chat completion responses (rate not publicly documented). unit: tokens aggregation: sum dimensions: - account - model principles: - name: Visibility description: Pull glhf account usage; inspect GPU and token burn per model. - name: Allocation description: Tag API keys per workload/team; map to internal cost centers. - name: Optimization description: Choose smaller open-source models when sufficient; avoid keeping large models warm unnecessarily since usage drives cost. - name: Accountability description: Assign owners per project/key; review usage monthly against budget. maintainers: - FN: Kin Lane email: kin@apievangelist.com