specification: FinOps Framework specificationVersion: '1.0' schema: https://www.finops.org/framework/ provider: Moonshot AI providerId: moonshot-ai created: '2026-05-08' # Provenance stamped 2026-08-11: this artifact was written by the API Evangelist # bulk sweep dated 2026-05-08, not harvested from the provider. See roadmap#35. method: generated modified: '2026-05-08' reconciled: true tags: - AI - LLM - Inference - Long Context - Kimi - FinOps - Cost Management - FOCUS description: >- FOCUS-aligned FinOps profile for Moonshot AI / Kimi. Billing model is prepaid recharge with pay-as-you-go consumption metered against per-million-token rates per model and per direction (input cache-hit, input cache-miss, output). Batch jobs are billed at 50% of online rates. notes: >- Prepaid balance maps to FOCUS Purchase / Adjustment categories; consumption maps to Usage. Document extraction and storage are free at reconciliation; expect them to monetize later. sources: - https://platform.kimi.ai/docs/pricing/chat - https://platform.kimi.ai/docs/pricing/batch - https://platform.kimi.ai/docs/pricing/limits - https://focus.finops.org/focus-specification/v1-3/ alignedWith: framework: FinOps Foundation Framework frameworkUrl: https://www.finops.org/framework/ dataSpec: FOCUS dataSpecVersion: '1.3' dataSpecUrl: https://focus.finops.org/focus-specification/v1-3/ publisherName: Moonshot AI serviceCategory: AI and Machine Learning billingModel: pricingCategory: Usage Based billingFrequency: On-Demand (Recharge) billingCurrency: CNY/USD chargeCategories: - Usage - Purchase - Adjustment focusColumns: ServiceName: Moonshot AI Platform ServiceCategory: AI and Machine Learning ProviderName: Moonshot AI PublisherName: Moonshot AI InvoiceIssuerName: Moonshot AI BillingCurrency: CNY ChargeCategory: Usage meters: - name: input_tokens_cache_miss description: Input tokens billed when cache miss occurs. unit: tokens aggregation: sum dimensions: - model - account - name: input_tokens_cache_hit description: Discounted input tokens served from prompt cache. unit: tokens aggregation: sum dimensions: - model - account - name: output_tokens description: Generated output tokens. unit: tokens aggregation: sum dimensions: - model - account - name: batch_tokens description: Batch endpoint tokens (50% discount versus online). unit: tokens aggregation: sum dimensions: - model - account - name: web_search_calls description: Web search tool invocations. unit: calls aggregation: sum dimensions: - account principles: - name: Visibility description: Use the Balance API and console usage exports to track recharge balance and consumption. - name: Allocation description: Tag API keys per workload/team to attribute model spend. - name: Optimization description: Prefer prompt caching, batch endpoints (50% off), shorter contexts, and right-sized models. - name: Accountability description: Monitor tier ceiling utilization and renegotiate at $3K/$10K cumulative-recharge tier breaks. maintainers: - FN: Kin Lane email: kin@apievangelist.com