slug: scalable-inference-serving provider: Scalable Inference Serving generated_by: planning/capability-mapping/scripts/classify_capabilities.py model: claude-opus-5 frame: - Software & Technology min_confidence: 0.7 capability_model: source: https://github.com/vincentmakes/turbo-ea-capabilities license: CC-BY-4.0 attribution: Turbo EA Capabilities by Vincent Verdet — Turbo EA, https://github.com/vincentmakes/turbo-ea-capabilities, CC BY 4.0 notice: NOTICE edge_count: 2 edges: - tag: Inference spec_file: scalable-inference-serving-inference-api-openapi.yml capability_id: BC-610.60 capability_id_l1: BC-610 capability_name: Artificial Intelligence Management confidence: 0.78 evidence: POST /v2/models/{model_name}/infer RunInference Run Model Inference reason: The operations execute ML model inference against served model versions (InferenceRequest/InferenceResponse, TensorDatatype), which is the serving leg of the AI/ML model lifecycle under Artificial Intelligence Management. Some ambiguity remains as runtime inference could also be read as application operations. - tag: Models spec_file: scalable-inference-serving-models-api-openapi.yml capability_id: BC-610.60 capability_id_l1: BC-610 capability_name: Artificial Intelligence Management confidence: 0.72 evidence: GET /v2/models/{model_name}/versions/{model_version} GetModelVersionMetadata Get Model Version Metadata reason: Operations expose served ML models and their versions, including readiness and metadata, which is model lifecycle/MLOps territory under Artificial Intelligence Management.