aid: cumulus-labs name: Cumulus Labs description: >- Cumulus Labs is a Y Combinator (W26) company building a unified inference platform for production AI. Cumulus consolidates the eight subsystems teams normally assemble from separate vendors — an OpenAI-compatible gateway, a per-workflow router, a layered prompt/KV cache, request-level observability, continuous shadow evaluation, one-click LoRA fine-tuning, custom open-weight hosting, and the proprietary Ion inference engine running on NVIDIA Grace and Blackwell GPUs — behind a single OpenAI-compatible API at api.cumuluslabs.io/v1. Ion's custom attention kernels deliver 30-50% more throughput than stock vLLM and SGLang, and the gateway is a drop-in replacement for the OpenAI, Anthropic, LangChain, LlamaIndex, and Vercel AI SDKs — change one line of configuration and keep your existing code. url: https://raw.githubusercontent.com/api-evangelist/cumulus-labs/refs/heads/main/apis.yml x-type: company x-source: vc-portfolio x-backed-by: - y-combinator - nvidia-inception x-tier: profiled x-tier-reason: enrichment-pass accessModel: pricing: unknown onboarding: self-serve trial: false try_now: false public: false label: Self-serve signup confidence: medium source: - authentication generated: '2026-07-22' method: derived specificationVersion: '0.20' created: '2026-07-17' modified: '2026-07-18' tags: - Company - Inference - LLM - AI Infrastructure - GPU - Machine Learning - Model Serving - Fine-Tuning - API Gateway - Y Combinator image: https://cumuluslabs.io/cumulus-logo/Cumulus-White.svg apis: - name: Cumulus Inference Gateway description: >- OpenAI-compatible HTTP inference gateway. One client works against every upstream provider; per-workflow routing rules pick the model, provider, and infrastructure, with a layered exact/prefix/semantic cache, request-level observability, and the Ion engine serving open-weight models and fine-tunes. humanURL: https://docs.cumuluslabs.io/ baseURL: https://api.cumuluslabs.io/v1 tags: - Inference - LLM - OpenAI-Compatible - Gateway - GPU properties: - type: Documentation url: https://docs.cumuluslabs.io/ - type: GettingStarted url: https://docs.cumuluslabs.io/getting-started/ - type: APIReference url: https://docs.cumuluslabs.io/inference/overview/ - type: Authentication url: authentication/cumulus-labs-authentication.yml maintainers: - FN: Kin Lane email: kin@apievangelist.com - FN: APIs.json email: info@apis.io common: - type: DomainSecurity url: security/cumulus-labs-domain-security.yml - type: Website url: https://cumuluslabs.io/ - type: DeveloperPortal url: https://docs.cumuluslabs.io/ - type: Documentation url: https://docs.cumuluslabs.io/ - type: GettingStarted url: https://docs.cumuluslabs.io/getting-started/ - type: APIReference url: https://docs.cumuluslabs.io/inference/overview/ - type: Blog url: https://cumulus.blog/ - type: GitHubOrganization url: https://github.com/cumulus-compute-labs - type: Support url: mailto:founders@cumuluslabs.io - type: TermsOfService url: https://cumuluslabs.io/terms-of-service.pdf - type: PrivacyPolicy url: https://cumuluslabs.io/privacy-policy.pdf - type: LinkedIn url: https://www.linkedin.com/company/cumuluscomputelabs/ - type: Twitter url: https://x.com/cumuluslabsio - type: Authentication url: authentication/cumulus-labs-authentication.yml - type: Conventions url: conventions/cumulus-labs-conventions.yml - type: Conformance url: conformance/cumulus-labs-conformance.yml - type: Lifecycle url: lifecycle/cumulus-labs-lifecycle.yml - type: LLMsTxt url: llms/cumulus-labs-llms.txt x-enrichment: date: '2026-07-19' status: backfilled pass: local-v1 note: backfilled from .gitignore signal + verified work evidence