| AI |
- Inference runtime integration (vLLM/SGLang), domestic GPU support
- Model asset center MVP (user/project/repo management, model & dataset upload/download, CLI)
- Pre-integrated domestic model repos (Qwen/GLM/Baichuan)
- Inference acceleration: multi-level KV Cache, topology-aware scheduling (Kueue/Gang)
- Training-inference co-location basics
- AI fault diagnosis (multi-source log correlation + root cause analysis)
- Predictive alerting (time-series anomaly detection, resource exhaustion warnings)
|
- DCE AI Runtime GA
- Unified inference API (OpenAI API/Llama Stack compatible)
- Fine-tuning/LoRA support
- Multi-modal inference (text-image, audio-video)
- Model asset center enhancements (remote replication/sync, security scanning, pre-warming, i18n)
- MatrixHub CNCF submission1
- AI Agent infrastructure Beta (sandbox, memory & context, semantic routing)
- Fault self-healing (integrated training/inference framework auto-recovery)
- Alert noise reduction (automatic correlated alert grouping)
- LLM security (model access control, inference content safety policies)
|
- Distributed inference
- Training-inference co-location optimization
- Full-stack AI automation (AutoML + Agent)
|
| Infra |
- MetaX GPU onboarding (network topology, Lustre GDS)
- Ascend 910C NPU scheduling (CANN driver)
- Hygon DCU GPU scheduling
- AI high-performance storage (Lustre file system)
- Kueue/Gang Scheduling/LWS/DRA integration
- HAMi commercial edition integration2
- containerd enhancements (container disk limits)
|
- Domestic GPU full GA (MetaX/Ascend/Hygon/Biren)
- MetaX supernode release
- Supernode solution (8/16-card high-density, GPU sharing scheduler)
- GPU Operator hybrid scheduling (CPU + GPU + NPU), utilization → 80%+
- Distributed storage solution (cloud scenarios)
|
- DPU/NPU unified scheduling
- Computing network, multi-cluster compute federation
- InfiniBand topology discovery (via UFM)
|
| Plat |
- One-click install (Web UI + CLI, auto environment detection)
- Preflight check framework (plugin-based, network/storage/permission checks)
- Gateway API migration start (Ingress retirement)
- Log aggregation enhancements
- Compute cloud operations platform admin console
- Compute baseline review & billing model optimization
- Ghippo admin console UI
- CSP user two-factor authentication (2FA)
|
- Rolling upgrades (zero-downtime, canary + rollback)
- Gateway API migration complete
- Deployment time → 15 min (from ~2 hours)
- Compute cloud platform enhancements (tenant isolation, inventory management, billing conversion, GPU up/downgrade)
- Bare-metal deployer (cluster provisioning, automated testing, single-node troubleshooting)
|
- Lightweight kernel, edge-native
- Self-adaptive platform (auto-tuning + self-healing)
|
| Eco |
- Kueue/LWS/Gang Scheduling K8s AI/ML SIG contributions
- Spiderpool DRA implementation, DRANet
- Spiderpool MetaX GPU support
- GAIE/NIXL/LMCache inference optimization project participation
|
- MatrixHub Sandbox
- unifabric 1.0 (network health check, disaster marking, KV Cache sync monitoring)
- metal-deployer engineering delivery
- GAIE/NIXL community seats
|
- unifabric Sandbox, InfiniBand support
- Low-code orchestration, natural language operations
|