Crusoe Cloud Launches Managed AI Services for Scalable Inference and Reliable Orchestration

Crusoe, the first vertically integrated AI infrastructure provider, has announced two innovative managed services on its Crusoe Cloud platform—Crusoe Managed Inference and Crusoe AutoClusters. These offerings, powered by NVIDIA technology, were previewed at the NVIDIA GTC AI Conference.

Crusoe Managed Inference

Crusoe Managed Inference simplifies AI model deployment by allowing enterprise developers to run and automatically scale machine learning models without managing complex infrastructure. Users can send prompts directly to a Crusoe API and receive responses from an advanced AI model of their choice.

Key features and benefits include:

  • Rapid development and optimization: Accelerate AI solution creation without infrastructure overhead.
  • Agentic AI workflows: Easily embed AI responses into automated systems and intelligent applications.
  • User-friendly interface: Test and interact with models through an intuitive chat UI.

“Crusoe Managed Inference enables developers to focus on building intelligent applications instead of managing servers. I like to think of it as intelligence as a service,” said Nadav Eiron, SVP of cloud engineering at Crusoe.

Crusoe AutoClusters

Crusoe AutoClusters is a fault-tolerant orchestration platform for AI training. It simplifies deployment, orchestration, and maintenance of AI services, allowing users to concentrate on innovation. This service supports Slurm, Kubernetes, and other orchestration platforms, optimizing administration of high-performance computing environments.

Key features and benefits include:

  • Effortless provisioning: Deploy GPU clusters via API, CLI, or UI, leveraging NVIDIA Quantum-2 InfiniBand and VAST Data-backed file systems.
  • Proactive monitoring: Monitor performance using NVIDIA DCGM and Crusoe’s proprietary tools.
  • Automated node replacement: Detect errors and auto-replace failed nodes with minimal disruption.
  • Intelligent orchestration: Manage Slurm clusters with topology-aware job scheduling and automated job re-queueing.

“We’re eliminating the operational burdens that often hinder AI innovation,” added Eiron. “Crusoe AutoClusters ensure AI training workloads recover smoothly from hardware issues, delivering a reliable experience.”

“It’s pretty incredible that we were able to rapidly spin up 1600 GPUs, submit a job via Slurm on Crusoe Cloud and it just worked,” said Less Wright, PyTorch Partner Engineer at Meta.

Availability

Explore Crusoe’s latest services at GTC at booth #1633. Interested users can contact Crusoe to join the Q2 private preview programs.