Skip to content

Blog

GPU as a Service: how to use GPUs in the cloud

How to use cloud GPUs with Hikube GPU as a Service: use cases, cost and a practical path. NVIDIA L40S, A100 and H100 are in the 14-day trial.

Hidora article published 23 February 2026. Figures, prices and comparisons are as of that date.

Introduction

Demand for compute is accelerating under the pressure of generative AI, machine learning and intensive processing workloads. GPU access, long confined to specialised environments or dedicated servers, has become a strategic question for companies. As models grow more complex and hardware cycles shorten, GPU as a Service (GPUaaS) offerings appear as a way to obtain fitted compute capacity immediately, with no upfront hardware investment.

With GPU availability fluctuating from market to market, and demand sometimes outstripping supply, organisations are looking for a flexible approach that lets them run high-performance accelerators while keeping costs and availability under control.

GPUs available on demand

GPUaaS services provision one or several GPUs instantly through an API, a portal or a Kubernetes cluster. Providers - sovereign clouds, hyperscalers or specialised operators - offer infrastructure able to run AI workloads, simulation or large-scale data analysis.

These offerings rest on recent accelerators such as NVIDIA A100, H100, L40S, or AMD Instinct and Gaudi/TPU alternatives depending on the environment. Access is usually possible in three ways: through GPU virtual machines, through Kubernetes pods (Device Plugin, GPU Operator), or by calling compute APIs directly.

The on-demand model, billed by usage, gives access to recent GPUs for a few hours or a few days without tying up capital. That format suits experimentation phases as well as production workloads.

An architecture designed for intensive workloads

Technically, GPUaaS services run on nodes fitted with modern GPUs, interconnected over PCIe Gen4/Gen5 or through NVLink. Distributed environments rely on high-throughput networks supporting RDMA or RoCEv2, essential for synchronising large models. Local NVMe volumes or low-latency distributed storage matter just as much, because training performance depends on data throughput as much as on raw GPU power.

AI operators such as Kubeflow Training Operator, DeepSpeed, Megatron-LM, Ray Serve or the MPI Operator orchestrate training, inference and task distribution across several GPUs. Providers support advanced features such as Multi-Instance GPU (MIG), or forms of GPU sharing through MPS or third-party extensions.

Use cases include LLM fine-tuning, video analysis, scientific computing, simulation and multimodal generation.

Optimising cost and performance

The GPUaaS approach stands out for its elasticity and usage-based pricing. Companies can allocate resources only during intensive training periods, while benefiting from:

  • hourly or à la carte billing,

  • automatic scaling to adjust GPU capacity,

  • the option to reserve GPUs over a longer term,

  • logical partitioning through MIG to share resources.

Performance varies with network quality, storage and how well the ML frameworks are tuned (PyTorch, TensorFlow, JAX). GPUaaS solutions offer an agile alternative to on-premises architectures, while avoiding the rapid depreciation that comes with successive GPU generations.

Comparison with the alternatives

Against dedicated GPU servers

  • GPUaaS advantage: flexibility, fast access to recent GPUs, no hardware to manage.

  • Limit, dependence on the provider, and costs that vary with usage.

Against hyperscalers

  • Sovereign or specialised GPUaaS can offer:

    • lower latency,

    • more predictable costs,

    • AI-oriented support,

    • hosting under known geographic control.

Against on-premises

  • Lower capex,

  • no hardware purchasing cycle,

  • dynamic allocation according to load,

  • the option to move between GPU generations as needs change.

What it means for companies, and what comes next

Organisations today are looking for environments that can run intensive AI workloads while guaranteeing sovereignty, flexibility and predictable performance. Cloud-native platforms such as Hikube, operated by Hidora SA on three Swiss datacenters (Geneva, Gland, Lucerne) with NVIDIA GPU nodes (L40S, A100, H100, included in the 14-day trial), offer a model suited to teams that want serious compute without excessive operational complexity. This kind of infrastructure supports distributed training, hosts production inference workloads, and answers location and compliance constraints.

The next developments in the market should include:

  • topology-aware scheduling to optimise multi-GPU placement,

  • broader support for heterogeneous accelerators (TPU, NPU, RDU),

  • automatic GPU scalability mechanisms,

  • improvements in GPU slicing and logical partitioning,

  • tighter integration with MLOps workflows and distributed training frameworks.

In a landscape where demand for compute keeps rising, GPUaaS solutions are a major lever for exploiting advances in AI while keeping costs under control and delivering the flexibility technical teams expect.

Ready to run on 100% Swiss infrastructure?

14-day trial, no credit card. GPUs included.