GPU clusters & on-prem AI

Your own AI compute, specified, installed and handed over: sizing, procurement, power, cooling, scheduling and support.

01

Who it's for

You are probably here because

  • 01

    Research data cannot legally or contractually leave the country or the building.

  • 02

    Cloud GPU hours have become the largest line in your budget.

  • 03

    Jobs queue for days and researchers have stopped bothering.

  • 04

    You have been quoted for hardware and want someone independent to check the spec.

02

What's included

What you actually get.

Very few firms in this market will do the whole chain, from working out how much VRAM your workload actually needs to racking the nodes and training your admin to run them. We do, and we write the code that runs on top.

01

Workload sizing

What you intend to run, whether inference, fine-tuning, rendering or simulation, converted into VRAM, throughput and node count.

02

Specification and procurement

GPU selection, server chassis, CPU, RAM, NVMe and interconnect. Sourcing, import handling and warranty terms.

03

Site readiness

Power draw and phase, UPS and generator sizing, cooling load, rack space, floor loading and physical security.

04

Build and install

Racking, cabling, driver and CUDA stack, cluster networking, storage tier.

05

Orchestration

Kubernetes or Slurm scheduling, container images, multi-tenant quotas, job queues.

06

Model serving

vLLM or Triton endpoints, a model registry, an internal API gateway and authentication.

07

Monitoring and handover

Utilisation and thermal dashboards, alerting, a runbook, admin training and a support contract.

03

Published shapes

Three shapes, not three prices.

We publish what each tier physically is and who it suits. The number comes after we understand the workload.

Tier Shape Suits

Workstation

Fastest to stand up

1 to 2 GPUs, single tower or 4U node A team piloting AI, one researcher, or render work Brief this tier

Departmental

The common shape

2 to 4 nodes, 8 to 16 GPUs, shared storage, scheduler A university lab, a data team, private model serving Brief this tier

Production

Site survey required

Multi-rack, high-speed interconnect, redundant power Company-wide inference, regulated data, continuous training Brief this tier
04

How it runs

Each stage, and what it produces.

  1. 01

    Sizing

    We start from the workload, not from a vendor's catalogue.

    Output

    A sizing note: VRAM, throughput, node count, power draw in kW.

  2. 02

    Specification

    Full bill of materials with alternatives and lead times.

    Output

    A quote you could take to another supplier.

  3. 03

    Site readiness

    Power, cooling, rack space, floor loading and security.

    Output

    A site readiness report and a list of what must change first.

  4. 04

    Install

    Racking, cabling, drivers, CUDA stack, cluster networking, storage.

    Output

    A cluster running a benchmark you watched.

  5. 05

    Orchestration and serving

    Slurm or Kubernetes, quotas, queues, model endpoints.

    Output

    Users submitting jobs and getting results.

  6. 06

    Handover

    Dashboards, alerting, runbook, admin training.

    Output

    Your team running it without us in the room.

What you are handed

  • A rack elevation and an as-built cabling diagram
  • A written runbook covering restart, failure and expansion
  • Utilisation and thermal dashboards with alerting
  • Administrator training, and a named engineer on support

Argument panel (dark)

Why on-premise

Data that cannot legally or contractually leave the country or the building.

Predictable cost once utilisation is steady, against per-token or per-hour cloud billing that scales with your success.

No dependence on international bandwidth for inference.

Hardware you own, on your balance sheet, that keeps working when a contract lapses.

Worth saying

When cloud is the right answer

Your load is spiky and rare. You are paying for idle silicon the rest of the month.

You do not yet know what you will run. Sizing hardware for an unknown workload is how expensive mistakes get made.

You have nowhere to put it. Power and cooling are real constraints and we will not pretend otherwise.

Technology

NVIDIACUDASlurmKubernetesvLLMTritonUbuntuNVLinkInfiniBandPrometheusGrafana
05

Proof

One project, in full.

On-premise AI cluster for a research institution Schools and institutions Example

On-premise AI cluster for a research institution

Research data that could not leave the country, on cloud GPU hours nobody could afford.

Overnight

Job turnaround

SlurmvLLMUbuntuNVIDIA

Research data could not be sent to overseas cloud providers, and cloud GPU hours were unaffordable at the volume required. Jobs queued for days, and researchers had begun to shrink their experiments to fit …

Overnight

Job turnaround

Eight

GPUs on site

On premise

Data residency

Read the full case
06

Commercials

How buying this works.

Engagement
Supply and install: hardware at quoted cost plus a stated installation fee.

Price
On request, after scope

How pricing works

What sets the price

Equipment, site conditions and lead time. Sizing comes first and everything else follows from it, which is why we will not quote a cluster over the phone.

We publish the shapes, not the prices, because the same node count costs very differently depending on interconnect, power and what your building can already carry.

Every project is quoted after we understand the scope. You will have a written, fixed quote before any build work begins. No open-ended billing.

07

Questions

Answered plainly.

That is the first engagement and it is deliberately separate. We size from the workload you intend to run and give you the number in writing, whether or not you then buy through us.

Part of the site readiness stage. We calculate draw and thermal load, size UPS and cooling, and tell you plainly if the building cannot take it.

Yes. Sourcing, import handling and warranty terms are part of supply and install, and lead times are stated in the quote.

That is the goal. Handover includes a runbook and administrator training, and the support contract is optional rather than assumed.

Yes, that is a different service. See GPU optimisation.

Infrastructure

Get your workload sized

Tell us what you intend to run. We will come back with VRAM, node count and power draw before anyone mentions a price.

We reply to every enquiry within one working day.

WhatsApp us