
GPU clusters & on-prem AI
Your own AI compute, specified, installed and handed over: sizing, procurement, power, cooling, scheduling and support.
Who it's for
You are probably here because
-
01
Research data cannot legally or contractually leave the country or the building.
-
02
Cloud GPU hours have become the largest line in your budget.
-
03
Jobs queue for days and researchers have stopped bothering.
-
04
You have been quoted for hardware and want someone independent to check the spec.
What's included
What you actually get.
Very few firms in this market will do the whole chain, from working out how much VRAM your workload actually needs to racking the nodes and training your admin to run them. We do, and we write the code that runs on top.
Workload sizing
What you intend to run, whether inference, fine-tuning, rendering or simulation, converted into VRAM, throughput and node count.
Specification and procurement
GPU selection, server chassis, CPU, RAM, NVMe and interconnect. Sourcing, import handling and warranty terms.
Site readiness
Power draw and phase, UPS and generator sizing, cooling load, rack space, floor loading and physical security.
Build and install
Racking, cabling, driver and CUDA stack, cluster networking, storage tier.
Orchestration
Kubernetes or Slurm scheduling, container images, multi-tenant quotas, job queues.
Model serving
vLLM or Triton endpoints, a model registry, an internal API gateway and authentication.
Monitoring and handover
Utilisation and thermal dashboards, alerting, a runbook, admin training and a support contract.
Published shapes
Three shapes, not three prices.
We publish what each tier physically is and who it suits. The number comes after we understand the workload.
| Tier | Shape | Suits | |
|---|---|---|---|
|
Workstation Fastest to stand up |
1 to 2 GPUs, single tower or 4U node | A team piloting AI, one researcher, or render work | Brief this tier |
|
Departmental The common shape |
2 to 4 nodes, 8 to 16 GPUs, shared storage, scheduler | A university lab, a data team, private model serving | Brief this tier |
|
Production Site survey required |
Multi-rack, high-speed interconnect, redundant power | Company-wide inference, regulated data, continuous training | Brief this tier |
How it runs
Each stage, and what it produces.
-
01
Sizing
We start from the workload, not from a vendor's catalogue.
Output
A sizing note: VRAM, throughput, node count, power draw in kW.
-
02
Specification
Full bill of materials with alternatives and lead times.
Output
A quote you could take to another supplier.
-
03
Site readiness
Power, cooling, rack space, floor loading and security.
Output
A site readiness report and a list of what must change first.
-
04
Install
Racking, cabling, drivers, CUDA stack, cluster networking, storage.
Output
A cluster running a benchmark you watched.
-
05
Orchestration and serving
Slurm or Kubernetes, quotas, queues, model endpoints.
Output
Users submitting jobs and getting results.
-
06
Handover
Dashboards, alerting, runbook, admin training.
Output
Your team running it without us in the room.
What you are handed
- A rack elevation and an as-built cabling diagram
- A written runbook covering restart, failure and expansion
- Utilisation and thermal dashboards with alerting
- Administrator training, and a named engineer on support
Argument panel (dark)
Why on-premise
Data that cannot legally or contractually leave the country or the building.
Predictable cost once utilisation is steady, against per-token or per-hour cloud billing that scales with your success.
No dependence on international bandwidth for inference.
Hardware you own, on your balance sheet, that keeps working when a contract lapses.
Worth saying
When cloud is the right answer
Your load is spiky and rare. You are paying for idle silicon the rest of the month.
You do not yet know what you will run. Sizing hardware for an unknown workload is how expensive mistakes get made.
You have nowhere to put it. Power and cooling are real constraints and we will not pretend otherwise.
Technology
Proof
One project, in full.

On-premise AI cluster for a research institution
Research data that could not leave the country, on cloud GPU hours nobody could afford.
Overnight
Job turnaround
Research data could not be sent to overseas cloud providers, and cloud GPU hours were unaffordable at the volume required. Jobs queued for days, and researchers had begun to shrink their experiments to fit …
Overnight
Job turnaround
Eight
GPUs on site
On premise
Data residency
Commercials
How buying this works.
Engagement
Supply and install: hardware at quoted cost plus a stated installation fee.
Price
On request, after scope
What sets the price
Equipment, site conditions and lead time. Sizing comes first and everything else follows from it, which is why we will not quote a cluster over the phone.
We publish the shapes, not the prices, because the same node count costs very differently depending on interconnect, power and what your building can already carry.
Every project is quoted after we understand the scope. You will have a written, fixed quote before any build work begins. No open-ended billing.
Questions
Answered plainly.
That is the first engagement and it is deliberately separate. We size from the workload you intend to run and give you the number in writing, whether or not you then buy through us.
Part of the site readiness stage. We calculate draw and thermal load, size UPS and cooling, and tell you plainly if the building cannot take it.
Yes. Sourcing, import handling and warranty terms are part of supply and install, and lead times are stated in the quote.
That is the goal. Handover includes a runbook and administrator training, and the support contract is optional rather than assumed.
Yes, that is a different service. See GPU optimisation.
Often bought alongside
Infrastructure
Get your workload sized
Tell us what you intend to run. We will come back with VRAM, node count and power draw before anyone mentions a price.
We reply to every enquiry within one working day.