Rent a GPU server that stays up when the grid doesn't.
For LLM inference, fine-tuning and rendering — behind two independent fiber carriers, a solar array and a battery bank. Full root, nothing shared, one flat monthly rate.
No setup fee · Month to month · The person who built it answers the phone.
The monitoring agent hasn't reported recently. We'd rather show nothing than a stale reading.
Everyone promises uptime. We built for it.
An SLA is a refund policy — it pays you back for the hours your model was offline. We'd rather it never went offline.
Two carriers, not two cables
Most hosts advertising redundant network mean two drops from one provider — when that provider has a bad day, both drop together. We run two genuinely independent fiber carriers. One can fail completely and the rack doesn't notice.
Solar into the load
A solar array feeds the facility and keeps the battery bank topped up. It isn't a green talking point — it's the reason a utility interruption is a non-event instead of a scramble to start a generator.
Battery bank underneath it all
The batteries carry the load through a cut and the panels put the charge back. No transfer switch waiting on a diesel to spin up, no thirty-second gap where your inference queue dies.
The builder answers the phone
No ticket tiers, no offshore first line reading a script. You get the person who racked the hardware and can physically get hands on it. That's the real advantage of being small, and we'd rather stay small than lose it.
One flat rate. The whole machine.
No setup fee, no egress billing, no hourly meter. Month to month, cancel with 30 days' notice.
Ryzen AI Max+ 395
128 GB shared between CPU and GPU — a 70B at 8-bit, or a 120B-class model at 4-bit, in memory that would cost 5× this in discrete cards.
One 128 GB pool shared by CPU and GPU — that's the whole point of this box. The catch is bandwidth: 256 GB/s means tokens generate slower than on a discrete card, so it suits large models, batch work and dev rather than high-throughput serving. Runs on ROCm, not CUDA.
2× AMD Radeon AI PRO R9700
A 70B at 4-bit across both cards, or a 32B at fp8 with plenty of context headroom.
Runs on ROCm, not CUDA. PyTorch, vLLM, llama.cpp and Ollama all work — but check your stack before you commit, and ask us if you're unsure.
RTX 5090
A 32B at fp8 with real context, or SDXL / Flux / ComfyUI without queueing.
2× RTX 5090
A 70B at 4-bit across both cards, or two independent 32B endpoints.
RTX PRO 6000 Blackwell
A 70B at fp8 with long context on ONE card — no tensor-parallel setup to debug.
2× RTX PRO 6000 Blackwell
A 120B-class open model at fp8, or a full fine-tune of a 35B.
Every plan: full root on bare metal · dedicated card, nothing shared · 1 dedicated IPv4 · Ubuntu, Windows or your own ISO
Three steps to a running box.
No console to learn, no quota requests, no sales cycle.
Tell us the workload
Say what you're running — model, context length, how many requests a second. We'll tell you which box fits, including when the cheaper one is enough.
We rack and provision it
Ubuntu 24.04 with the right driver stack already in place — CUDA on the NVIDIA boxes, ROCm on the AMD ones — plus the container toolkit. In-stock configurations are typically handed over the same day.
SSH in and it's yours
Root on bare metal, full VRAM, one flat monthly rate. No hourly meter, no egress bill, nothing shared with anybody else.
Built for workloads that don't stop.
LLM inference & serving
vLLM, Ollama, llama.cpp or TGI on a card that's yours around the clock. No cold starts, no rate limits, no queue behind someone else's job.
Fine-tuning & LoRA
Train against the whole card instead of a sliced allocation, and leave a job running for a week without a spot instance evaporating underneath it.
Image & video generation
SDXL, Flux, ComfyUI and video models with the VRAM headroom they actually want, on hardware that doesn't bill you by the second while it loads.
Rendering & batch
Blender, V-Ray or Redshift on a dedicated card, plus any long-running batch job that just needs a machine nobody else is going to touch.
Where we win, and where we don't.
Including the row where we lose. If per-second billing is what you need, we'll tell you to go elsewhere rather than sell you the wrong thing.
| GPURack | Hyperscaler | Typical GPU host | |
|---|---|---|---|
| Dedicated card, nothing shared | |||
| Full root on bare metal | |||
| Flat monthly rate | |||
| No egress / bandwidth billing | |||
| Two independent fiber carriers | |||
| Solar + battery behind the load | |||
| Talk to the person who built it | |||
| Per-second billing We're honest: if your workload is genuinely bursty, spot pricing beats us. |
means it depends on the provider and the plan.
Or don't touch the server at all.
Renting the box is the easy half. The hard half is picking a quantisation that fits, getting throughput out of vLLM, and noticing at 3am when the endpoint stops answering.
On the managed tier we do that part. You send us the model — or tell us what the thing needs to do — and you get back an endpoint and a number to call. The hardware underneath is the same dedicated card, on the same power.
On top of any plan.
- We pick the model and quantisation for your box
- vLLM or Ollama deployed, tuned and benchmarked
- An OpenAI-compatible endpoint, ready for your code
- Keep-warm so the first request isn't the slow one
- Driver, CUDA and framework upgrades handled
- We watch it — you hear from us before you notice
The things people ask first.
Why is this cheaper than AWS or Google Cloud?
What actually happens during a power outage?
Is the GPU shared with anyone else?
Which card do I need for my model?
Do I get root access? Can I run Docker?
You rent AMD boxes. Does my stack actually run on them?
How fast is deployment?
Do you offer hourly billing?
What's the contract and how do I pay?
Who do I talk to when something breaks?
Tell us what you're running.
Model, context length, requests per second — that's enough for us to tell you which box you need, and whether the cheaper one would do.
No setup fee · Month to month · sales@gpurack.net