GRGPURackAI infrastructureGet a box
Dedicated NVIDIA & AMD GPUs · Flat monthly rate

Rent a GPU server that stays up when the grid doesn't.

For LLM inference, fine-tuning and rendering — behind two independent fiber carriers, a solar array and a battery bank. Full root, nothing shared, one flat monthly rate.

No setup fee · Month to month · The person who built it answers the phone.

Status unavailable
Florida, USA
Solar
Battery
Grid
Rack
Fiber A — Carrier 1
Fiber B — Carrier 2

The monitoring agent hasn't reported recently. We'd rather show nothing than a stale reading.

2
Independent fiber carriers
Solar
Backed by a battery bank
192 GB
VRAM in a single box
30 min
Typical handover
Why we don't go down

Everyone promises uptime. We built for it.

An SLA is a refund policy — it pays you back for the hours your model was offline. We'd rather it never went offline.

Two carriers, not two cables

Most hosts advertising redundant network mean two drops from one provider — when that provider has a bad day, both drop together. We run two genuinely independent fiber carriers. One can fail completely and the rack doesn't notice.

Solar into the load

A solar array feeds the facility and keeps the battery bank topped up. It isn't a green talking point — it's the reason a utility interruption is a non-event instead of a scramble to start a generator.

Battery bank underneath it all

The batteries carry the load through a cut and the panels put the charge back. No transfer switch waiting on a diesel to spin up, no thirty-second gap where your inference queue dies.

The builder answers the phone

No ticket tiers, no offshore first line reading a script. You get the person who racked the hardware and can physically get hands on it. That's the real advantage of being small, and we'd rather stay small than lose it.

Plans

One flat rate. The whole machine.

No setup fee, no egress billing, no hourly meter. Month to month, cancel with 30 days' notice.

Starter
ROCm

Ryzen AI Max+ 395

128 GB unified LPDDR5X

128 GB shared between CPU and GPU — a 70B at 8-bit, or a 120B-class model at 4-bit, in memory that would cost 5× this in discrete cards.

Stream processors2,560 (40 CU)
FP3229.7 TFLOPS peak
Mem bandwidth256 GB/s
CPU16C / 32T Zen 5
RAM128 GB LPDDR5X (shared CPU+GPU)
Storage2 TB NVMe
Network1 Gbps unmetered

One 128 GB pool shared by CPU and GPU — that's the whole point of this box. The catch is bandwidth: 256 GB/s means tokens generate slower than on a discrete card, so it suits large models, batch work and dev rather than high-throughput serving. Runs on ROCm, not CUDA.

The most addressable memory per dollar we rent
$249/month
Reserve this box
2 available now
Standard
ROCm

AMD Radeon AI PRO R9700

64 GB total GDDR6

A 70B at 4-bit across both cards, or a 32B at fp8 with plenty of context headroom.

Stream processors4,096 per card
FP3247.8 TFLOPS per card
Mem bandwidth640 GB/s per card
CPU24 vCPU
RAM128 GB DDR5
Storage2 TB NVMe
Network1 Gbps unmetered

Runs on ROCm, not CUDA. PyTorch, vLLM, llama.cpp and Ollama all work — but check your stack before you commit, and ask us if you're unsure.

The cheapest route into 64 GB of VRAM
$399/month
Reserve this box
1 available now
Most popular
Standard
CUDA

RTX 5090

32 GB GDDR7

A 32B at fp8 with real context, or SDXL / Flux / ComfyUI without queueing.

CUDA cores21,760
FP32105 TFLOPS
Mem bandwidth1.79 TB/s
CPU24 vCPU
RAM128 GB DDR5
Storage2 TB NVMe
Network1 Gbps unmetered
Production inference, image and video generation
$449/month
Reserve this box
3 available now
Multi-GPU
CUDA

RTX 5090

64 GB total GDDR7

A 70B at 4-bit across both cards, or two independent 32B endpoints.

CUDA cores21,760 per card
FP32105 TFLOPS per card
Mem bandwidth1.79 TB/s per card
CPU32 vCPU
RAM192 GB DDR5
Storage2 TB NVMe
Network1 Gbps unmetered
Tensor-parallel serving, parallel render pipelines
$849/month
Reserve this box
1 available now
Pro
CUDA

RTX PRO 6000 Blackwell

96 GB GDDR7

A 70B at fp8 with long context on ONE card — no tensor-parallel setup to debug.

CUDA cores24,064
FP32126 TFLOPS
Mem bandwidth1.8 TB/s
CPU32 vCPU
RAM256 GB DDR5
Storage4 TB NVMe
Network1 Gbps unmetered
Large-model serving on a single card, LoRA fine-tuning
$949/month
Reserve this box
1 available now
Multi-GPU
CUDA

RTX PRO 6000 Blackwell

192 GB total GDDR7

A 120B-class open model at fp8, or a full fine-tune of a 35B.

CUDA cores24,064 per card
FP32126 TFLOPS per card
Mem bandwidth1.8 TB/s per card
CPU48 vCPU
RAM512 GB DDR5
Storage8 TB NVMe
Network1 Gbps unmetered
Frontier open-weight models, multi-tenant inference
$1,799/month
Reserve this box

Every plan: full root on bare metal · dedicated card, nothing shared · 1 dedicated IPv4 · Ubuntu, Windows or your own ISO

How it works

Three steps to a running box.

No console to learn, no quota requests, no sales cycle.

Step 1

Tell us the workload

Say what you're running — model, context length, how many requests a second. We'll tell you which box fits, including when the cheaper one is enough.

Step 2

We rack and provision it

Ubuntu 24.04 with the right driver stack already in place — CUDA on the NVIDIA boxes, ROCm on the AMD ones — plus the container toolkit. In-stock configurations are typically handed over the same day.

Step 3

SSH in and it's yours

Root on bare metal, full VRAM, one flat monthly rate. No hourly meter, no egress bill, nothing shared with anybody else.

What people run on them

Built for workloads that don't stop.

Always-on

LLM inference & serving

vLLM, Ollama, llama.cpp or TGI on a card that's yours around the clock. No cold starts, no rate limits, no queue behind someone else's job.

vLLMOllamallama.cppTGISGLang
Full VRAM

Fine-tuning & LoRA

Train against the whole card instead of a sliced allocation, and leave a job running for a week without a spot instance evaporating underneath it.

PyTorchAxolotlUnslothPEFTTRL
High VRAM

Image & video generation

SDXL, Flux, ComfyUI and video models with the VRAM headroom they actually want, on hardware that doesn't bill you by the second while it loads.

ComfyUISDXLFluxAUTOMATIC1111
No queues

Rendering & batch

Blender, V-Ray or Redshift on a dedicated card, plus any long-running batch job that just needs a machine nobody else is going to touch.

BlenderV-RayRedshiftFFmpeg
How we compare

Where we win, and where we don't.

Including the row where we lose. If per-second billing is what you need, we'll tell you to go elsewhere rather than sell you the wrong thing.

 GPURackHyperscalerTypical GPU host
Dedicated card, nothing shared
Full root on bare metal
Flat monthly rate
No egress / bandwidth billing
Two independent fiber carriers
Solar + battery behind the load
Talk to the person who built it
Per-second billing
We're honest: if your workload is genuinely bursty, spot pricing beats us.

means it depends on the provider and the plan.

Managed AI

Or don't touch the server at all.

Renting the box is the easy half. The hard half is picking a quantisation that fits, getting throughput out of vLLM, and noticing at 3am when the endpoint stops answering.

On the managed tier we do that part. You send us the model — or tell us what the thing needs to do — and you get back an endpoint and a number to call. The hardware underneath is the same dedicated card, on the same power.

Managed add-on
from $400/month

On top of any plan.

  • We pick the model and quantisation for your box
  • vLLM or Ollama deployed, tuned and benchmarked
  • An OpenAI-compatible endpoint, ready for your code
  • Keep-warm so the first request isn't the slow one
  • Driver, CUDA and framework upgrades handled
  • We watch it — you hear from us before you notice
Questions

The things people ask first.

Why is this cheaper than AWS or Google Cloud?
Because you're renting the actual machine, not a slice of one with a hyperscaler's margin stacked on top. There's no egress billing, no per-hour meter, no support tier to upgrade into. One flat monthly number for the whole box, and the GPU is yours alone for the month.
What actually happens during a power outage?
Nothing you'd notice. The facility runs behind a battery bank with a solar array feeding it, so a utility cut is a transfer, not an interruption — the batteries carry the load and the panels recharge them. Network is the same story: two independent fiber carriers, not two drops from one provider, so a single carrier's outage doesn't take the rack with it. This is the reason the business exists.
Is the GPU shared with anyone else?
No. Every plan is a dedicated card in a dedicated box. You get the full VRAM, the full clock, and root. Nothing is virtualised, oversold, or time-sliced, so throughput doesn't move around based on who else is on the machine.
Which card do I need for my model?
Rough rule for inference: VRAM in GB should be about double the parameter count in billions at fp16, or roughly equal to it at fp8. A 32B fits an RTX 5090 at fp8; a 70B wants the 96 GB RTX PRO 6000 to stay on one card. Tell us the model and the context length you're targeting and we'll tell you honestly which box to take — including when the cheaper one is enough.
Do I get root access? Can I run Docker?
Full root on bare metal, and yes to Docker. Ubuntu 24.04 by default with the driver stack already in place — CUDA on the NVIDIA boxes, ROCm on the AMD ones — plus the container toolkit. vLLM, Ollama, llama.cpp, ComfyUI, PyTorch and TensorFlow all run as they do on any other machine you own. Bring your own ISO if you'd rather.
You rent AMD boxes. Does my stack actually run on them?
Usually, but check first — this is the one thing worth five minutes before you order. PyTorch, vLLM, llama.cpp and Ollama all have working ROCm support, and for straightforward LLM inference the AMD boxes are the cheapest VRAM we rent. Where it gets thin is custom CUDA kernels, some quantisation libraries, and anything depending on a niche NVIDIA-only package. Tell us what you're running and we'll give you a straight answer rather than a maybe — and if the answer is that you need CUDA, we'll point you at the NVIDIA plans instead.
How fast is deployment?
In-stock configurations are typically handed over the same day — you get an IP, root credentials, and a working driver stack. Custom builds depend on parts and we'll give you a real date rather than a hopeful one.
Do you offer hourly billing?
Not today. Hourly billing means keeping cards idle between customers, and the cost of that idle time gets priced back into every hour you do use. Flat monthly is how the rate stays where it is. If your workload is genuinely bursty, a hyperscaler's spot market will beat us and we'll say so.
What's the contract and how do I pay?
Month to month. No setup fee, no minimum term, cancel with 30 days' notice. Invoiced monthly by card or ACH.
Who do I talk to when something breaks?
The person who built the rack. There's no ticket tier and no offshore first line — you get a direct line to someone who can actually get hands on the hardware. That's a genuine advantage of being small, and it stops being true if we ever get big enough to need a call centre.

Tell us what you're running.

Model, context length, requests per second — that's enough for us to tell you which box you need, and whether the cheaper one would do.

No setup fee · Month to month · sales@gpurack.net