GRGPURackAI infrastructureGet a box
Managed AI

You send the model. We send back an endpoint.

Renting the box is the easy half. Getting throughput out of it, and keeping it up, is the half that eats your week. On the managed tier that's our job.

How it goes

Four steps, then it's running.

Step 1

You tell us the job

A model you want served, or just a description of what it has to do. If you don't have a model picked, that's fine — choosing one is part of this.

Step 2

We fit it to the hardware

Quantisation, context length, batch size and engine choice tuned against the actual card, then benchmarked so you get real tokens-per-second numbers rather than a guess.

Step 3

It goes live and stays warm

An OpenAI-compatible endpoint, kept loaded so the first request of the day isn't the slow one. Driver and framework upgrades happen on our side.

Step 4

We watch it, not you

Throughput, latency and error rates monitored. If something drifts you hear from us — ideally before you'd have noticed.

Fit

When this is worth paying for.

Good fit

  • You want an endpoint, not a server to administer
  • Your model has to be up during business hours, every day
  • Nobody on your team wants to become a vLLM expert
  • You need the data to stay on hardware you can point at

Don't buy this if

  • You're happy running your own stack — take the bare-metal plan instead and save the fee
  • The workload is a few hours a month — a per-second API will be cheaper
  • You need a frontier closed model — we serve open weights, so use the vendor's API
Questions

The things people ask first.

Why is this cheaper than AWS or Google Cloud?
Because you're renting the actual machine, not a slice of one with a hyperscaler's margin stacked on top. There's no egress billing, no per-hour meter, no support tier to upgrade into. One flat monthly number for the whole box, and the GPU is yours alone for the month.
What actually happens during a power outage?
Nothing you'd notice. The facility runs behind a battery bank with a solar array feeding it, so a utility cut is a transfer, not an interruption — the batteries carry the load and the panels recharge them. Network is the same story: two independent fiber carriers, not two drops from one provider, so a single carrier's outage doesn't take the rack with it. This is the reason the business exists.
Is the GPU shared with anyone else?
No. Every plan is a dedicated card in a dedicated box. You get the full VRAM, the full clock, and root. Nothing is virtualised, oversold, or time-sliced, so throughput doesn't move around based on who else is on the machine.
Which card do I need for my model?
Rough rule for inference: VRAM in GB should be about double the parameter count in billions at fp16, or roughly equal to it at fp8. A 32B fits an RTX 5090 at fp8; a 70B wants the 96 GB RTX PRO 6000 to stay on one card. Tell us the model and the context length you're targeting and we'll tell you honestly which box to take — including when the cheaper one is enough.
Do I get root access? Can I run Docker?
Full root on bare metal, and yes to Docker. Ubuntu 24.04 by default with the driver stack already in place — CUDA on the NVIDIA boxes, ROCm on the AMD ones — plus the container toolkit. vLLM, Ollama, llama.cpp, ComfyUI, PyTorch and TensorFlow all run as they do on any other machine you own. Bring your own ISO if you'd rather.
You rent AMD boxes. Does my stack actually run on them?
Usually, but check first — this is the one thing worth five minutes before you order. PyTorch, vLLM, llama.cpp and Ollama all have working ROCm support, and for straightforward LLM inference the AMD boxes are the cheapest VRAM we rent. Where it gets thin is custom CUDA kernels, some quantisation libraries, and anything depending on a niche NVIDIA-only package. Tell us what you're running and we'll give you a straight answer rather than a maybe — and if the answer is that you need CUDA, we'll point you at the NVIDIA plans instead.
How fast is deployment?
In-stock configurations are typically handed over the same day — you get an IP, root credentials, and a working driver stack. Custom builds depend on parts and we'll give you a real date rather than a hopeful one.
Do you offer hourly billing?
Not today. Hourly billing means keeping cards idle between customers, and the cost of that idle time gets priced back into every hour you do use. Flat monthly is how the rate stays where it is. If your workload is genuinely bursty, a hyperscaler's spot market will beat us and we'll say so.
What's the contract and how do I pay?
Month to month. No setup fee, no minimum term, cancel with 30 days' notice. Invoiced monthly by card or ACH.
Who do I talk to when something breaks?
The person who built the rack. There's no ticket tier and no offshore first line — you get a direct line to someone who can actually get hands on the hardware. That's a genuine advantage of being small, and it stops being true if we ever get big enough to need a call centre.

Tell us what you're running.

Model, context length, requests per second — that's enough for us to tell you which box you need, and whether the cheaper one would do.

No setup fee · Month to month · sales@gpurack.net