You send the model. We send back an endpoint.
Renting the box is the easy half. Getting throughput out of it, and keeping it up, is the half that eats your week. On the managed tier that's our job.
Four steps, then it's running.
You tell us the job
A model you want served, or just a description of what it has to do. If you don't have a model picked, that's fine — choosing one is part of this.
We fit it to the hardware
Quantisation, context length, batch size and engine choice tuned against the actual card, then benchmarked so you get real tokens-per-second numbers rather than a guess.
It goes live and stays warm
An OpenAI-compatible endpoint, kept loaded so the first request of the day isn't the slow one. Driver and framework upgrades happen on our side.
We watch it, not you
Throughput, latency and error rates monitored. If something drifts you hear from us — ideally before you'd have noticed.
When this is worth paying for.
Good fit
- You want an endpoint, not a server to administer
- Your model has to be up during business hours, every day
- Nobody on your team wants to become a vLLM expert
- You need the data to stay on hardware you can point at
Don't buy this if
- You're happy running your own stack — take the bare-metal plan instead and save the fee
- The workload is a few hours a month — a per-second API will be cheaper
- You need a frontier closed model — we serve open weights, so use the vendor's API
The things people ask first.
Why is this cheaper than AWS or Google Cloud?
What actually happens during a power outage?
Is the GPU shared with anyone else?
Which card do I need for my model?
Do I get root access? Can I run Docker?
You rent AMD boxes. Does my stack actually run on them?
How fast is deployment?
Do you offer hourly billing?
What's the contract and how do I pay?
Who do I talk to when something breaks?
Tell us what you're running.
Model, context length, requests per second — that's enough for us to tell you which box you need, and whether the cheaper one would do.
No setup fee · Month to month · sales@gpurack.net