Your models. Your hardware. Your data.
We design, build, and deploy on-premise NVIDIA servers for training and inference — then stay on as the team that keeps them running, tuned, and current as the stack moves.
Process
How it works
Three steps.
Intake call
What you need to run, what's blocking you, what your constraints are. If self-hosting is the wrong answer for your situation, we'll say so on this call.
We design and build it
Or pick up where your stalled project left off. Workload analysis, benchmarking, hardware selection, procurement, build, tuning, installation. You get working hardware, not a parts list.
We keep it running
Support contract begins. You focus on using your hardware; we keep everything on it current — models included.
Why self-host
Which of these sounds familiar?
Do you want control over your own data?
If your workload touches medical records, legal documents, financial data, or anything under a confidentiality obligation, running it through someone else's inference endpoint is a conversation you'd rather not have — and one your customers may eventually start asking you about.
Self-hosting removes the question entirely. Your data never leaves your premises: no shared tenancy, no third-party endpoint, no ambiguity in your DPA.
Is regulation keeping you out of the cloud?
For some organisations this isn't a preference. Sector rules, procurement requirements, or a data processing agreement you've already signed can rule out sending data to a third-party provider — no matter how good their security posture is. And with most of the widely used inference APIs operated by non-EU companies, every request is a transfer question you have to be able to answer.
If that's your situation, on-premise isn't the cautious option, it's the only one. Hardware in your own building, in your own jurisdiction, makes the answer trivial — we build for that constraint rather than working around it, and we'll document the setup so your compliance officer has something concrete to sign off.
Are your token costs stacking up?
Per-token pricing is excellent while you're experimenting and less comfortable once a workload runs continuously and predictably. Some teams would simply rather make one capital investment than watch a meter every month.
Whether owning works out cheaper for you depends entirely on how steady your usage is — so we'll model it against your actual numbers and tell you honestly if the answer is no. We're not going to quote you a payback period on a website.
Is your provider changing models underneath you?
Hosted models get deprecated, retired, and quietly updated. Your prompts were tuned against behaviour that may no longer exist, your evals shift for reasons you can't see, and the migration timeline is somebody else's decision.
An open-weights model on your own hardware changes when you decide it changes — and not a day sooner. When something genuinely better arrives, we'll measure it and tell you.
Do you know whether you're getting the most out of the hardware you already have?
Most self-hosted deployments run on defaults. The BIOS settings are whatever shipped, the serving runtime is whatever got installed first, and the configuration was tuned once — if at all — against a workload that has since changed. It still works, which is exactly why nobody goes back to look at it.
The distance between a default configuration and a properly tuned one on identical hardware is often substantial. We find it with isobench, our own benchmarking framework: we measure across BIOS and firmware settings, driver versions, serving stacks, and model configurations, then show you what your hardware can actually do next to what it's doing now.
Sometimes the conclusion is that you need another node. More often it's that you don't.
What self-hosting hands you in exchange is a hardware selection problem, a power and cooling problem, a firmware and driver problem, and a "who fixes this when it breaks" problem.
That's the part we do.
Deployment
Getting it running
You tell us what you need to train or serve. We hand back working hardware with your workload already on it, performing the way it should.
| 1. Identify the workload | Model sizes, batch shapes, throughput and latency targets, concurrency, context lengths, and how any of that is expected to grow. Before any hardware is discussed. |
| 2. Design the system | The GPU server is one component of several. Inference gateway and request routing, authentication and rate limiting, storage tiers for weights, datasets and checkpoints, network topology and bandwidth between nodes, observability, backup and failover — and how all of it meets the environment you already run. |
| 3. Select the hardware | GPU choice, host platform, CPU, RAM, NVMe tiering, interconnect, power and cooling envelope — sized to the workload and the design above, and backed by measured benchmarks rather than datasheets. We take no margin on hardware, so the recommendation is the one that fits, not the one that pays best. |
| 4. Procure and build | Nothing is ordered until you've approved the design and the bill of materials. Then sourcing, lead-time management, assembly, firmware, thermal validation, and sustained load testing before anything sees production traffic. |
| ⋯ | Stack tuning, OS and driver configuration, serving runtime selection and flags, orchestration, deployment, integration, benchmark baselining, runbooks, handover — and a fair amount more. The real process is considerably longer than a table on a website; we'll walk you through the parts that matter for your situation. |
Typical builds
Serving a large open-weights model at real concurrency takes more hardware than most teams expect.
8-GPU datacenter node (HGX-class, H200 / B200)
The largest open-weights models at real concurrency, and serious single-node training. This is what it takes to serve them properly.
Multi-node cluster
Several nodes plus a high-speed fabric and shared storage, for distributed training or an inference fleet that can't go down.
Single-node server (1–4 GPUs)
Smaller models, fine-tuning, on-prem RAG, or a pilot before committing to a full node.
Three shapes, not a menu. If your workload doesn't look like any of them, that's exactly what the intake call is for — and sometimes the honest answer is smaller than all three.
Timeline. We work to your timeline. GPU lead times are the main variable and they're outside anyone's control, but we can often streamline sourcing — and you'll get a realistic estimate on the intake call rather than an optimistic one on a website.
Stalled projects
Already bought the hardware and the project stalled?
It's a common situation. The servers arrived, the person driving it moved on or moved teams, and now there's expensive hardware sitting at a fraction of what it can do — or not in production at all. Nobody's quite sure what was decided or why.
We pick these up. We work out what you have, what was configured and what wasn't, where the performance is going, and what the fastest route to production looks like from here. You get a clear picture of where the project actually stands and a plan to finish it.
No judgement about how it got there. This is most of what goes wrong in self-hosted AI infrastructure, and it's rarely anyone's fault — the stack is genuinely hard and it moves faster than most teams can track.
Support contract
Keeping it current
Hardware is the easy half. New open-weights models and serving software ship constantly, and a server nobody maintains quietly becomes a liability.
The configuration that was optimal at launch won't be optimal in three months. A new model lands that beats what you're running at the same cost. SGLang ships a feature that materially speeds up your specific workload. A driver release changes throughput.
Almost nobody has time to track this, benchmark it, and decide whether it's worth the migration risk. We do it as a matter of course — and we tell you when it's worth moving, not every time something is released.
One contract. Not tiers, not an advisory upsell, no per-incident invoicing — because the most valuable thing we do is tell you when something better exists, and that shouldn't sit behind a paywall.
- Monitoring and alerting GPU utilisation, memory, thermals, ECC errors, throughput, and drift from your benchmark baseline. We watch it so your engineers don't.
- Patching and upgrades Firmware, drivers, CUDA, kernel, container runtime, serving stack. Tested before it touches production.
- Recalibration and tuning As your traffic shape changes, the configuration that was right at launch stops being right. We adjust it.
- Runtime and model advisory When a newer model or a new runtime release measurably beats your current setup, you hear it from us with numbers and a recommendation on timing. Not a newsletter.
- Migration support When it is worth moving, we do the move.
- On-call incident response Included, not an add-on. Response terms agreed in your contract.
- Capacity planning We tell you when you're months from needing another node, before it's urgent.
- Hardware lifecycle Warranty and RMA handling, a spares strategy sized to your tolerance for downtime, and an honest answer on when to refresh versus hold.
We run this stack across multiple customers. When a new model, runtime, or driver actually earns its upgrade, we've usually already measured it somewhere else. You get the benefit of those experiments without paying for them.
We don't guess. We measure.
Performance is decided by a long list of choices that interact with each other — and the list starts at procurement, not at installation: GPU and host platform, interconnect, air or liquid cooling, BIOS and firmware settings, OS configuration, driver and CUDA versions, serving runtime and its flags, quantisation, batching strategy, and the model itself. Change one and the others may want to change too. Vendor benchmarks won't tell you what your workload does, and testing combinations by hand doesn't scale past a handful.
So we built isobench — our own framework for running benchmarks systematically across BIOS settings, OS and driver configurations, serving stacks, and models, then visualising the results side by side.
- Hardware recommendations grounded in measurement Not vendor datasheets or assumptions about what should be fast. If you want the evidence before the order goes in, we'll rent the candidate GPUs in the cloud and run your workload on them.
- A tuned configuration, not a default one We find the settings that suit your workload instead of shipping whatever the installer chose.
- A baseline you can hold us to You know what your hardware does when healthy, so regressions are visible rather than suspected.
- Upgrade decisions with evidence When a new runtime release or model looks promising, we measure it against your actual baseline before recommending a move.
Every customer deployment adds to what we've measured. That's the compounding part, and it's why our recommendations get better over time rather than staler.
Why us
Why isocline
- We do both halves. Integrators sell you a box. Consultants sell you a slide deck. We build the thing and then live with it.
- We measure instead of assuming. Our own benchmarking framework means the configuration you run is the one that tested best for your workload.
- No hardware margin. We don't make money on the parts, which is why we can be vendor-neutral about them and why we'll happily recommend fewer GPUs than you asked for.
- We size honestly. Overselling GPUs is the easiest margin in this business and the fastest way to lose a client.
- We'll talk you out of it. Some workloads genuinely belong in the cloud — bursty, unpredictable, or too small to justify capital. We'd rather tell you that than sell you a rack you don't need.
- Dutch, and staying that way. Your hardware, your premises, your jurisdiction.
Questions
FAQ
We don't have a datacentre. Can we still do this?
You don't need one. For anything HGX-class our default is a Dutch colocation facility — you still own the hardware outright, it's still in the Netherlands, still your jurisdiction. Self-hosting was never about the building.
We find the facility, tell you what to insist on before you sign, design the rack and manage the migration. Two things customers rarely expect: colocation operators sit in a lower energy-tax band than your own building does on identical electricity, and with the grid congested across much of the country, colocation is partly a way to buy a connection you can't otherwise get.
That holds for pilots too — a small deployment doesn't need a full rack, and we'll size the footprint to the build rather than pushing you into space you won't use.
Which models can we run?
Any open-weights model that fits the memory envelope we design for — and we design for the models you name plus headroom for what's next. We've worked with the Llama, Qwen, and Mistral/Mixtral families, including the large mixture-of-experts models.
We already have servers. Can you just help us get more out of them?
Yes. See above — picking up existing hardware and stalled projects is a normal engagement for us, and often the fastest win available.
Who owns the hardware?
You do, outright. We don't mark it up and we don't hold it.
What if we outgrow it?
Capacity planning is part of the support contract, so you'll know months ahead. Adding nodes to a cluster we built is straightforward because we planned the fabric for it.
Can you work alongside our existing infrastructure team?
That's the normal case. We handle the GPU layer and hand your team documentation they can actually use.
Do you support hybrid setups?
Yes, in both senses. Owned hardware for steady-state load with cloud burst for peaks is often the most sensible overall shape. So is a workstation-class development box on your premises with the production node in colocation — something physically present to work on, without putting a 10 kW machine in an office.
What happens if a GPU fails?
Warranty and RMA handling are ours under the support contract. We'll also design a spares strategy with you up front, sized to how much downtime you can tolerate.
Is self-hosting cheaper than renting cloud GPUs?
Sometimes, and it depends entirely on how steady your workload is. We won't promise you a number on a website — we'll model it against your actual usage and tell you honestly if the answer is no.
About us
Assembled around the whole problem.
Self-hosted AI infrastructure usually fails at the seams. The hardware gets specified by people who won't run the workload. The OS, firmware, and drivers get configured by people who didn't choose the hardware. The serving stack gets deployed by people who inherit both and can't change either. Every part is defensible on its own, and the result still leaves half the performance on the floor.
isocline is put together to cover that whole span.
Erik
Accountable for the assessment, the sizing, and the recommendation that comes out of it — including when the honest recommendation is a smaller build than you came for.
David
Infrastructure, across every build. Operating system, firmware, drivers, kernel updates, storage, monitoring, machine load. Everything between the metal and the model, and the part that decides whether the machine is still fast a year from now.
Bjorn
Application layer, across every build. The serving stack selected and tuned until the configuration is the one that measures fastest for your workload, and the gateways in front of it — redundant, so no single failure takes the service down.
A dedicated engineer
For the software side, each project gets an engineer assigned to it for its duration, under Bjorn's lead: the serving and training stack, the gateway and routing layer in front of it, benchmark design, and the tuning that turns a correctly built machine into a fast one. One engineer, one project — not a queue.
The bench
Datacenter installation crews, and two engineers with embedded and applied-mathematics ML backgrounds for work that goes below the framework — hardware-specific inference engines down to the kernels, for customers who want the last measurable margin out of what they've bought.
One point of contact
The person you speak to first stays with it — through the assessment, the build, and on into the support years. Someone who knows your setup, who you reach when something matters, and who doesn't change halfway through.
Everyone here is an engineer. There's no account manager in between, and nobody in a meeting who has to check with the technical team before answering you.
Where we work
We work in the datacenters where the hardware lives, and otherwise distributed. We don't keep an office to receive customers in — for this work, the address that matters is the one the machines are at.
Certifications
The team is certified by the manufacturer of the hardware we build with: in designing and operating AI infrastructure, and in working with generative models and LLMs. Both certificates can be checked online — infrastructure and operations and generative AI and LLMs.
Contact
Tell us what you're trying to run.
Bring your workload, your constraints, and your questions — or the hardware you've already got and can't get moving. You'll get a straight answer on whether self-hosting makes sense for you, and if it doesn't, we'll say so.
Or email → intake@isocline.nl
Any other questions → info@isocline.nl