Services
AI that stays yours
Local LLMs, governed inference, costs you actually understand. When your data can't leave the building and your API bill has stopped making sense, I help you bring AI in-house: deployed, governed, and run with the same discipline as the rest of your platform. One person, hands on, working alongside your team rather than above it.
What I won't do is sell you a transformation. You don't need a GPU cluster sized for a company a thousand times yours, and I won't spec one to pad an invoice. I size the work to the problem: most of the value is in boring, well-governed automation that just runs. And when a cloud API is genuinely the right call, I'll be the first to say so.
Why run AI on your own infrastructure
Open models have reached parity with proprietary ones for most of the work organisations automate day to day: classification, routing, summarisation, extraction. Roughly 80% of real workloads are boring, high-volume jobs that don't need frontier reasoning. They need governance, cost control, and data that stays put. None of that requires paying cloud rates per token.
Cloud AI billing compounds with usage, and the maths gets worse as adoption spreads: subscription plans give way to per-token consumption, and what looked like tooling spend starts to look like infrastructure spend. You also can't govern what you can't see. If your data carries PII, NDA obligations, or regulatory constraints, sending it to a third-party API is a risk you're carrying whether or not anyone has written it down. A private endpoint in your cloud tenancy softens that, but you are still renting the meter, choosing from someone else's model menu, and auditing only as deep as the provider allows.
Here's the part most AI consultancies miss: running AI in-house is a platform problem, not a research problem. Deployment, monitoring, capacity, access control. Your infrastructure team already runs reliable systems every day. What's missing is the bridge between "we want AI" and "we run it ourselves", and that bridge is what I build.
Local LLM deployment
Open models running on your hardware or in your own cloud tenancy, deployed like any other production service: versioned, monitored, and predictable.
Inference governance
Policies, audit trails, and cost controls around every request, so you know who is asking what, with which data, and what it costs. Guardrails your auditors can inspect.
Cost analysis and migration
A clear model of what your cloud AI usage really costs against what running it yourself would, then a migration path for the workloads where the numbers say to move.
Platform engineering for AI
Inference treated like any other critical system: monitoring, scaling, failover, and capacity planning. Reliability engineering, not demo engineering.
Data sovereignty consulting
For financial, healthcare, legal, and research constraints: what can leave the building, what can't, and how to prove it. Systems designed so sensitive data never reaches a third party.
Team upskilling
Getting your infrastructure and ops people confident running local models, so the capability stays in-house instead of depending on me or a vendor.
Who this is for
This work fits organisations where AI has stopped being an experiment and started being a line item. CTOs and infrastructure leaders questioning what their AI spend is buying. CFOs watching API bills compound quarter on quarter. Data teams whose compliance constraints rule out third-party APIs. Platform engineers who are rightly sceptical of adding another SaaS dependency to the stack.
And who it isn't for: if you want a chatbot on your website or a modest-volume AI feature, a cloud API like Claude's is excellent and you should just use it. This page is for when volume, governance, or sovereignty makes that stop being true. A useful rule of thumb: if your total AI spend is under about £1,000 a month, the numbers rarely justify moving, and I'll tell you so.
How I work
I take on a small number of projects at a time. That selectivity is deliberate: it means each engagement gets sustained attention, not the divided focus of someone running ten clients at once.
Every engagement starts with a conversation about the problem, not a quote. I need to understand what you're trying to do before I can say whether I'm the right person to help. If I can't, I'll tell you that clearly, and point you toward what will.
What a typical engagement looks like
- An AI cost and sovereignty assessment: what you're spending, where your data goes, and which workloads local models can already handle
- Scoping document with a clear description of what will be built and what won't
- Build phase with regular check-ins, not silence followed by a delivery
- Deployment with monitoring, alerting, and documented runbooks
- Handover that sticks: your team runs it, and I stay reachable
What's different
AI consulting usually goes wrong in one of two ways. Either you get a pilot that never ships, a demo that impressed everyone and changed nothing. Or you get a platform you'll be paying to feed for years, built for a scale you'll never reach.
I avoid both, because I come at this as a platform engineer rather than a salesperson for a model. I care about cost, observability, and reliability before anything else. I start with the problem, not the technology. I upskill your team as we go, so what I leave behind is capability, not dependency. Every engagement is built to end with your team running the system without me, which is also my answer to the fair question any one-person vendor should face: what happens if I'm not around? You keep everything, and you already know how to run it. And when a cloud API serves you better, I'll say so. You can see the approach in the Cycle Kirklees case study: a production platform I built and still operate, with full monitoring, a workflow non-technical staff can use, and no surprises.
In their words
A few candid things the people I've worked with on platform and delivery engagements have said.
"When things run well, no one notices. It's when they don't that it counts, and in our moment of need, Rutger flagged our issue straight away, was honest with stakeholders, and set about putting it right."
What it costs
I'll be straight with you: the answer starts with an assessment, not a proposal. Before anyone commits to hardware or a migration, I work out what your current AI usage costs, what your data constraints require, and which workloads local models can already handle. That's a scoped piece of work with one fixed price, and it pays for itself by telling you whether bringing AI in-house is worth doing at all. Sometimes the honest answer is that it isn't yet, and you'll have that in writing before spending anything more.
Larger engagements are priced the same way: I scope the problem first and give you one fixed price, with any ongoing support as a separate monthly fee, and only when the work genuinely needs one. You should only ever pay for what you need, and you can trust me to size the work, quote you the real number before anything begins, and never sell you more than the job requires.
Let's work out what you actually need
Start with the assistant below. It's an AI trained on how I work: describe what you're weighing up, whether that's an API bill that keeps growing, data that can't leave the building, or a workload you suspect could run in-house, and it will help you see whether an assessment is the sensible next step. It won't quote you or commit me to anything; it's here to help us work out, together, whether this is worth a proper conversation. It also practises what this page preaches: the assistant runs on an open model, not a frontier API.
Tap to describe what you're weighing up →
Prefer a callback? Leave your details instead email or phone
Prefer to reach out directly? Email [email protected] or connect on LinkedIn.
Frequently asked questions
Common questions before the first conversation.
Frequently asked questions
Common questions before the first conversation.
Is local AI actually good enough?
For most of what organisations automate, yes. Classification, routing, extraction, and summarisation are well within reach of current open models, and the gap with proprietary models has closed for that kind of work. For genuinely hard reasoning work, cloud models still win, and I'll tell you when your workload is in that bucket. Sorting your workloads into those two buckets is exactly what the assessment is for.
We already use Claude or OpenAI. Do we have to rip that out?
No, and I wouldn't recommend it. Hybrid is the normal end state: high-volume and high-sensitivity workloads move in-house where the economics or the compliance case is clear, and cloud APIs stay where they're the right tool. The goal is knowing which workload belongs where, not ideological purity.
What about Azure OpenAI or Bedrock in our own tenancy?
Sometimes that's the right call, and the assessment treats it as an option to price, not a competitor to beat. A private endpoint answers part of the sovereignty question, but you're still paying per token at volume, choosing from the models your provider licenses, and auditing only as deep as they allow, and the exit gets harder the more you build on it. Where none of that bites, use it. Where cost at volume, model choice, or provable auditability matters, open models on your own infrastructure are the stronger position.
What hardware do we need?
Usually less than the hype suggests. The workloads that make sense to run locally mostly fit on a single GPU server or a modest node, and many organisations have more suitable capacity than they think. Exact sizing comes out of the assessment, priced against your current API spend rather than in the abstract, so the comparison is real numbers, not vendor benchmarks.
Do we need a data science team?
No. Deploying and governing existing open models is a platform engineering problem: deployment, monitoring, capacity, access control. That's work your infrastructure and ops team already understands, and they're exactly who should own it. Part of every engagement is getting them confident with the AI-specific parts, so the capability stays with you.
Who owns what at the end?
You own everything: the infrastructure code, the configuration, the model weights under their open licences, the runbooks, and the dashboards. There are no proprietary platforms and nothing designed to make leaving difficult. Success, as I define it, is your team running the system without needing me.
How long does an engagement take?
The assessment takes a few weeks. A deployment engagement usually runs a few months, depending on scope. Either way, every engagement starts with a written scope and one fixed price before anything begins, and regular check-ins are part of how I work; I don't go quiet mid-build.