I build and repair the infrastructure underneath AI products — GPU inference that costs
too much, video calls that drop and nobody knows why, RAG pipelines that demo beautifully and
fall over in production. I've spent the last two years doing exactly this on a platform serving
500k+ monthly active users.
Engagements are fixed-scope and fixed-price. You know the number before I start.
what I do
Three specialist engagements below. If your problem isn't one of them but you need senior
capacity across a stack, skip to
the retainer.
GPU Inference Cost Audit
$800–1,500 · 1 week
You're burning money on GPUs and you suspect most of it is waste. I profile your serving
stack, then hand you a written plan with measured numbers: batching strategy, quantization,
model server choice, where your utilization is actually going.
I did this at Vooz — moved standalone containers onto Triton with dynamic batching and
took throughput to 1,280+ RPS
on the same hardware.
If I can't find a meaningful reduction, you don't pay.
WebRTC Rescue
$1,500–3,000 · 2 weeks
Your calls drop, reconnection is unreliable, and swapping a headset mid-call kills the
session. WebRTC failures are miserable to debug because they're a state machine problem
wearing a networking costume.
I built and maintain a video coordinator handling handshakes, hardware hot-swap, TURN/STUN
fallback, and recovery from abnormal socket closes —
99.9% session continuity in production.
RAG & Agent Systems, Production-Grade
$2,000–5,000 · 3–4 weeks
Retrieval that holds up on real corpora, evaluation you can trust, caching that keeps the
bill sane, and tracing so you can see why a given answer came out wrong. Built with Claude
and OpenAI APIs, vector and graph stores, and OpenTelemetry throughout. The gap between a
RAG demo and a RAG product is mostly retrieval quality and observability — that's the part I build.
Fractional Senior Engineer
$40–70/hr · weekly retainer
Ongoing senior capacity for a small team: Next.js and TypeScript, React Native (iOS and
Android, including native modules), Python/FastAPI and Node backends, Kubernetes and GitOps
deployment. Useful when you need someone who can move across the whole stack without
a two-week ramp.
what I've shipped
- 500k+
- monthly active users served
- 1,280+
- inference requests/sec
- 99.9%
- WebRTC session continuity
- 5+ yrs
- shipping production systems
Full background on the portfolio
or in the resume.
how it works
-
01
Send me the problem.
A paragraph is enough. I'll tell you honestly whether it's something I can fix —
and if it isn't, I'll say so rather than take the work.
-
02
Short call, then a written scope.
Fixed deliverables, fixed price, fixed end date. No hourly surprises.
-
03
50% up front, 50% on delivery.
You get working code, a written handover, and the reasoning behind the decisions —
not a black box your team can't maintain.
the practical questions
- Will our working hours overlap?
-
Yes. I keep a late schedule and hold a reliable overlap with US Eastern mornings and the
full European working day. I've worked remote-only for four years across distributed teams
— async by default, on a call within a few hours when something is actually on fire.
- Who owns the code?
-
You do, entirely, on final payment. I'll sign your contract and NDA, or provide a simple
one if you'd rather not draft it.
- What if the work turns out to be bigger than we scoped?
-
I tell you as soon as I know, and we decide together — either we cut scope to hold the
date, or we agree a change in writing before I do the extra work. What doesn't happen is a
surprise invoice or a deadline that quietly slips.
- What do you need from me to start?
-
Repository access, whatever observability you already have, and one person on your side who
can answer questions. For inference work, a week of GPU utilization metrics. For WebRTC, logs
from a handful of sessions that actually failed.
- When can you start?
-
This week. I currently have capacity for roughly 30 hours a week and I take on a small number
of engagements at a time, so the work gets real attention rather than a slot in a queue.
Probably not a fit if
you need a full product built from nothing on a fixed budget, ongoing 40-hour-a-week coverage,
or someone to own a system long-term after handover. I'll say so early rather than take the
contract and disappoint you.
start here
Tell me what's broken or what you're trying to ship. I reply to everything within 24 hours.