Service

Custom AI integrations for the systems you already run

A custom AI integration puts a language model inside a workflow you already run: sorting incoming email, summarising a call, pulling fields out of a document, drafting a reply for someone to approve. The model is the easy part. What keeps it working is a set of test cases with known right answers, a backup plan for when the provider goes down, and a spending cap so a stuck loop cannot bill you for a fortnight.

How the process works

Sound familiar?

The signs this is your problem

If two or more of these are true, there is almost certainly a worthwhile automation hiding in it.

  • An AI feature looked convincing when you trialled it, then broke in real use
  • Nobody can tell you how accurate the AI step actually is
  • Your API bill is unpredictable and nobody knows what drives it
  • The AI does something odd occasionally and there is no audit trail
  • It works until the provider has an outage, and then everything stops

What gets built

What an integration build includes

Four things, in roughly this order. The proportions shift by project — the shape does not.

The integration itself

Sorting, pulling out data, summarising, drafting or routing — placed at the exact point in your workflow where it saves work. The output comes back in a fixed shape, so your other systems can rely on it.

An evaluation harness

A test set drawn from your real data with known-correct answers, so accuracy is a measured number you can see rather than a feeling. It re-runs whenever a prompt or model changes, which is how you catch a regression before your customers do.

Cost controls and rate limits

Hard spend ceilings, per-request token budgets, caching for repeated inputs, and alerts long before a bill becomes a problem. You get a written estimate of monthly running cost before the build starts.

Fallbacks and audit logging

A planned response to each way it can go wrong: the provider goes down, the answer comes back malformed, the model is unsure. Every call is logged with what went in, what came out and what it cost, somewhere you can look at it.

What you get

Handed over, documented, yours.

  • The working integration, running in your infrastructure or a provider account you own
  • An evaluation report with measured accuracy on your own data
  • A written monthly running-cost estimate with the assumptions shown
  • Audit logging of every model call
  • Documentation covering prompts, fallbacks and how to change them safely

Typically built with

OpenAIAnthropic ClaudeGoogle GeminiPythonn8nPlaywright

Chosen per project, not by habit. Everything runs in accounts registered to you — you hold the keys and see the bills directly.

When this is not the right service

If you need a model trained or fine-tuned from scratch on proprietary data at scale, you want an ML engineer rather than an automation consultant. I will say so and, where I can, point you somewhere sensible.

How we'd work

From audit to autopilot in weeks.

Fixed scope. One fixed price. Working software every week. You always know what ships next and what it saves.

  1. Week 0 · Free

    Audit

    Thirty minutes on where your hours actually go. I map the repetitive workflows, time them, and score each one by what automating it would return. You keep that map either way — including the honest note on which processes to leave alone.

  2. Week 1

    Blueprint

    A written proposal: what gets built, which systems it touches, what it costs to run each month, the timeline, and one fixed price for the whole engagement. Nothing starts until you approve it. The price does not move unless you change the scope.

  3. Weeks 2–4

    Build

    I design, build and test inside your own accounts, with evaluations and guardrails on anything using a language model. You see working software every week rather than a status update — which means you can redirect early, while redirecting is still cheap.

  4. Ongoing

    Run & improve

    Launch, monitor, iterate. Thirty days of support included with every build. After that: take a care plan, hand it to your own team using the documentation, or run it yourself. All three are genuinely fine.

Questions

Custom AI integrations — the usual questions

Primarily OpenAI, Anthropic and Google models, chosen per task rather than by loyalty — different models are genuinely better at different jobs, and cost per task varies by an order of magnitude. Everything runs through API accounts you own, so you keep control of the keys, the billing and the data-handling terms.

Three things, all in place before launch. A hard monthly spending cap on the account. A limit per request, so one bad input cannot set off a long, expensive loop. And caching, so the same input is never paid for twice. You also get a written estimate of the monthly cost before the build starts, with the assumptions shown so you can check them yourself.

Every integration I build has a defined fallback. Depending on the workflow that means failing over to a second provider, queueing the work for retry, or routing to a human with a clear notice — but never silently dropping the task or writing a wrong answer into your system of record.

Ready to get the hours back?

A free 30-minute audit call. You leave with a written map of your automatable workflows and what each one is worth — whether we end up working together or not.

Replies within 24 hours · No sales team — you talk to the builder

Chat on WhatsApp