AI implementation
AI doing real work, with a person checking it.
Most of what a business repeats follows a rule and belongs in ordinary software. A smaller share needs a decision each time: what to tell a customer, which request to take first, where an answer is buried in old records. That share is where AI belongs, and only under conditions — proven against your own history before it starts, watched after launch, and never in front of a customer without a person approving the words.
The fit
Situations we build for
This engagement is for businesses where part of the weekly grind is thinking work — reading, wording, sorting, finding — and it lands on the people who have better uses for the hour. If several of these sound like your week, the fit is probably there.
- Customer replies take real thought each time, and the writing falls on the one or two people who know the business best.
- Everything arrives in a single pile, and somebody reads all of it just to work out what is urgent.
- The answer to most questions exists in past emails, quotes or job notes, but finding it takes longer than the question deserves.
- You tried an off-the-shelf chatbot and switched it off after it told a customer something that was not true.
- The same judgment call gets made dozens of times a week, slightly differently each time, depending on who is on shift.
- You want AI doing measured, supervised work, not another demo that impressed everyone and changed nothing.
The method
What the build actually involves
Six decisions separate AI that quietly helps from AI that gets turned off after one bad week. This is how we make each of them.
Task selection
Not every step deserves a model. Work that follows a written rule — copy this field, send that reminder on day three — runs as ordinary software, which costs less and behaves the same on Friday as it did on Monday. AI is kept for the steps where someone currently has to think: wording a reply, judging what a request is really about, deciding what an old record means. Sorting the two apart is the first working session, and the AI list comes out of it shorter than it went in.
The review gate
Anything written for a customer is produced as a draft and stops at a named person on your team, who edits, approves or discards it before it goes anywhere. This is not a temporary training phase; it is the permanent design. The speed gain survives the checking: working from a competent draft is far quicker than writing from scratch, and the reviewer stays close enough to the output to notice the day quality slips.
Escalation
Every AI step carries a confidence threshold. When the model is unsure, or a request looks unlike anything it was tested on, the item routes to a person with the context gathered so far attached — no answer is forced out. We design against the confident wrong answer, because a polite "someone will get back to you" costs minutes, while a mistake delivered smoothly can cost the customer. The chain always ends with a person, never with a guess.
Testing first
Before an AI step touches live work, it is scored against examples drawn from your records: replies your team actually sent, requests they actually sorted, questions with answers we can check. Each task gets its own test set and its own passing bar, and a step that misses the bar does not launch — the task stays with a person or gets a simpler design. When a failure shows up later in production, it joins that test set, so the same mistake cannot pass quietly twice.
After launch
Going live is the start of the measurement, not the end of it. The system keeps a record of what each AI step did and why, and we watch how often reviewers rewrite drafts — the honest signal of quality — alongside cost and volume. New models from the providers are treated as candidates, not upgrades: each one runs the existing test sets first and takes over a task only where it scores at least as well as the current one. The previous model stays installed, so a change that misbehaves is rolled straight back.
Right-sized models
The most capable models cost many times more per request than the small ones, and most tasks in a business do not need the most capable. Each step runs on the smallest model that passes its tests: sorting a message into six categories is small-model work even when drafting a sensitive reply is not. Because the tests are per task, moving down a size is a decision backed by evidence, not a hope. You see the projected monthly running cost per task before launch, so the operating bill is a figure you approved rather than one you discover.
The engagement
How it runs
Scope and sort
We go through the candidate work with you and split it: rule-following steps to ordinary software, genuine judgment calls to the AI list. You get that list with a projected running cost and a build order, and the fee for each phase is agreed in writing before that phase begins.
Prove it
For every AI step we assemble the test set from your records and measure candidate models against it. You see the results in plain terms — how often the draft matched what your team would have sent — and nothing moves to a build until it clears the bar.
Launch supervised
The system goes live with every review gate on. We watch the first weeks alongside your team, adjusting thresholds where real traffic differs from the history, until the volume reaching people is right.
Hand it over
Code, accounts, documentation and the test sets transfer to you, and the people who will run it are trained on real work. Support after handover is offered and priced on its own; the system keeps working whether we stay involved or not.
What transfers
What you own when we leave
Every engagement is built to be handed over. These are yours once the work is paid for, not licensed back to you.
- The running system itself — the automations, the AI steps, the review queues and the escalation routes — on accounts held in your name.
- The per-task test sets built from your records, which are what let you evaluate the next model without calling us.
- Documentation written for the people who operate it: what each step may do on its own, what needs sign-off, and where to look first when output goes wrong.
- Reviewers and operators trained on live traffic, plus the running-cost picture per task, so the monthly bill stays a known quantity.
Next step
Talk it through first.
Bring the piece of your week that needs judgment but eats time: the inbox, the replies, the hunting through records. On a short intro call we will tell you honestly whether AI belongs in it, and where we would begin.
Book a call