Agentic automation
Not chatbots. Systems that read the order, check stock, draft the reply and post the update — with a human in the loop exactly where the cost of being wrong is high.
NewEvery build now ships with its evaluation suite
We build and run the systems behind order exceptions, supplier email and catalogue work for commerce teams. Every task carries an accuracy bar agreed in writing.
Below its accuracy bar an agent drafts and waits for a person, so a mistake reaches your team and not your customer. Your data never trains a public model.
Trusted by operators at



Capabilities
We don't sell strategy decks or pilots that die in a slide. Every engagement ends with something running in your stack, owned by your team, with the numbers to prove it earned its place.
Not chatbots. Systems that read the order, check stock, draft the reply and post the update — with a human in the loop exactly where the cost of being wrong is high.
Catalogues, SOPs, tickets and supplier terms turned into something your team and your agents can query — every answer traced to the source clause.
What makes AI useful is plumbing. We connect the marketplaces, ERPs, WMS and helpdesks you run — and keep those connections alive.
You find out an agent is wrong before your customer does. Every deployment ships with a test set, thresholds and an escalation path someone watches.
Most AI projects fail on data, not models. We clean, join and version the operational data first, so the system has something true to reason over.
Models drift, catalogues change, policies get rewritten. We operate what we build and report monthly on the metrics that justify the spend.
Measured outcomes
Aggregate across deployments running twelve months or longer. We report these monthly — the same numbers you would use to decide whether to keep paying us.
Handling time
71%Less time per order exception, against the pre-deployment baseline.
Time to production
4 wksMedian from kickoff to the first agent doing real work in a live system.
Volume absorbed
128kTasks per month completed without a person touching them.
Cost per task
−83%Fully loaded cost against the manual process it replaced.
How we work
Each stage produces something you keep — a written scope, a working pilot, a deployed system, a runbook. No stage depends on you signing the next one.
01
One week · fixed fee
We sit with the team doing the work today and watch the actual process — not the documented one. You get a written scope naming what's worth automating and what isn't.
02
Two to three weeks
One task, one team, real data, shadow mode where mistakes cost nothing. The accuracy bar is agreed in writing before we start.
03
Two to four weeks
Live with escalation paths, full logging and an off switch your team controls. We integrate with what you already run.
04
Monthly · cancel anytime
We watch the evals, retrain on drift, tune cost per task, and report monthly. Or we hand over the runbook — that's a normal ending.
In their words
We ask clients to describe the change in their own terms — including what didn't work the first time.
Hundreds of thousands of parcels a month move without anyone touching them. It is the two percent that do not fit the pattern that consume a team's week — and that is the part every system on this page took over.
The work itself
Not a menu of possibilities. Every item below is running for at least one client right now, and each one started as a task someone was doing by hand at eleven at night.
Address mismatches, split shipments, stock that moved between checkout and pick. The agent reads the order, checks the WMS, and either resolves it or routes it with the context already attached.
Purchase-order confirmations, partial-fill notices and price-change emails parsed into structured updates, so nobody is copying dates out of a PDF into a spreadsheet.
Titles, attributes, sizing tables and category mapping generated per marketplace, against each platform's own rules — then held to a review queue until accuracy clears the bar.
Where-is-my-order, returns eligibility, sizing and stock questions answered from your real policy text, in the customer's language, with the order record in hand.
Photo evidence assessed against the policy, refund or replacement proposed with a stated confidence, and anything ambiguous pushed to a person with the reasoning shown.
Marketplace settlement files matched to orders, fees and refunds, with the unmatched residue isolated into a queue instead of hidden in a total.
Velocity, seasonality and lead-time variance turned into a reorder suggestion a buyer can accept, adjust or reject — and the rejections become training data.
Continuous checks for claims, restricted terms and category rules across marketplaces, so suspensions get caught as drafts rather than as emails from the platform.
Product and support copy adapted rather than translated — units, sizing conventions, payment names and the tone each market actually reads as normal.
Serial returners, address anomalies and reseller patterns surfaced with the evidence assembled, so the human decision takes thirty seconds instead of twenty minutes.
Delivery notes, customs paperwork and inbound manifests read on arrival, discrepancies flagged against the PO before the pallet is put away.
SOPs, supplier terms and past ticket resolutions made queryable for your own staff — the single highest-adoption thing we build, and the least glamorous.
Roughly a third of what clients ask us about does not need a model at all. When a scheduled job or a fixed rule solves it, we say so in the scoping week and you keep the finding.
Integrations
You are not migrating anything. If a system has an API, a database or a mailbox, it can participate — and when it has none of those, we have built the bridge more than once.
We are not a reseller for any model vendor and take no referral fees. Model choice is made per task on cost, latency and measured accuracy, and it gets revisited — three of our deployments have changed provider since launch without the client changing anything.



Honest comparison
We are the right answer for one of the four columns below, not all of them. If your situation fits another column better, we will tell you on the first call.
| Novycom | Hire in-house | Big consultancy | Off-the-shelf SaaS | |
|---|---|---|---|---|
| Time to first result | Four weeks to production | Three to six months to hire, then ramp | Six to nine months, discovery-heavy | Days — if your process matches theirs |
| Fits your actual process | Built around it; that is the whole point | Yes, eventually — they must learn it first | Often reshaped to fit their framework | You reshape your process to fit the tool |
| Cost, year one | Project fee plus a small monthly run | Two to three salaries before output | Usually the largest of the four | Lowest sticker, per-seat growth after |
| Who owns the code | You do, from day one, in your repo | You do | Negotiated; sometimes licensed back | The vendor |
| If it does not work | We say so in week one and you stop | A hiring decision to unwind | Change request, new statement of work | Cancel, and lose the configuration |
| Best when | The process is specific to you and worth real money | AI is core to your product, not your back office | Change must be driven across many business units | The problem is common and yours is not special |
Governance
Most of what goes wrong with production AI is not the model being wrong. It is nobody having decided in advance who is accountable when it is.
Nothing you give us trains a public model. Where a client requires it, inference runs inside their own cloud account and we never hold a copy at all.
Input, retrieved context, model, version, cost, output, and who or what approved it. Exportable, and readable by an auditor who has never met us.
Each task has a number agreed before build. Below it, the agent does not act autonomously — it drafts and waits. That line is in the contract, not the pitch.
One control, held by your team, that reverts to the manual process without a deploy, a support ticket, or a call to us at 2am.
Data minimisation at intake, retention windows set per field, deletion honoured downstream. We work to Malaysia's PDPA and to GDPR where an EU market is involved.
Runbooks, architecture notes and an onboarding session are deliverables, not favours. Clients ending the retainer and running it themselves is a success, not a churn event.
Engagement models
Indicative bands, not quotes. The binding number is the one in your scope document at the end of week one — we publish these so you know the order of magnitude before spending an hour on a call.
Start here
USD 3,500Fixed · one week
Most common
USD 18k–60kPer system · four to eight weeks
Optional
USD 2k–8kPer month · cancel any month
No retainer is required to start, and no engagement carries a minimum term. About one in five clients takes the scope and builds it internally. That is a fine outcome and we have said so before being asked.
Writing
Notes from deployments, including the ones that went badly. No gated PDFs and no email wall.
Field note
The demo runs on twenty hand-picked examples. Production runs on the twenty thousand nobody looked at. Here is what breaks in the gap, and the three checks that catch it early.
Method
A single percentage hides the only distinction that matters in operations: the mistakes you can absorb versus the ones that reach a customer. How to set thresholds that survive contact with a real queue.
Opinion
Six situations where a scheduled job, a lookup table or a two-hour conversation beats anything we could build — and how to recognise them before you have spent the budget.
Questions
If yours is not here, ask it on the call. We would rather answer it before you commit than after.
Our rough line is around twenty hours of human work per week on one repeatable task. Below that the build rarely pays back inside a year, and we will tell you so. Above it, the arithmetic usually makes itself.
No, and waiting for clean data is how these projects die of old age. We work with what exists and fix the specific fields the task depends on. A full data-quality programme is a different project with a different budget, and it is usually not the one you need.
It will. The design question is what happens next. Every task gets a stated accuracy bar, a confidence threshold below which the agent drafts instead of acts, an escalation path to a named person, and a log that shows exactly what it saw and why it decided that.
In the first month we run everything in shadow mode: the agent proposes, a human disposes, and we compare. Autonomy is earned per task, never granted by default.
In our deployments so far it has not. What changes is what the day is made of — the queue that was three hundred routine items and twelve hard ones becomes twelve hard ones. Teams generally redeploy rather than shrink, because the backlog they never had time for is still there.
We will not help you build a case for redundancies. Not a moral position so much as a practical one: the people who know the process are the people who make the system work, and they can tell when they are automating themselves out.
Whichever wins on the task, measured. We benchmark two or three candidates during the pilot on your data and choose on accuracy, latency and cost per task together. Systems are built so the model is a swappable component — three of our deployments have changed provider since launch without touching the surrounding code.
During the scoping week, about six hours total from two or three people. During the build, roughly two hours a week from one person who knows the process and can answer questions. If it needs more than that, we have scoped it wrong.
That is the normal case, and it is the better one. We work in your repo, follow your review process, and hand over as we go rather than at the end. Where an internal team wants to take the second system themselves, we will help them scope it and step back.
It stays yours and is never used to train a public model. Retention is set per field during scoping, and where a client requires it, everything runs inside their own cloud account so we never hold a copy. On exit we delete what we hold and send written confirmation of what was deleted and when.
Start here
Bring one process that's eating your team's week. We'll tell you on the call whether it's a good candidate — and we say no to roughly a third of what we're asked about.