Novycom
Book a call

NewEvery build now ships with its evaluation suite

Your team stops working the exception queue.

We build and run the systems behind order exceptions, supplier email and catalogue work for commerce teams. Every task carries an accuracy bar agreed in writing.

Below its accuracy bar an agent drafts and waits for a person, so a mistake reaches your team and not your customer. Your data never trains a public model.

Trusted by operators at

NORTHBOUNDKirana GroupVelo Logistics Meridian 3PLHalcyon BeautyTanaka Trading
novycom — order exception triage · live
Monitoring dashboard: accuracy split by error type, escalation rate over ninety days, cost per task, and a live decision log
24Agents in production
128,400Tasks handled / month
−71%Median handling time
6.2%Escalated to a human
7Languages in production
99.94%Uptime, trailing 90d

Capabilities

Six things we do, properly.

We don't sell strategy decks or pilots that die in a slide. Every engagement ends with something running in your stack, owned by your team, with the numbers to prove it earned its place.

Agentic automation

Not chatbots. Systems that read the order, check stock, draft the reply and post the update — with a human in the loop exactly where the cost of being wrong is high.

Retrieval & knowledge

Catalogues, SOPs, tickets and supplier terms turned into something your team and your agents can query — every answer traced to the source clause.

Integration engineering

What makes AI useful is plumbing. We connect the marketplaces, ERPs, WMS and helpdesks you run — and keep those connections alive.

Evaluation & guardrails

You find out an agent is wrong before your customer does. Every deployment ships with a test set, thresholds and an escalation path someone watches.

Data engineering

Most AI projects fail on data, not models. We clean, join and version the operational data first, so the system has something true to reason over.

Run & improve

Models drift, catalogues change, policies get rewritten. We operate what we build and report monthly on the metrics that justify the spend.

Measured outcomes

What clients got back.

Aggregate across deployments running twelve months or longer. We report these monthly — the same numbers you would use to decide whether to keep paying us.

Handling time

71%

Less time per order exception, against the pre-deployment baseline.

Time to production

4 wks

Median from kickoff to the first agent doing real work in a live system.

Volume absorbed

128k

Tasks per month completed without a person touching them.

Cost per task

−83%

Fully loaded cost against the manual process it replaced.

How we work

Four stages. You can stop after any of them.

Each stage produces something you keep — a written scope, a working pilot, a deployed system, a runbook. No stage depends on you signing the next one.

01

Scope

One week · fixed fee

We sit with the team doing the work today and watch the actual process — not the documented one. You get a written scope naming what's worth automating and what isn't.

02

Pilot

Two to three weeks

One task, one team, real data, shadow mode where mistakes cost nothing. The accuracy bar is agreed in writing before we start.

03

Production

Two to four weeks

Live with escalation paths, full logging and an off switch your team controls. We integrate with what you already run.

04

Operate

Monthly · cancel anytime

We watch the evals, retrain on drift, tune cost per task, and report monthly. Or we hand over the runbook — that's a normal ending.

In their words

What operators say.

We ask clients to describe the change in their own terms — including what didn't work the first time.

Two hands closing a carton in near-darkness

Volume was never the hard part.

Hundreds of thousands of parcels a month move without anyone touching them. It is the two percent that do not fit the pattern that consume a team's week — and that is the part every system on this page took over.

The work itself

Twelve jobs we have actually automated.

Not a menu of possibilities. Every item below is running for at least one client right now, and each one started as a task someone was doing by hand at eleven at night.

Order exception triage

Address mismatches, split shipments, stock that moved between checkout and pick. The agent reads the order, checks the WMS, and either resolves it or routes it with the context already attached.

Supplier email handling

Purchase-order confirmations, partial-fill notices and price-change emails parsed into structured updates, so nobody is copying dates out of a PDF into a spreadsheet.

Catalogue enrichment

Titles, attributes, sizing tables and category mapping generated per marketplace, against each platform's own rules — then held to a review queue until accuracy clears the bar.

First-line customer replies

Where-is-my-order, returns eligibility, sizing and stock questions answered from your real policy text, in the customer's language, with the order record in hand.

Returns and claims

Photo evidence assessed against the policy, refund or replacement proposed with a stated confidence, and anything ambiguous pushed to a person with the reasoning shown.

Payment reconciliation

Marketplace settlement files matched to orders, fees and refunds, with the unmatched residue isolated into a queue instead of hidden in a total.

Demand and reorder signals

Velocity, seasonality and lead-time variance turned into a reorder suggestion a buyer can accept, adjust or reject — and the rejections become training data.

Listing compliance sweeps

Continuous checks for claims, restricted terms and category rules across marketplaces, so suspensions get caught as drafts rather than as emails from the platform.

Content localisation

Product and support copy adapted rather than translated — units, sizing conventions, payment names and the tone each market actually reads as normal.

Fraud and abuse review

Serial returners, address anomalies and reseller patterns surfaced with the evidence assembled, so the human decision takes thirty seconds instead of twenty minutes.

Warehouse document intake

Delivery notes, customs paperwork and inbound manifests read on arrival, discrepancies flagged against the PO before the pallet is put away.

Internal knowledge answers

SOPs, supplier terms and past ticket resolutions made queryable for your own staff — the single highest-adoption thing we build, and the least glamorous.

Roughly a third of what clients ask us about does not need a model at all. When a scheduled job or a fixed rule solves it, we say so in the scoping week and you keep the finding.

Integrations

We plug into what you already run.

You are not migrating anything. If a system has an API, a database or a mailbox, it can participate — and when it has none of those, we have built the bridge more than once.

Marketplaces

  • Shopee · Lazada · TikTok Shop
  • Amazon SP-API
  • Shopify & Shopify Plus
  • WooCommerce · Magento
  • Tokopedia · Zalora
  • eBay · Etsy

ERP & finance

  • SAP Business One
  • Microsoft Dynamics 365
  • NetSuite
  • Odoo · Xero · QuickBooks
  • AutoCount · SQL Accounting
  • Custom SQL back-ends

Warehouse & logistics

  • Manhattan · Körber
  • EasyParcel · Ninja Van
  • DHL · FedEx · J&T
  • 3PL portals and SFTP drops
  • Barcode / RFID feeds
  • Spreadsheet-run warehouses

Service & comms

  • Zendesk · Freshdesk · Intercom
  • WhatsApp Business API
  • Gmail · Outlook / Graph
  • Slack · Microsoft Teams
  • Line · Telegram
  • Call-centre transcripts

Data & infrastructure

  • Postgres · MySQL · SQL Server
  • BigQuery · Snowflake
  • Airflow · dbt
  • AWS · GCP · Azure
  • Docker · Kubernetes
  • On-premise where required

Models

  • Anthropic Claude
  • OpenAI GPT family
  • Google Gemini
  • Open-weight models, self-hosted
  • Local embedding models
  • Classical ML where it wins

We are not a reseller for any model vendor and take no referral fees. Model choice is made per task on cost, latency and measured accuracy, and it gets revisited — three of our deployments have changed provider since launch without the client changing anything.

Honest comparison

Four ways to get this done.

We are the right answer for one of the four columns below, not all of them. If your situation fits another column better, we will tell you on the first call.

Novycom Hire in-house Big consultancy Off-the-shelf SaaS
Time to first result Four weeks to production Three to six months to hire, then ramp Six to nine months, discovery-heavy Days — if your process matches theirs
Fits your actual process Built around it; that is the whole point Yes, eventually — they must learn it first Often reshaped to fit their framework You reshape your process to fit the tool
Cost, year one Project fee plus a small monthly run Two to three salaries before output Usually the largest of the four Lowest sticker, per-seat growth after
Who owns the code You do, from day one, in your repo You do Negotiated; sometimes licensed back The vendor
If it does not work We say so in week one and you stop A hiring decision to unwind Change request, new statement of work Cancel, and lose the configuration
Best when The process is specific to you and worth real money AI is core to your product, not your back office Change must be driven across many business units The problem is common and yours is not special

Governance

The boring parts, in writing.

Most of what goes wrong with production AI is not the model being wrong. It is nobody having decided in advance who is accountable when it is.

Your data stays yours

Nothing you give us trains a public model. Where a client requires it, inference runs inside their own cloud account and we never hold a copy at all.

Every action is logged

Input, retrieved context, model, version, cost, output, and who or what approved it. Exportable, and readable by an auditor who has never met us.

Named accuracy thresholds

Each task has a number agreed before build. Below it, the agent does not act autonomously — it drafts and waits. That line is in the contract, not the pitch.

An off switch that works

One control, held by your team, that reverts to the manual process without a deploy, a support ticket, or a call to us at 2am.

PDPA and GDPR posture

Data minimisation at intake, retention windows set per field, deletion honoured downstream. We work to Malaysia's PDPA and to GDPR where an EU market is involved.

Handover is assumed

Runbooks, architecture notes and an onboarding session are deliverables, not favours. Clients ending the retainer and running it themselves is a success, not a churn event.

Engagement models

Three ways to start. All of them end somewhere useful.

Indicative bands, not quotes. The binding number is the one in your scope document at the end of week one — we publish these so you know the order of magnitude before spending an hour on a call.

Start here

Scoping week

USD 3,500Fixed · one week

  • Two days observing the process as it is really done
  • A written scope with effort and expected savings
  • An explicit list of what we would not automate
  • Yours to keep, and to take to another vendor

Most common

Build to production

USD 18k–60kPer system · four to eight weeks

  • Pilot in shadow mode against real data
  • Production deployment inside your stack
  • Evaluation suite, thresholds, escalation paths
  • Full source and IP transfer on delivery

Optional

Operate

USD 2k–8kPer month · cancel any month

  • Eval monitoring and drift response
  • Cost-per-task tuning as volume grows
  • A monthly report you could show a board
  • Or take the runbook and run it yourself

No retainer is required to start, and no engagement carries a minimum term. About one in five clients takes the scope and builds it internally. That is a fine outcome and we have said so before being asked.

Writing

What we have learned, written down.

Notes from deployments, including the ones that went badly. No gated PDFs and no email wall.

Questions

Asked on almost every first call.

If yours is not here, ask it on the call. We would rather answer it before you commit than after.

How small is too small to be worth it?

Our rough line is around twenty hours of human work per week on one repeatable task. Below that the build rarely pays back inside a year, and we will tell you so. Above it, the arithmetic usually makes itself.

Do we need clean data first?

No, and waiting for clean data is how these projects die of old age. We work with what exists and fix the specific fields the task depends on. A full data-quality programme is a different project with a different budget, and it is usually not the one you need.

What if the agent gets something wrong?

It will. The design question is what happens next. Every task gets a stated accuracy bar, a confidence threshold below which the agent drafts instead of acts, an escalation path to a named person, and a log that shows exactly what it saw and why it decided that.

In the first month we run everything in shadow mode: the agent proposes, a human disposes, and we compare. Autonomy is earned per task, never granted by default.

Will this replace our staff?

In our deployments so far it has not. What changes is what the day is made of — the queue that was three hundred routine items and twelve hard ones becomes twelve hard ones. Teams generally redeploy rather than shrink, because the backlog they never had time for is still there.

We will not help you build a case for redundancies. Not a moral position so much as a practical one: the people who know the process are the people who make the system work, and they can tell when they are automating themselves out.

Which model do you use?

Whichever wins on the task, measured. We benchmark two or three candidates during the pilot on your data and choose on accuracy, latency and cost per task together. Systems are built so the model is a swappable component — three of our deployments have changed provider since launch without touching the surrounding code.

How much of our team's time will this take?

During the scoping week, about six hours total from two or three people. During the build, roughly two hours a week from one person who knows the process and can answer questions. If it needs more than that, we have scoped it wrong.

Can you work with our existing IT team?

That is the normal case, and it is the better one. We work in your repo, follow your review process, and hand over as we go rather than at the end. Where an internal team wants to take the second system themselves, we will help them scope it and step back.

What happens to our data?

It stays yours and is never used to train a public model. Retention is set per field during scoping, and where a client requires it, everything runs inside their own cloud account so we never hold a copy. On exit we delete what we hold and send written confirmation of what was deleted and when.

Start here

A 30-minute call, then a written scope.

Bring one process that's eating your team's week. We'll tell you on the call whether it's a good candidate — and we say no to roughly a third of what we're asked about.

  • No NDA needed for the first conversation
  • Written scope with effort and expected savings
  • Fixed fee for the scoping week — no retainer to start
  • You keep the scope whether or not you continue
  • Full IP and source transfer on every build