Skip to content

AI systems

AI that does real work inside a business, not demonstrations.

Agents that act, retrieval that answers from your own records, and extraction that turns messy documents into data your systems can use. Scoped, tested and recorded, so you can show what each one did.

Why this matters now

40%+of agentic AI projects will be cancelled by the end of 2027 — due to escalating costs, unclear business value or inadequate risk controls.Source: Gartner, June 2025
33%of enterprise software applications will include agentic AI by 2028, up from less than 1% in 2024.Source: Gartner, June 2025
15%of day-to-day work decisions will be made autonomously through agentic AI by 2028, up from none in 2024.Source: Gartner, June 2025

AI is moving from answering questions to taking actions — and the projects that fail are not failing on the model. They fail on cost nobody controlled, value nobody defined, and controls nobody built. The technology is ready for production. Most implementations of it are not.

What we build

Systems that act

AI agents

Systems that take action — reading a queue, deciding, updating a record, and escalating what they cannot handle. Each agent has a defined set of permissions and records every action it takes. A person approves anything irreversible.

Where it applies: triaging a support queue and drafting replies for approval · processing refund or credit requests up to a set limit · updating the CRM from inbound emails and call notes.

Agent access layer

Agents need safe access to your core systems. We build the tools they call, a separate identity and credentials for each agent, and a gateway in front of your systems that records every call. We use the simplest interface that works — which is not always a new protocol.

Where it applies: several agents reading and writing the ERP without sharing one admin key · giving staff AI assistants controlled access to internal systems.

Voice agents

Voice agents connected to your phone lines and core systems, under the same permissions, approvals and action records as every other channel. Calls they cannot resolve are handed to a person with the context attached.

Where it applies: booking, rescheduling and confirming appointments · order and delivery status enquiries · first-line triage before a call reaches a person.

Systems that read and answer

Retrieval & knowledge systems

Question answering and search over your own documents, tickets and records, with citations back to the source. Keyword search, vector search and, where relationships matter, a knowledge graph. Access control applies at retrieval, so people only see what they are entitled to. Built into a workflow, not as a standalone chatbot.

Where it applies: support agents answering from policies and past tickets · legal teams finding clauses and obligations across contracts · engineers searching manuals, specifications and incident history.

Messy document processing

Invoices, receipts, contracts, scanned forms and drawings turned into structured data. OCR and vision-language models, with each field checked against business rules and given a confidence score. Low-confidence fields go to a person. Clean data goes straight into your ERP, CRM or database.

Where it applies: accounts payable processing supplier invoices in any layout · customer onboarding and KYC forms, including handwritten scans · bills of lading and customs forms.

LLM applications

Classification, drafting, summarisation and triage — built where the task is language work a person currently does by hand. Each one measured against an evaluation set built from your own examples.

Where it applies: routing inbound email to the right team with a priority · drafting first-pass proposals from structured data · summarising call transcripts into CRM notes.

Small models & cost routing

Each request sent to the cheapest model that passes your evaluation set. Where volume justifies it, a small open-weight model fine-tuned or distilled to run faster, cost less, and run privately on infrastructure you control.

Where it applies: high-volume classification where API costs are the main expense · latency-sensitive steps inside a larger pipeline · workloads whose data cannot leave your environment.

What comes with it

Every AI system we build ships with these, beyond the system itself:

  • An evaluation set built from your own cases — so whether it works is a measured answer, not an impression
  • Adversarial testing before release — deliberate attempts to make it disclose, act or overreach
  • Rate, token and spend limits configured before anything reaches production
  • A record of every action taken, readable by a person
  • Human approval gates on anything irreversible, costly or externally visible

Built secure and evidenced, as standard

Built so you can show what it does.

  • Security and compliance are not a separate service we sell. They are how the work is built, on every engagement.
  • For AI systems, that means three things in particular. Every agent is scoped to the narrowest capability that does the job, with its own identity and an explicitly enumerated set of tools — never a shared admin key. Retrieval systems enforce access control at the moment of retrieval, so a query can only reach documents the person asking is entitled to see. And every action a system takes is recorded, so when someone asks what it did and who allowed it, the answer already exists.
  • Everything we build in this area is delivered against our published engineering standard, and ships with the evidence to prove it: a threat model, the evaluation results, and signed build provenance.

How the work runs

  1. 01

    Scoping — what it will do, what data it touches, what it may act on.

  2. 02

    Design — architecture and threat model, written down.

  3. 03

    Build — against the published standard, every merge reviewed.

  4. 04

    Handover — in repositories you own, with documentation and a runbook.

Typical duration: 3 to 12 weeks, depending on the system. Focused language applications sit at the short end; agents with write access to core systems at the long end.

How we work

What we will tell you

A substantial share of the requests we get for custom AI are better solved with retrieval, a clearer process, or a smaller model than the one the client had in mind. Plenty of what is sold as agentic today does not need to be an agent at all.

We say that during scoping, not after invoicing.

Tell us what you’re building.

Describe the problem in your own words. We will tell you honestly whether we are the right people for it.

modularitiEngineering  :  
Available