(Service)

AI-Powered Applications & Tools

AI systems that run on infrastructure you control: air-gapped inside an isolated network, on your own servers, or in your private cloud. Open-weight models fine-tuned on your documents, retrieval over your archives, automation pipelines, and agents scoped to permissions you set.

What we build

Air-gapped LLM deployment

Weights, inference server, vector database, and interface installed inside an isolated network with no outbound route.

Self-hosted and private cloud

The same stack on your own servers or in your VPC, where full isolation is more than the work requires.

Fine-tuning on your domain

Open-weight models adapted to your vocabulary, document formats, and correct-answer standard, using LoRA or full fine-tuning depending on the data you have.

Retrieval over your archive (RAG)

Contracts, tickets, manuals, and wikis indexed and searchable, with answers cited back to the source paragraph.

Automation pipelines

Document intake, classification, extraction, routing, and the downstream action, end to end rather than one step in the middle.

AI agents with guardrails

Multi-step tool use bounded by explicit permissions, with audit logging and human approval on anything costly to get wrong.

Internal assistants and copilots

Chat and in-app assistants for your own staff, connected to your systems and scoped to each user's access.

Hybrid routing

Sensitive work pinned to the local model, the rest routed to a frontier API when it earns the trip. You set the boundary, the router enforces it.

Evaluation harnesses

A test set drawn from your real cases, scored on accuracy, cost per run, and latency, run again on every change.

Integration into existing software

Adding these into a product that already has users, without taking it out of service.

How a deployment runs

1. Workflow review

The task observed being done by hand, timed, and counted for errors, with an agreed definition of a good result.

2. Data and model selection

Reviewing the documents and records available, then benchmarking candidate open-weight models against a test set built from real cases.

3. Infrastructure sizing

GPU and memory requirements calculated against real concurrency, with the numbers written down before any hardware is bought or rented.

4. Build and tune

Retrieval, prompts, fine-tuning where it helps, guardrails, and the interface or API the system is used through.

5. Evaluation

Scored against the agreed threshold. Below it, nothing ships, and the failure cases are handed over rather than a summary.

6. Deployment and handover

Installed in the target environment, monitoring in place, runbooks written, and training on how to retrain and update it.

What stays yours

The weights

The fine-tuned model is a file on your infrastructure, not a subscription. It keeps working if you stop working with us.

The training data

Your material never enters a third-party training corpus. In a self-hosted deployment it never reaches a third party at all.

The evaluation set

The test cases and scores are handed over, so you can measure any future model against the same bar.

The code and infrastructure

Pipelines, prompts, and deployment configuration in your repositories and your accounts.

The trade-offs, stated plainly

Running your own models costs GPU capacity, bought or rented, and someone to keep the stack patched. A tuned open model won't match a frontier model at open-ended reasoning across every domain, though on a defined task with your own data behind it the gap is usually small. Where a public API is the better tool and nothing prevents you using one, we'll say so rather than build the private version for its own sake.

Who this is for

Organisations under confidentiality, regulatory, or client constraints that rule out sending data to a public API. Teams whose API bill has outgrown the value it returns. Companies with years of documents that should be answering questions and currently answer none.

Frequently asked

Can this run with no internet connection at all?

Yes. In an air-gapped deployment the weights, inference server, vector database, and interface all sit inside your isolated network, and there is no outbound route for the system to use. Updates come in through your existing change process.

What hardware do we need?

For a departmental workload, usually a single GPU server. Quantised models run useful workloads on hardware many teams already have. We size it against your real concurrency and write the numbers down before anything is bought.

Which models do you use?

Open-weight families: Llama, Mistral, Qwen, Gemma, and their successors. The choice comes from how they score on your evaluation set, not from benchmarks.

Is our data used to train anything?

Only your own model, and that model stays with you. Nothing goes into a third-party training corpus. The tuned weights, the training data, and the evaluation set are all yours.

How long does an AI project take?

Two to six weeks for one workflow with evaluation included. Air-gapped installations take longer because of hardware lead time, network approval, and the security review. Agent systems vary, and we won't quote until we've watched the task done manually.

How do you stop it making things up?

Answers grounded in your documents with citations, an evaluation set built from real cases, guardrails on what the system can do, and human approval where a wrong answer is expensive. Hallucination is reduced this way, not eliminated.

How will we know it's working?

You get a number before and after: tickets routed correctly, contracts flagged, documents processed without a person touching them. The measure gets agreed in the first week.

Can you add this to our existing product?

Yes, and most of this work is exactly that. Software that already has users and can't be taken offline for a rewrite.

Who maintains it afterwards?

Either side. We hand over runbooks and train your team to update models and retrain on new data, or we keep running it under a support arrangement.