AI-Powered Applications & Tools
AI systems that run on infrastructure you control: air-gapped inside an isolated network, on your own servers, or in your private cloud. Open-weight models fine-tuned on your documents, retrieval over your archives, automation pipelines, and agents scoped to permissions you set.
What we build
Air-gapped LLM deployment
Weights, inference server, vector database, and interface installed inside an isolated network with no outbound route.
Self-hosted and private cloud
The same stack on your own servers or in your VPC, where full isolation is more than the work requires.
Fine-tuning on your domain
Open-weight models adapted to your vocabulary, document formats, and correct-answer standard, using LoRA or full fine-tuning depending on the data you have.
Retrieval over your archive (RAG)
Contracts, tickets, manuals, and wikis indexed and searchable, with answers cited back to the source paragraph.
Automation pipelines
Document intake, classification, extraction, routing, and the downstream action, end to end rather than one step in the middle.
AI agents with guardrails
Multi-step tool use bounded by explicit permissions, with audit logging and human approval on anything costly to get wrong.
Internal assistants and copilots
Chat and in-app assistants for your own staff, connected to your systems and scoped to each user's access.
Hybrid routing
Sensitive work pinned to the local model, the rest routed to a frontier API when it earns the trip. You set the boundary, the router enforces it.
Evaluation harnesses
A test set drawn from your real cases, scored on accuracy, cost per run, and latency, run again on every change.
Integration into existing software
Adding these into a product that already has users, without taking it out of service.
How a deployment runs
1. Workflow review
The task observed being done by hand, timed, and counted for errors, with an agreed definition of a good result.
2. Data and model selection
Reviewing the documents and records available, then benchmarking candidate open-weight models against a test set built from real cases.
3. Infrastructure sizing
GPU and memory requirements calculated against real concurrency, with the numbers written down before any hardware is bought or rented.
4. Build and tune
Retrieval, prompts, fine-tuning where it helps, guardrails, and the interface or API the system is used through.
5. Evaluation
Scored against the agreed threshold. Below it, nothing ships, and the failure cases are handed over rather than a summary.
6. Deployment and handover
Installed in the target environment, monitoring in place, runbooks written, and training on how to retrain and update it.
What stays yours
The weights
The fine-tuned model is a file on your infrastructure, not a subscription. It keeps working if you stop working with us.
The training data
Your material never enters a third-party training corpus. In a self-hosted deployment it never reaches a third party at all.
The evaluation set
The test cases and scores are handed over, so you can measure any future model against the same bar.
The code and infrastructure
Pipelines, prompts, and deployment configuration in your repositories and your accounts.
The trade-offs, stated plainly
Running your own models costs GPU capacity, bought or rented, and someone to keep the stack patched. A tuned open model won't match a frontier model at open-ended reasoning across every domain, though on a defined task with your own data behind it the gap is usually small. Where a public API is the better tool and nothing prevents you using one, we'll say so rather than build the private version for its own sake.
Who this is for
Organisations under confidentiality, regulatory, or client constraints that rule out sending data to a public API. Teams whose API bill has outgrown the value it returns. Companies with years of documents that should be answering questions and currently answer none.
Frequently asked
Can this run with no internet connection at all?
Yes. In an air-gapped deployment the weights, inference server, vector database, and interface all sit inside your isolated network, and there is no outbound route for the system to use. Updates come in through your existing change process.
What hardware do we need?
For a departmental workload, usually a single GPU server. Quantised models run useful workloads on hardware many teams already have. We size it against your real concurrency and write the numbers down before anything is bought.
Which models do you use?
Open-weight families: Llama, Mistral, Qwen, Gemma, and their successors. The choice comes from how they score on your evaluation set, not from benchmarks.
Is our data used to train anything?
Only your own model, and that model stays with you. Nothing goes into a third-party training corpus. The tuned weights, the training data, and the evaluation set are all yours.
How long does an AI project take?
Two to six weeks for one workflow with evaluation included. Air-gapped installations take longer because of hardware lead time, network approval, and the security review. Agent systems vary, and we won't quote until we've watched the task done manually.
How do you stop it making things up?
Answers grounded in your documents with citations, an evaluation set built from real cases, guardrails on what the system can do, and human approval where a wrong answer is expensive. Hallucination is reduced this way, not eliminated.
How will we know it's working?
You get a number before and after: tickets routed correctly, contracts flagged, documents processed without a person touching them. The measure gets agreed in the first week.
Can you add this to our existing product?
Yes, and most of this work is exactly that. Software that already has users and can't be taken offline for a rewrite.
Who maintains it afterwards?
Either side. We hand over runbooks and train your team to update models and retrain on new data, or we keep running it under a support arrangement.