Put AI agents to work

Local AI Agents on Your Own Hardware. Your AI works in-house – your data never leaves.

Setting up local AI hardware with open language models and AI agents in-house

  1. Workshop
  2. Setup
  3. Sprint
  4. Build & Support

Fixed price after scoping · proposal within 48 h

Reviewed: September 2026

Local AI Agents on Your Own Hardware
Local AI Agents on Your Own Hardware
Local AI Agents on Your Own Hardware

Local AI agents are AI assistants whose language model and agent software run on hardware inside the company rather than at a cloud provider.

We set up AI infrastructure on your premises: a Mac Studio, an NVIDIA DGX Spark or a machine with AMD Ryzen AI Max – depending on which models and how many users you need. It runs open language models such as Qwen, DeepSeek, GLM or gpt-oss through a local runtime like llama.cpp, Ollama, MLX or vLLM. For the agents we use Hermes Agent, an open-source agent framework by Nous Research: with memory, reusable skills, a scheduler and connections to Slack, Telegram or email. Via MCP the agents access your systems – with approvals you define. Requests and documents are processed on your hardware; no cloud provider sees them.

Who it's for

Ideal for companies that want to use AI with confidential data – customer data, contracts, engineering drawings, HR files – and cannot or do not want to hand that data to a cloud provider. A strong fit for law firms, medical practices, manufacturing and SMEs with high requirements for data protection and trade secrets.

Key Features

  • On-premise
  • Data-sovereign
  • Secure
  • Protected with Guardrails
  • Least-Privilege Principle
  • Documented
  • Hermes Agent
  • llama.cpp
  • Ollama
  • MLX
  • vLLM
  • Qwen
  • DeepSeek
  • GLM
  • gpt-oss
  • MCP
  • Docker

We select the optimal tech stack for your specific requirements

(01)

Hardware

The machine on your premises. Memory is what matters: it decides how large a model can run.

ClassDevicesMemoryWhat realistically runs
CompactSingle teamsNVIDIA DGX Spark · machines with AMD Ryzen AI Max+128 GBModels up to about 120B parameters, e.g. gpt-oss-120b
WorkstationDepartmentsApple Mac Studio with M5 Ultraup to 512 GBLarge mixture-of-experts models with several hundred billion parameters
ClusterSeveral departments2–4 networked Mac Studios or DGX Sparks256 GB and more, distributedVery large models spread across several devices
EnterpriseCompany-wideNVIDIA DGX Station · workstations with RTX PRO 600096 GB to about 780 GBMany concurrent users in production

Rule of thumb: in 4-bit quantisation a model needs about 0.6 GB of memory per billion parameters, plus room for the context. We make the concrete choice in the workshop, based on your tasks.

(02)

Runtime

The software that runs the model and exposes an OpenAI-compatible interface.

  • llama.cpp
  • Ollama
  • LM Studio
  • MLX
  • vLLM
  • SGLang

MLX on the Mac, vLLM or SGLang on NVIDIA, llama.cpp on AMD – whatever runs fastest on your hardware.

(03)

Models

Open models whose weights live on your hardware. We choose by task, language and licence.

FamilyStrengthLicence
Qwen (Alibaba)Versatile, strong at coding and tool useApache 2.0 for small and medium variants
DeepSeekReasoning, very large variantsMIT
GLM (Zhipu)Agentic tasks and codingMIT or own licence, per variant
gpt-oss (OpenAI)Compact, with tool callingApache 2.0
MistralEuropean vendor, multilingualApache 2.0
Gemma (Google)Small to medium models, images tooApache 2.0

New versions appear every month. We test with your own data before rollout and check the licence for each model.

(04)

Agents

Hermes Agent plans tasks, uses tools and remembers earlier conversations.

  • Memory across conversations
  • Skills the agent derives from completed tasks
  • Scheduler for recurring tasks
  • Reachable via Slack, Telegram, Discord or email
  • Access to your systems via MCP
  • Sub-agents for parallel work

Hermes Agent is open source (MIT licence) and made by Nous Research.

Security and operations

Local does not automatically mean secure. That is why this is part of the setup:

  • Agents run in containers, not directly on the host
  • Sensitive actions need your approval
  • Skills are vetted before installation
  • Guardrails for model behaviour, especially with customer contact
  • A managed update process for models, runtime and agents
  • On request, air-gapped operation without an internet connection

How we work on it

  1. Workshop
  2. Setup
  3. Sprint
  4. Build & Support
(01)

Workshop

A facilitated session that ends with a ranked list of your projects.

½–2 days · three fixed prices
(02)

Setup

We set up one clearly bounded system and hand it over ready to use.

1–2 weeks · fixed price after scoping
(03)

Build & Support

We build the project out and stay alongside you once it is live.

after scoping, ongoing · fixed price after scoping; support billed monthly

Included

(01)

Workshop on tasks, number of users and data as the basis for choosing hardware

(02)

Hardware recommendation and setup: Mac Studio, NVIDIA DGX Spark or a machine with AMD Ryzen AI Max

(03)

Local runtime (e.g. llama.cpp, Ollama, MLX or vLLM) with an OpenAI-compatible interface

(04)

Selection of open models by task, language and licence, such as Qwen, DeepSeek, GLM or gpt-oss

(05)

Agents with Hermes Agent: memory, reusable skills and scheduler

(06)

Access rights, logging and onboarding for your team

(07)

30 days of free bug fixing from final delivery

Not included

(01)

Hardware purchase costs

(02)

Connecting further systems beyond the agreed scope (separately after scoping)

(03)

Legal advice on the EU AI Act or GDPR

(04)

Ongoing support after the 30 days of bug fixing – available as Build & Support

How it runs

(01)

Workshop

In a workshop (½–2 days) we clarify tasks, users and data and decide on hardware, models and licences.

T1
(02)

Hardware & runtime

Setting up the machine on your premises, the runtime and the selected open models.

W1
(03)

Agents & handover

Hermes Agent with skills and connections to your systems, access rights, onboarding and documentation.

W2
(04)

Expansion

Optional: further agents, integrations and user groups step by step, scope after scoping.

Build
(01)What hardware do we need?
That depends on model size and the number of users. For models up to roughly 120 billion parameters in 4-bit quantisation, devices with 128 GB of unified memory are enough, such as an NVIDIA DGX Spark or a machine with AMD Ryzen AI Max+ 395. Larger models need a Mac Studio with M5 Ultra (up to 512 GB; these configurations ship from late October 2026) or several networked devices. In the workshop we decide which class fits your tasks.
(02)Are local models as good as ChatGPT or Claude?
For many business tasks – summarising, searching your own documents, drafting, classifying, coding – current open models are good enough. For very complex tasks the large cloud models are often still ahead. That is why we test the models in advance with your own data, including German-language data, instead of relying on benchmarks.
(03)Does no data really leave the building?
The model's processing runs entirely on your hardware. Whether an agent may also access the internet, for example for a web search, is your decision. Fully air-gapped operation without an internet connection is possible but requires additional hardening for installation and updates – we can plan that in on request.
(04)Why Hermes Agent?
Hermes Agent is an open-source agent framework by Nous Research under the MIT licence. It works with any local model server, comes with memory, skills, a scheduler and connections to Slack, Telegram or email, and connects your systems via MCP. Because it evolves quickly, a managed update process is part of our handover.
(05)May we use models like Qwen, DeepSeek or GLM?
Run locally, these models send no data to their vendor – the model weights are files on your hardware. Many variants are licensed under Apache 2.0 or MIT and can be used commercially; we check the licence for each model. Like any model they need guardrails for their behaviour, and for customer-facing applications we choose model and safeguards with particular care.
(06)Does the EU AI Act apply to local AI too?
Yes. Which obligations apply depends on the purpose, not on where the model runs. Running locally does make control and documentation easier, because you know which model in which version processes which data.

Ready for your project?

Talk to us for 30 minutes with no obligation, or write to us directly.

Fixed price after scoping · proposal within 48 h