Fine-tuning for language models

LLM Fine-Tuning Services

LLM fine-tuning further trains a pre-trained language model on your examples so that it reliably matches your terminology, tone and output formats, where prompts and RAG alone are not enough. Context Studios, an AI-native development studio in Berlin, handles data preparation, training, evaluation and operation, fully on your own infrastructure if you wish.

Observatory with a patinated copper dome under an overcast sky, photographed from belowAI-generated image
Domain-specific adaptationOpen-source and commercial modelsQuality measured by evaluationGDPR-compliant training
  1. Workshop
  2. Setup
  3. Sprint
  4. Build & Support

Fixed price after scoping · proposal within 48 h

Last updated:

(01)

What is LLM fine-tuning?

AI technology

LLM fine-tuning is the further training of a pre-trained language model on domain-specific examples to specialise it for particular tasks, terminology or output formats. While RAG adds knowledge at runtime, fine-tuning changes the model's behaviour itself, such as the style, terminology and structure of its answers.

Specialisation
LoRA, QLoRA, full fine-tuning, DPO
Technologies
Hugging Face Transformers, Axolotl, Unsloth, vLLM
Target group
Companies with specific terminology, fixed output formats or data protection requirements
Project duration
Typically 4–12 weeks including data preparation
Compliance
GDPR, data sovereignty, on-premise training possible

LLM developmentRAG developmentMachine learning developmentGenerative AI development

(02)

Which fine-tuning services do we offer?

From curating training data to an optimised production model

(01)

Training data preparation

We curate, clean and structure your examples, remove personal data and fill gaps in coverage; data quality determines the result.

(02)

LoRA and QLoRA fine-tuning

Parameter-efficient methods adapt only a small part of the model weights. That makes training and operation considerably cheaper and allows several specialised variants of one base model.

(03)

Domain specialisation

The model learns your terminology, style and answer structures, for example for reports, product copy, classification or structured extraction.

(04)

Evaluation & benchmarking

Automated test sets, human review and A/B tests measure whether the adapted model beats the base model and prompting, objectively and reproducibly.

(05)

Preference training & optimisation

With DPO or comparable methods we align answers with your preferences; quantisation then reduces memory requirements and latency.

(06)

Self-hosted deployment

From training to production use with vLLM on your own or European infrastructure, with pipelines for regular retraining.

(03)

How does a fine-tuning project work?

  1. (01)

    Initial call

    A free 30-minute video call with Michael Kerkhoff. We get to know your project, assess where AI adds value and give you a first estimate of feasibility, effort and timeframe.

    Step 1
  2. (02)

    Proposal & planning

    A detailed feature breakdown, a technical architecture plan and a written proposal covering scope, schedule and a fixed price.

    Step 2
  3. (03)

    AI-accelerated development

    Agile development with weekly demos and production-ready code backed by automated tests. Goal: a working MVP in about 4 weeks.

    Step 3
  4. (04)

    Launch & operation

    Production deployment with complete documentation and handover. 30 days of free bug fixing from final delivery; maintenance and further development by agreement.

    Step 4

Frequently asked questions about LLM fine-tuning

(01)When is fine-tuning better than RAG?
Fine-tuning pays off when a model needs to reliably hit a particular style, terminology or fixed output format, or when a smaller model should handle a task more cheaply. RAG is better for current, changing knowledge with source citations. Often, combining both approaches is the best solution.
(02)How much training data do we need?
Typically a few hundred to several thousand high-quality examples, depending on the task and method; quality and variety matter more than sheer volume. We help with curation, generate additional examples where needed and run a small experiment first to check whether the effort is worthwhile.
(03)How much does LLM fine-tuning cost?
Costs depend on the method, data volume, model size and evaluation effort; LoRA training on an open-source model with existing data is much smaller than a project with a newly created dataset. Fixed price after scoping, proposal within 48 hours. Compute costs for training and operation come on top.
(04)Which models are best suited to fine-tuning?
Open-source models such as Llama, Mistral, Qwen or DeepSeek offer full control and can be self-hosted. Commercial providers offer fine-tuning for selected models via their APIs, for example for GPT. In the analysis we recommend the right base model based on quality, cost, licence and data protection.
(05)Do we keep the rights to the fine-tuned model?
For the results created in the project, you receive the exclusive rights of use under section 5 of our terms. For open-source models their licence terms also apply, and for commercial APIs the provider's terms. We set out the rights to training data and model weights in writing before the project starts.
(06)How long does a fine-tuning project take?
A project typically takes 2–6 weeks including data preparation, training and evaluation. More complex projects with several iterations, a custom dataset or preference training usually need 6–12 weeks. The actual training is usually the shortest part; most of the time goes into data and testing.
(07)Can a fine-tuned model also use RAG?
Yes, and that is often the strongest combination: through fine-tuning the model learns terminology, tone and output formats, while RAG supplies current knowledge with sources at runtime. That keeps answers stylistically consistent and up to date without retraining the model every time knowledge changes.
(08)What happens when our data changes?
Training can be continued with new examples. We set up pipelines that collect and check new data and trigger retraining at regular intervals; every new model version goes through the same tests before it goes live. Rapidly changing knowledge, however, is better placed in a RAG system.
(09)Can we run the model on our own hardware?
For open-source models, yes: we run the model on your infrastructure or with a European provider, optimised with vLLM and quantisation. That gives you full data sovereignty and predictable costs. The hardware required depends on model size, load and response times, which we clarify in advance.
(10)How do we measure whether fine-tuning was successful?
With a test set of real tasks defined before training. We compare the adapted model with the base model and with good prompting, both automatically and through human review, for example on correctness, consistency, terminology and tone. We only recommend deployment if the gain is clear.
(04)

Technology stack for fine-tuning

(01)

AI & ML

Open-source LLMs (Llama, Mistral, Qwen, DeepSeek)Fine-tuning APIs of commercial providers (e.g. OpenAI GPT)Hugging Face Transformers & PEFTLoRA, QLoRA, DPOAxolotl & Unsloth (training)vLLM (inference)Evaluation with custom test sets
(02)

Web & Mobile

Next.js & ReactTypeScriptReact Native & ExpoTailwind CSSshadcn/uiVercel Edge Runtime
(03)

Backend & Data

Node.js & HonoPythonPostgreSQL & SupabaseConvex (Real-Time DB)RedistRPC & GraphQLOpenAPI
(04)

DevOps & Infrastructure

Vercel & AWSDocker & KubernetesCI/CD (GitHub Actions)OpenTelemetry & GrafanaLangfuse (LLM Monitoring)
(05)

Fine-tuning by industry

Legal

Models that handle legal terminology and citation styles reliably, for contract analysis, research summaries and draft briefs.

Healthcare & pharma

Language models for medical documentation and professional communication, precise in terminology and always reviewed by experts.

Finance

Models for risk reports, compliance texts and structured extraction from financial documents with fixed formats.

Technical documentation

Consistent manuals, data sheets and translations that follow your terminology and style guides.

Customer service

Answers in your brand voice, with fixed structures and correct classification of requests.

Insurance

Claims descriptions, classification and summaries that follow internal guidelines and terminology.

(06)

Fine-tuning: example projects

Examples we can build for you

Customer service

Classifying service requests

A small open-source model is trained to assign incoming requests to the right categories and extract structured fields, more cheaply than a large model.

Custom test set · Self-hosting · Goal: lower cost per request
Consulting

Report drafts in house style

A model learns the structure, terminology and tone of internal reports and produces drafts that experts only need to review and complete.

LoRA adapter · Human approval · Consistent style
Industry

Domain-specific extraction

An adapted model reads technical documents and extracts key figures into a fixed schema, combined with RAG for current product data.

Fixed output format · Combined with RAG · Regular retraining
(07)

LLM fine-tuning: consulting in Berlin

Founder AI-native since
2024
Email
info [at] contextstudios [dot] ai

Adapt language models to your business

Talk to us for 30 minutes about your data and goals, and we will tell you honestly whether fine-tuning, RAG or both is the right fit.