# LLM App and RAG Hosting UK

Source: https://dijituldns.co.uk/llm-app-hosting/
Updated: 2026-10-08

> dijitul hosts LLM-powered applications and RAG apps for UK businesses, quoted to your workload. We run your app on a managed VPS or dedicated server with a vector database such as pgvector or Qdrant, Redis-backed queue workers for embedding and long AI calls, and API keys held server-side. GPU capacity is added only if you self-host models.

## Key facts

- Quoted per application after a short technical review
- Managed VPS or dedicated server sized for app, database and workers
- Vector database options: PostgreSQL with pgvector, Qdrant and others
- Queue workers for embeddings and long-running AI calls
- Streaming responses supported through Nginx and Cloudflare
- API keys in server environment variables only
- Python, Node.js and PHP stacks

## What an LLM app needs from hosting

An LLM app is ordinary software with one unusual dependency: calls to a language model that are slow, sometimes expensive, and occasionally fail. Good hosting allows for that:

- **An app server** running Python (FastAPI, Django, Flask), Node.js, or PHP (Laravel).
- **A database** for users, conversations and settings.
- **A queue and workers**, so long jobs run in the background rather than holding a web request open.
- **Streaming support**, so tokens reach the user as they're generated. That needs the right proxy buffering and timeout settings.
- **Secrets management** for model provider keys.
- **Logging** of usage and cost per user or feature.

## RAG: retrieval over your own content

Retrieval-augmented generation lets an AI answer from your documents: policies, manuals, product data, knowledge bases. The pipeline usually looks like this:

- Documents are split into chunks.
- Each chunk is turned into an embedding (a list of numbers) by an embedding model.
- Embeddings are stored in a **vector database**.
- When a user asks a question, the closest chunks are retrieved and sent to the LLM with the question.

For many projects, **PostgreSQL with pgvector** is enough and keeps everything in one database. Larger or search-heavy projects may suit a dedicated vector engine such as **Qdrant**. We size RAM for the index, because vector search is memory-hungry.

## Hosted model or self-hosted model

| | Hosted model API | Self-hosted open-weight model |
| Hardware | Normal VPS | GPU server |
| Cost pattern | Pay per token | Fixed server cost |
| Data | Prompts sent to the provider | Stays on your server |
| Model quality | Access to frontier models | Depends on model size you can fit |

Many apps mix both: a hosted API for answers and a small self-hosted embedding model. See [GPU servers](https://dijituldns.co.uk/gpu-servers/).

## Questions we'll ask

- What language and framework is the app written in?
- Which model providers or open-weight models will you use?
- How many documents for RAG, and how often do they change?
- Expected users and requests per day?
- Are there confidentiality rules on the data going into prompts?
- Do you need separate staging and production?

## Build it with us

dijitul's agency team at dijitul.uk builds custom software, including LLM and RAG applications. If you'd like us to build as well as host, we'll design both together. Call 01623 650333 or email info@dijitul.uk. Related: [managed VPS](https://dijituldns.co.uk/managed-vps/), [Python hosting](https://dijituldns.co.uk/python-django-hosting/), [SaaS hosting](https://dijituldns.co.uk/saas-hosting/).

## FAQs

### How do I host a RAG application?

You need an app server, a vector database such as pgvector or Qdrant, a worker to create embeddings, and access to an LLM via API or a GPU server. dijitul sets up and manages RAG hosting for UK businesses on a quoted basis.

### What is the best vector database for a small project?

For many small and medium projects, PostgreSQL with the pgvector extension is enough and keeps all your data in one database. Dedicated engines such as Qdrant suit larger collections or heavy search loads.

### Why does my AI app time out?

LLM calls can take tens of seconds, which can exceed PHP, proxy or load balancer timeouts. Use streaming responses or move the call into a background queue worker. dijitul configures timeouts and workers for LLM apps it hosts.

### Should I self-host an LLM or use an API?

Use an API for the best model quality and no hardware to manage. Self-host an open-weight model if data must stay on your server or you have high, steady volumes that make a fixed GPU cost worthwhile.

### Can I host a Python AI app with FastAPI?

Yes. A FastAPI app typically runs under Uvicorn or Gunicorn behind Nginx, managed by systemd or Docker, with Cloudflare in front. dijitul hosts Python AI apps on managed VPS and dedicated servers.

### How much RAM does a vector database need?

It depends on the number of vectors, their dimensions and the index type. Many vector indexes perform best when held in memory. dijitul estimates RAM from your document count and embedding model before quoting.

## Pricing and ordering

This is quoted to fit. Request a quote at https://dijituldns.co.uk/quote/ or call 01623 650333.
