RMM Labs
Integrations 3 min read 28 September 2026 By Fred

Run local models with Ollama, vLLM or any OpenAI-compatible server in Xcellerate AIG

Run models on your own hardware and put them behind the same governed gateway as the cloud vendors, with budgets, guardrails and logs.

With Xcellerate AIG you can put models you run yourself, on a laptop, a GPU server or a Kubernetes cluster, behind the same governed gateway as the cloud vendors. Prompts stay in your network, and budgets, guardrails and logs still apply. This guide is for admins at MSPs, SMEs and service businesses who want to keep AI local.

TL;DR

  • Ollama has its own provider type; vLLM, SGLang, Triton, KServe and similar servers use the OpenAI-compatible type.
  • You need a base URL ending in /v1, and usually no key.
  • Tick what the server really offers, list the model names in Allowed models and add your own price.

Before you start

  • A running server that AIG can reach over the network:
    • Ollama: the OpenAI surface under /v1, for example http://<your-host>:11434/v1.
    • vLLM, SGLang, Triton (OpenAI frontend) or KServe: a base URL ending in /v1, for example http://vllm:8000/v1.
  • If the server was started with an API key (for example vLLM with --api-key), that key. Otherwise nothing.
  • Recommended: start vLLM or SGLang with --served-model-name <name>, so the model ID is a readable name.
  • The Admin role in AIG, or a custom role with permission to create and edit providers.

Set up the local provider

1. Choose the type

Go to Connect → Providers and click Connect a provider (or Add provider). Choose Ollama or OpenAI-compatible.

2. Fill in the Base URL

For OpenAI-compatible, the Base URL is required: this type has no default endpoint. For Ollama, enter the /v1 URL of your own host.

3. Leave the key empty if the server has none

If your server has no key, leave the field empty. AIG then sends no Authorization header at all.

4. Narrow what the provider serves

Under What this provider serves, tick only what the server really offers, for example chat only. Other requests are then refused before they leave AIG.

5. Models and prices

List the served model names in Allowed models on the key. Add a pricing override, per token or per request: your own models are not in any public price list.

6. Test and enable

Click Test connection, then Enable provider. Callers use <provider-name>/<served-model-name>.

What happens after you connect

  • Ollama: chat and embeddings.
  • OpenAI-compatible: chat, embeddings, images, speech, transcription and rerank, as far as you tick them.
  • AIG has been tested with vLLM, SGLang, Triton's OpenAI frontend and the KServe Hugging Face runtime. Their error formats are turned into readable messages.
  • Traffic stays inside your network, while governance and logging apply as usual.

Good to know

  • Streamed requests are costed at zero unless the caller sends stream_options: {"include_usage": true}.
  • Triton reports no token usage, so price Triton models per request. Triton embeddings through the KServe v2 protocol are not reachable.
  • vLLM on a CPU needs --dtype float32.
  • Ollama's native (non-/v1) API is not supported. Always use the /v1 URL.
  • Without a licence, AIG runs in free mode: 5 virtual keys with 5 requests per key per day.

Troubleshooting

  • This provider type has no default endpoint — a base URL is required. Fill in the Base URL for the OpenAI-compatible type.
  • No model to try. Sync the catalog first… Your own model is not in the public catalogue. Add it under Allowed models and create a pricing override.
  • KServe validation errors show up as, for example, Field required: body.messages.
  • A request is refused straight away? The client asked for a capability you did not tick. The message names a provider that can serve it.

Talk to us about AIG

Want to bring local and cloud models under one set of rules? Talk to us about AIG. Read more about every provider in one catalogue and self-hosted AIG.

Frequently asked questions

Do my prompts leave the network?
No. Traffic runs between AIG and your own server, inside your network.
Do I need an API key?
Only if the server was started with one. Otherwise leave the field empty and AIG sends no Authorization header.
Which servers can I use?
Ollama, and any OpenAI-compatible server such as vLLM, SGLang, Triton's OpenAI frontend or KServe.
Sources: Verified against the Xcellerate AIG source code by the product team on 2026-09-28. Feature pages: https://rmmlabs.io/en/contact; https://rmmlabs.io/en/products/aig/features/providers; https://rmmlabs.io/en/products/aig/features/self-hosted.

Ready to solve time registration compliance?

Xcellerate OPS covers Belgian 2027 time registration requirements out of the box — no extra module needed.

Related articles