With Xcellerate AIG you can put models you run yourself, on a laptop, a GPU server or a Kubernetes cluster, behind the same governed gateway as the cloud vendors. Prompts stay in your network, and budgets, guardrails and logs still apply. This guide is for admins at MSPs, SMEs and service businesses who want to keep AI local.
TL;DR
- Ollama has its own provider type; vLLM, SGLang, Triton, KServe and similar servers use the OpenAI-compatible type.
- You need a base URL ending in
/v1, and usually no key.- Tick what the server really offers, list the model names in Allowed models and add your own price.
Before you start
- A running server that AIG can reach over the network:
- Ollama: the OpenAI surface under
/v1, for examplehttp://<your-host>:11434/v1. - vLLM, SGLang, Triton (OpenAI frontend) or KServe: a base URL ending in
/v1, for examplehttp://vllm:8000/v1.
- Ollama: the OpenAI surface under
- If the server was started with an API key (for example vLLM with
--api-key), that key. Otherwise nothing. - Recommended: start vLLM or SGLang with
--served-model-name <name>, so the model ID is a readable name. - The Admin role in AIG, or a custom role with permission to create and edit providers.
Set up the local provider
1. Choose the type
Go to Connect → Providers and click Connect a provider (or Add provider). Choose Ollama or OpenAI-compatible.
2. Fill in the Base URL
For OpenAI-compatible, the Base URL is required: this type has no default endpoint. For Ollama, enter the /v1 URL of your own host.
3. Leave the key empty if the server has none
If your server has no key, leave the field empty. AIG then sends no Authorization header at all.
4. Narrow what the provider serves
Under What this provider serves, tick only what the server really offers, for example chat only. Other requests are then refused before they leave AIG.
5. Models and prices
List the served model names in Allowed models on the key. Add a pricing override, per token or per request: your own models are not in any public price list.
6. Test and enable
Click Test connection, then Enable provider. Callers use <provider-name>/<served-model-name>.
What happens after you connect
- Ollama: chat and embeddings.
- OpenAI-compatible: chat, embeddings, images, speech, transcription and rerank, as far as you tick them.
- AIG has been tested with vLLM, SGLang, Triton's OpenAI frontend and the KServe Hugging Face runtime. Their error formats are turned into readable messages.
- Traffic stays inside your network, while governance and logging apply as usual.
Good to know
- Streamed requests are costed at zero unless the caller sends
stream_options: {"include_usage": true}. - Triton reports no token usage, so price Triton models per request. Triton embeddings through the KServe v2 protocol are not reachable.
- vLLM on a CPU needs
--dtype float32. - Ollama's native (non-
/v1) API is not supported. Always use the/v1URL. - Without a licence, AIG runs in free mode: 5 virtual keys with 5 requests per key per day.
Troubleshooting
This provider type has no default endpoint — a base URL is required.Fill in the Base URL for the OpenAI-compatible type.No model to try. Sync the catalog first…Your own model is not in the public catalogue. Add it under Allowed models and create a pricing override.- KServe validation errors show up as, for example,
Field required: body.messages. - A request is refused straight away? The client asked for a capability you did not tick. The message names a provider that can serve it.
Talk to us about AIG
Want to bring local and cloud models under one set of rules? Talk to us about AIG. Read more about every provider in one catalogue and self-hosted AIG.
