Tutorial: LLM API Gateway with LiteLLM¶
What you'll build¶
In this guide we will build one API endpoint that applications can use to talk to a language model running on two Verda Serverless Container deployments. The application sends a prompt to the gateway and the gateway chooses which deployment handles it.
An application calling a model directly needs the deployment's Endpoint Address and its inference key. As you add applications and deployments you also need a way to control who can call the model, how much traffic each application can send and what happens when an endpoint becomes unavailable. The gateway puts those controls in one place.
By the end of the guide:
- Applications call one model name
verda-qwenwithout choosing a GPU deployment. - Each application can have its own API key and request limit.
- You can compare the tokens returned to the client with the usage stored in PostgreSQL.
- Requests are distributed across two model endpoints and you test routing when one endpoint is stopped and restored.
We use LiteLLM because it understands OpenAI-compatible model requests and provides virtual keys, token accounting and model-aware routing in the same service. The GPU deployments generate the text. LiteLLM runs on a CPU instance and manages access to them.
Architecture¶
The request travels through two layers. LiteLLM checks the application's key and limits, then sends the request to one of the two Verda endpoints. PostgreSQL keeps the gateway's key records and usage history. In this walkthrough the gateway listens only on the CPU instance itself. The application and the gateway share that machine.
CPU instance
+--------------------------------------------------------------+
| |
| Application --- localhost:4000 ---> LiteLLM ---> PostgreSQL |
| | keys + usage |
+-----------------------------------------|--------------------+
|
+--------------+--------------+
HTTPS + key A | | HTTPS + key B
v v
+------------------+ +------------------+
| Endpoint A | | Endpoint B |
| Serverless | | Serverless |
| Container (GPU) | | Container (GPU) |
+------------------+ +------------------+
Both deployments serve Qwen/Qwen2.5-1.5B-Instruct. The client-facing name verda-qwen is an
alias for that pool.
| Component | Runs on | Purpose |
|---|---|---|
| LiteLLM | CPU instance | Client authentication, rate limits and model routing |
| PostgreSQL | The same CPU instance in this walkthrough | Persistent virtual-key records and usage logs |
| Model A and model B | Two Serverless Container deployments | Generate model responses |
The gateway is a single point of failure. Two deployments protect against one endpoint going down but if the gateway instance stops every request fails.
Prerequisites¶
- Have a Hugging Face User Access Token with
READpermission. The deployments use it to download the model weights. - Have ready an Inference API Key from Credentials → Inference API Keys in the Verda console.
Step 1: Create the two model deployments¶
The gateway needs two deployments that serve the same model. The configuration in this guide
uses Qwen/Qwen2.5-1.5B-Instruct, so create both deployments with the settings below. They
differ only in their name.
In the Verda cloud console, go to Containers → New deployment and use these settings:
| Setting | Value |
|---|---|
| Deployment name | llm-gateway-a for the first deployment, llm-gateway-b for the second |
| Compute type | L40S, which this guide was tested on. General Compute (24 GB VRAM) also works: the vLLM quickstart runs the same model and tag on it |
| Container image | docker.io/vllm/vllm-openai, with Public toggled on |
| Tag | v0.29.0 |
| Exposed HTTP port | 8000 |
| Healthcheck port | 8000 |
| Healthcheck path | /health |
| Start Command | On |
| CMD | --model Qwen/Qwen2.5-1.5B-Instruct --gpu-memory-utilization 0.9 --model-loader-extra-config '{"enable_multithread_load": true}' |
| Environment variables | HF_TOKEN, set to your Hugging Face User Access Token |
| Scaling | Minimum number of replicas 1, so each deployment stays ready during the test |
Wait until both deployments are healthy, then copy each one's Endpoint Address from its deployment page. You paste them in Step 3.
For an explanation of each setting and a test request, see the vLLM quickstart. To serve the model with a different framework, see the other Serverless Container tutorials, and keep the model the same on both deployments.
Step 2: Prepare the gateway instance¶
Log in to the Verda cloud console and create a CPU instance with these settings:
| Setting | Value | Notes |
|---|---|---|
| Instance type | CPU.4V.16G |
4 vCPU and 16 GB RAM. |
| Image | Ubuntu 26.04 CUDA 13.1 + Docker | Docker and Docker Compose are preinstalled |
| SSH key | Your registered key |
This size handled the short verification workload in this guide. Measure your own request volume and response sizes before treating it as a production sizing recommendation.
Connect over SSH.
Create the working directory with a folder for test results and make your user its owner:
## New files are readable by you only
umask 077
## /opt belongs to root so creating the directory needs sudo
sudo mkdir -p /opt/llm-gateway/evidence
## Hand the directory to your user
sudo chown -R "$USER": /opt/llm-gateway
cd /opt/llm-gateway
This directory will hold the gateway credentials and test keys, which is why umask 077
keeps new files private. evidence/ collects the responses and summaries saved by the tests.
Check that the tools the guide uses are installed:
Each command should print a version number. If Docker or Compose is missing, complete Docker's Ubuntu installation before continuing.
Pull the gateway and database images¶
Download the LiteLLM and PostgreSQL 16 images:
main-stable is LiteLLM's stable release channel, as used in its
official quickstart Compose file.
Both tags move to newer builds over time, so the image you pull may be newer than the one this
guide was tested with. postgres:16 keeps PostgreSQL on major version 16, which matters
because a database created by one major version cannot be opened by another.
Step 3: Set the Endpoint Addresses and credentials¶
LiteLLM needs to know where each model endpoint is and which inference key to use. It also
needs a master key for administration, a salt key for encryption and a password for
PostgreSQL. Docker Compose reads those values from a private .env file.
The setup script below prompts for the two Endpoint Addresses and Inference API Keys. It then generates
the local credentials. It checks that the images were downloaded and refuses to replace an
existing .env, so running it again cannot silently change the secrets used by your
database.
Create configure.py in /opt/llm-gateway by copying the whole block below into the
terminal. cat > ... <<'EOF' writes every line up to the closing EOF into the file and the
quotes around 'EOF' keep the shell from changing any $ signs in the code. The other files
in this guide are created the same way.
# configure.py
cat > /opt/llm-gateway/configure.py <<'EOF'
"""Create private Compose settings for a fresh gateway. Never overwrite secrets."""
import getpass
import os
import re
import secrets
import subprocess
from pathlib import Path
from urllib.parse import urlsplit
IMAGES = {
"LITELLM_IMAGE": "docker.litellm.ai/berriai/litellm:main-stable",
"POSTGRES_IMAGE": "postgres:16",
}
def read_base(label):
value = input(f"Endpoint {label}: paste its Endpoint Address followed by /v1: ").strip().rstrip("/")
parsed = urlsplit(value)
if (
parsed.scheme != "https" or not parsed.hostname
or parsed.username or parsed.password or parsed.query or parsed.fragment
or not parsed.path.endswith("/v1")
or any(c.isspace() or c in "'\"\\$" for c in value)
):
raise SystemExit("Use the HTTPS Endpoint Address followed by /v1, without credentials, query, or fragment.")
return value
def read_key(label):
value = getpass.getpass(f"Endpoint {label}: paste its Inference API Key (hidden): ").strip()
if not re.fullmatch(r"[A-Za-z0-9._~+/\-]+=*", value):
raise SystemExit("Expected a non-empty Bearer token; paste the token only, without 'Bearer '.")
return value
def main():
os.umask(0o077)
if Path(".env").exists():
raise SystemExit(".env already exists. Keep it; this script does not rotate existing secrets.")
for reference in IMAGES.values():
subprocess.run(["docker", "image", "inspect", reference], check=True, stdout=subprocess.DEVNULL)
values = dict(IMAGES)
values.update(
POSTGRES_PASSWORD=secrets.token_hex(24),
LITELLM_MASTER_KEY="sk-" + secrets.token_hex(32),
LITELLM_SALT_KEY="sk-" + secrets.token_hex(32),
)
for label in ("A", "B"):
values[f"VERDA_{label}_API_BASE"] = read_base(label)
values[f"VERDA_{label}_KEY"] = read_key(label)
if values["VERDA_A_API_BASE"] == values["VERDA_B_API_BASE"]:
raise SystemExit("A and B must be different Endpoint Addresses.")
# Validated endpoint values are single-quoted for literal Compose dotenv parsing.
lines = [f"{k}='{v}'" if k.startswith("VERDA_") else f"{k}={v}" for k, v in values.items()]
with Path(".env").open("x") as handle:
handle.write("\n".join(lines) + "\n")
print("Created .env with private permissions. No keys were printed.")
if __name__ == "__main__":
main()
EOF
Run it:
Example
Example: what the terminal asks
Endpoint A: paste its Endpoint Address followed by /v1: <endpoint-address-a>/v1
Endpoint A: paste its Inference API Key (hidden):
Endpoint B: paste its Endpoint Address followed by /v1: <endpoint-address-b>/v1
Endpoint B: paste its Inference API Key (hidden):
Created .env with private permissions. No keys were printed.
For security reasons, the key prompts stay blank as you paste.
Check that .env was created before continuing. You should see one file with permissions
-rw------- owned by root. If you see No such file or directory the script stopped before writing it.
Scroll up to the message it printed, fix that problem, and run the script again.
These credentials have different jobs:
| Credential | Who uses it | Why it exists |
|---|---|---|
VERDA_A_KEY, VERDA_B_KEY |
LiteLLM | The Inference API Keys you pasted for endpoints A and B. LiteLLM sends them to authenticate to each deployment |
LITELLM_MASTER_KEY |
The operator | Create and manage gateway keys |
| Virtual key (created later) | An application | Call the allowed model through the gateway |
LITELLM_SALT_KEY |
LiteLLM | Encrypt stored secrets |
POSTGRES_PASSWORD |
LiteLLM and PostgreSQL | Authenticate the database connection |
Applications receive virtual keys. They do not need the Verda inference keys or the
master key. Keep .env private, preserve the salt for the lifetime of the database and
provision these values through your secret-management system for ongoing operation.
Warning
Editing the password or salt in .env is not a complete rotation procedure. See LiteLLM's database-backed setup.
Info
The gateway only needs Inference API Keys. Cloud API credentials used to manage Verda resources are not needed here. If a GPU container itself needs a secret such as a model-download token, set it on the deployment by following the Serverless Container secrets guide. Those secrets stay with the container and are not available on the gateway instance.
Step 4: Define the model pool and services¶
We give both deployments the same client-facing model name. That tells LiteLLM they are two places to serve the same request. Their deployment IDs remain different so that we can see where traffic went.
Create config.yaml in /opt/llm-gateway:
# config.yaml
cat > /opt/llm-gateway/config.yaml <<'EOF'
model_list:
- model_name: verda-qwen
litellm_params:
model: openai/Qwen/Qwen2.5-1.5B-Instruct
api_base: os.environ/VERDA_A_API_BASE
api_key: os.environ/VERDA_A_KEY
model_info:
id: verda-a
mode: chat
input_cost_per_token: 0
output_cost_per_token: 0
health_check_max_tokens: 8
health_check_timeout: 10
- model_name: verda-qwen
litellm_params:
model: openai/Qwen/Qwen2.5-1.5B-Instruct
api_base: os.environ/VERDA_B_API_BASE
api_key: os.environ/VERDA_B_KEY
model_info:
id: verda-b
mode: chat
input_cost_per_token: 0
output_cost_per_token: 0
health_check_max_tokens: 8
health_check_timeout: 10
router_settings:
routing_strategy: simple-shuffle
num_retries: 2
timeout: 30
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: os.environ/DATABASE_URL
background_health_checks: true
health_check_interval: 30
enable_health_check_routing: true
EOF
model_name is the client-facing alias. The openai/ prefix selects LiteLLM's
OpenAI-compatible adapter; the remainder must match the upstream model's served name. If you
chose another model, change both upstream model entries. See the
configuration reference.
simple-shuffle picks one of the eligible deployments at random for each request. Because
the choice is random, a handful of requests can land unevenly on A and B; the split evens out
over many requests. The model_info.id values identify which deployment served a request. See
load balancing and
response headers.
The input_cost_per_token: 0 and output_cost_per_token: 0 lines deliberately make this a
token-accounting example: LiteLLM counts tokens but assigns them no price. They do not make
GPU use free, calculate your Verda bill, or enforce monetary budgets. LiteLLM documents that
setting both prices to zero bypasses its monetary budget enforcement. Configure an explicit
pricing policy if you need monetary accounting. See
custom pricing.
Define the containers¶
PostgreSQL makes virtual keys and usage records persistent. LiteLLM uses the database
connection supplied by Compose and reads the model pool from config.yaml.
Create compose.yaml in /opt/llm-gateway:
# compose.yaml
cat > /opt/llm-gateway/compose.yaml <<'EOF'
name: llm-gateway
services:
db:
image: ${POSTGRES_IMAGE:?Missing POSTGRES_IMAGE}
restart: unless-stopped
environment:
POSTGRES_USER: litellm
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?Missing database password}
POSTGRES_DB: litellm
volumes:
- postgres_data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U litellm -d litellm"]
interval: 5s
timeout: 5s
retries: 20
litellm:
image: ${LITELLM_IMAGE:?Missing LITELLM_IMAGE}
restart: unless-stopped
env_file:
- .env
environment:
DATABASE_URL: postgresql://litellm:${POSTGRES_PASSWORD}@db:5432/litellm
volumes:
- ./config.yaml:/app/config.yaml:ro
command:
- "--config"
- "/app/config.yaml"
- "--host"
- "0.0.0.0"
- "--port"
- "4000"
- "--num_workers"
- "1"
ports:
- "127.0.0.1:4000:4000"
depends_on:
db:
condition: service_healthy
healthcheck:
test:
- CMD
- python
- -c
- "import urllib.request; urllib.request.urlopen('http://127.0.0.1:4000/health/readiness', timeout=5)"
interval: 10s
timeout: 10s
retries: 18
start_period: 60s
volumes:
postgres_data:
EOF
The database uses a named Docker volume, so replacing its container does not discard its data. That volume is local to this host; it is not a backup or a second copy on another machine.
The gateway publishes port 4000 on 127.0.0.1, so it is reachable only from the instance
itself during this walkthrough. PostgreSQL has no published host port. The single worker keeps
the first rate-limit test straightforward; shared state becomes necessary when adding workers
or replicas.
Step 5: Start the gateway and check model health¶
Start the two services from /opt/llm-gateway. Compose waits for PostgreSQL to be healthy
before starting LiteLLM:
sudo docker compose config --quiet
sudo docker compose up -d --wait --wait-timeout 240
sudo docker compose ps
curl --silent --show-error --fail-with-body \
--max-time 15 http://127.0.0.1:4000/health/readiness
Both services should become healthy and readiness should return HTTP 200. config --quiet
validates the files without printing resolved secrets.
Readiness checks the gateway and its database; it does not establish that either model
endpoint is healthy. The authenticated /health route reports model health. With background
checks enabled, it returns the latest cached probe results. See
LiteLLM health checks.
Load the master key for the administrative commands that follow, without printing it:
LITELLM_MASTER_KEY=$(sudo sed -n 's/^LITELLM_MASTER_KEY=//p' .env)
: "${LITELLM_MASTER_KEY:?Master key is missing}"
Define this helper in the same Bash session:
check_gateway_health() {
local label="$1"
curl --silent --show-error --fail-with-body --max-time 60 \
http://127.0.0.1:4000/health \
--header "Authorization: Bearer ${LITELLM_MASTER_KEY:?Master key missing}" \
--output "evidence/health-${label}.json" || return
python3 - "evidence/health-${label}.json" <<'PY'
import json
import sys
with open(sys.argv[1]) as handle:
data = json.load(handle)
for name in ("healthy_endpoints", "unhealthy_endpoints"):
endpoints = data.get(name, [])
print(f"{name}: {len(endpoints)}")
for endpoint in endpoints:
print(" ", endpoint.get("api_base", "(URL not reported)"))
PY
}
check_gateway_health baseline
Wait for two healthy endpoints and zero unhealthy endpoints before continuing. If the
first background cycle has not completed, wait 30 seconds and call the helper again. A 200
status from /health alone is not enough; inspect its endpoint lists.
If you reconnect
LITELLM_MASTER_KEY, CLIENT_KEY (from Step 7) and check_gateway_health exist only in the current terminal session. After logging out or opening a new terminal, commands fail with HTTP 401 Malformed API Key. Reload the keys from their files:
cd /opt/llm-gateway
LITELLM_MASTER_KEY=$(sudo sed -n 's/^LITELLM_MASTER_KEY=//p' .env)
[ -f client-key.json ] && CLIENT_KEY=$(python3 -c 'import json; print(json.load(open("client-key.json"))["key"])')
Then paste the check_gateway_health block above again before running Step 10.
Step 6: Send the first request¶
We can now send a request through the gateway. The payload uses verda-qwen, so LiteLLM
chooses the underlying deployment.
Create gateway-request.json in /opt/llm-gateway:
# gateway-request.json
cat > /opt/llm-gateway/gateway-request.json <<'EOF'
{
"model": "verda-qwen",
"messages": [{"role": "user", "content": "Reply with one short greeting."}],
"max_tokens": 32,
"stream": false
}
EOF
Send a request. For this first check, use the master key already loaded in the shell. We create an application key in the next step:
curl --silent --show-error --fail-with-body \
--connect-timeout 5 --max-time 120 \
http://127.0.0.1:4000/v1/chat/completions \
--header "Authorization: Bearer $LITELLM_MASTER_KEY" \
--header "Content-Type: application/json" \
--data-binary @gateway-request.json \
--dump-header evidence/gateway-response.headers \
--output evidence/gateway-response.json \
--write-out 'Gateway request: HTTP %{http_code}\n'
python3 -m json.tool evidence/gateway-response.json
Expect HTTP 200, an assistant message, and a usage object containing input and output token
counts.
Check that unauthenticated requests are rejected:
curl --silent --show-error --max-time 30 \
http://127.0.0.1:4000/v1/chat/completions \
--header "Content-Type: application/json" \
--data-binary @gateway-request.json \
--output evidence/auth-missing-key.json \
--write-out 'No API key: HTTP %{http_code}\n'
curl --silent --show-error --max-time 30 \
http://127.0.0.1:4000/v1/chat/completions \
--header "Authorization: Bearer sk-deliberately-invalid" \
--header "Content-Type: application/json" \
--data-binary @gateway-request.json \
--output evidence/auth-invalid-key.json \
--write-out 'Invalid API key: HTTP %{http_code}\n'
Both requests should return HTTP 401.
Step 7: Give an application its own key¶
A virtual key lets you control one application without sharing an upstream inference
credential. Here we allow only verda-qwen and set a deliberately low limit of three requests
per minute, so we can see the rejection happen.
Create the test key once. It expires after 24 hours, and its returned secret is saved in
client-key.json. Running key generation again creates another key:
curl --silent --show-error --fail-with-body --max-time 30 \
http://127.0.0.1:4000/key/generate \
--header "Authorization: Bearer $LITELLM_MASTER_KEY" \
--header "Content-Type: application/json" \
--data '{"key_alias":"llm-gateway-rpm-test","models":["verda-qwen"],"rpm_limit":3,"duration":"24h"}' \
--output client-key.json \
--write-out 'Create client key: HTTP %{http_code}\n'
CLIENT_KEY=$(python3 -c 'import json; print(json.load(open("client-key.json"))["key"])')
Use virtual keys in applications. Administrative keys can bypass limits, so using the master key would not test the client policy. See virtual keys and rate limits.
Send six requests immediately with the new client key:
for i in 1 2 3 4 5 6; do
curl --silent --show-error --connect-timeout 5 --max-time 30 \
http://127.0.0.1:4000/v1/chat/completions \
--header "Authorization: Bearer ${CLIENT_KEY:?Client key missing}" \
--header "Content-Type: application/json" \
--data-binary @gateway-request.json \
--dump-header "evidence/rpm-${i}.headers" \
--output "evidence/rpm-${i}.json" \
--write-out "Request ${i}: HTTP %{http_code}, duration %{time_total}s\n"
done | tee evidence/rpm-summary.txt
Expect three HTTP 200 responses followed by HTTP 429 responses. The gateway accepted the
first three calls and rejected the rest at the configured limit. Run the loop only once. If you
need to repeat it, for example because the requests spanned a one-minute window boundary,
first create a new key with a new key_alias such as llm-gateway-rpm-test-2, and use that
alias in Step 8's query.
The reference stack uses one worker so its in-memory counters have one owner, and a restart resets that state. Add shared Redis coordination before adding workers or gateway replicas; PostgreSQL alone does not synchronize the rate limiter. See LiteLLM Redis requirements.
Step 8: Check the tokens used¶
Each successful response includes usage: tokens in the prompt, tokens in the completion,
and their total. LiteLLM also records that usage in PostgreSQL, where it is associated with
the application key.
We check the count from both sides: first what the client received, then what the gateway stored.
What the client received. Step 7 saved each response as evidence/rpm-1.json to
evidence/rpm-6.json. The following script opens those six files, skips the rejected ones,
and adds up the usage of the rest:
python3 - <<'PY'
import json
from pathlib import Path
names = ("prompt_tokens", "completion_tokens", "total_tokens")
totals = dict.fromkeys(names, 0)
successful = 0
for i in range(1, 7):
data = json.loads(Path(f"evidence/rpm-{i}.json").read_text())
if data.get("choices") and data.get("usage"):
successful += 1
for name in names:
totals[name] += data["usage"].get(name, 0)
print("Responses with usage:", successful)
print(json.dumps(totals, indent=2))
PY
range(1, 7) means files 1 to 6, one per request sent in Step 7. If you sent a different
number of requests, change 7 to that number plus one.
In the reference test, it printed:
What the gateway stored. Now query PostgreSQL for the usage LiteLLM recorded against this test key. Usage writes can be asynchronous, so give them a few seconds to finish first. The query matches requests to the key by its alias, without displaying the key itself:
sudo docker compose exec -T db psql -U litellm -d litellm -v ON_ERROR_STOP=1 <<'SQL'
SELECT
COUNT(*) FILTER (WHERE s.total_tokens > 0) AS requests_with_usage,
COALESCE(SUM(s.prompt_tokens), 0) AS prompt_tokens,
COALESCE(SUM(s.completion_tokens), 0) AS completion_tokens,
COALESCE(SUM(s.total_tokens), 0) AS total_tokens
FROM "LiteLLM_SpendLogs" AS s
JOIN "LiteLLM_VerificationToken" AS k ON s.api_key = k.token
WHERE k.key_alias = 'llm-gateway-rpm-test';
SQL
In the reference test, it returned:
requests_with_usage | prompt_tokens | completion_tokens | total_tokens
---------------------+---------------+-------------------+--------------
3 | 105 | 30 | 135
(1 row)
The two totals match: three successful requests, 105 prompt tokens plus 30 completion
tokens, 135 in total. The three rejected requests are stored as zero-token failure records, so
they do not count towards requests_with_usage. Your numbers depend on the model, prompt and
output, so compare your own two totals rather than expecting 135.
If you ran the Step 7 loop more than once with the same key, the database total includes every run while the client total covers only the last one, so the two won't match. This read-only query matches the tested schema; recheck it when upgrading LiteLLM.
Token counts describe model usage. They are separate from time-based infrastructure charges. For logging configuration and retention, see LiteLLM spend tracking. This walkthrough verifies non-streaming accounting; validate streaming separately if your application uses it.
Step 9: Send traffic to both deployments¶
One successful response shows that the gateway can reach a model. To see whether both endpoints serve traffic, we send 100 short requests with ten concurrent clients and count the deployment IDs in the response headers.
Create load-test.py in /opt/llm-gateway. It reads the same gateway-request.json and
writes its results to evidence/:
# load-test.py
cat > /opt/llm-gateway/load-test.py <<'EOF'
import json
import math
import statistics
import sys
import time
from collections import Counter
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from urllib.error import HTTPError
from urllib.request import Request, urlopen
label = sys.argv[1] if len(sys.argv) > 1 else "baseline"
count = int(sys.argv[2]) if len(sys.argv) > 2 else 100
workers = int(sys.argv[3]) if len(sys.argv) > 3 else 10
key = json.loads(Path("load-key.json").read_text())["key"]
payload = Path("gateway-request.json").read_bytes()
stamp = time.strftime("%Y%m%dT%H%M%SZ", time.gmtime())
folder = Path("evidence") / f"{label}-{stamp}"
folder.mkdir(parents=True, exist_ok=True)
def send_request(number):
started = time.perf_counter()
row = {"number": number, "status": 0}
request = Request(
"http://127.0.0.1:4000/v1/chat/completions",
data=payload,
headers={
"Authorization": f"Bearer {key}",
"Content-Type": "application/json",
},
)
try:
with urlopen(request, timeout=120) as response:
row["status"] = response.status
row["endpoint"] = response.headers.get("x-litellm-model-id")
row["call_id"] = response.headers.get("x-litellm-call-id")
row["retries"] = response.headers.get("x-litellm-attempted-retries")
body = json.loads(response.read())
row["response_id"] = body.get("id")
row["usage"] = body.get("usage", {})
if not body.get("choices") or not row["usage"]:
row["error"] = "Missing completion or usage"
except HTTPError as error:
row["status"] = error.code
row["error"] = "HTTPError"
except Exception as error:
row["error"] = type(error).__name__
row["seconds"] = round(time.perf_counter() - started, 6)
return row
started = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
rows = list(pool.map(send_request, range(1, count + 1)))
elapsed = time.perf_counter() - started
good = [r for r in rows if r["status"] == 200 and "error" not in r]
latencies = sorted(r["seconds"] for r in good)
summary = {
"label": label,
"requests": count,
"concurrency": workers,
"successful": len(good),
"failed": count - len(good),
"http_statuses": dict(Counter(str(r["status"]) for r in rows)),
"successful_requests_by_endpoint": dict(
Counter(r.get("endpoint") or "unknown" for r in good)
),
"elapsed_seconds": round(elapsed, 3),
"successful_requests_per_second": round(len(good) / elapsed, 2),
"p50_seconds": round(statistics.median(latencies), 4) if good else None,
"p95_seconds": round(latencies[math.ceil(0.95 * len(latencies)) - 1], 4) if good else None,
"tokens": {
name: sum(r["usage"].get(name, 0) for r in good)
for name in ("prompt_tokens", "completion_tokens", "total_tokens")
},
}
(folder / "requests.json").write_text(json.dumps(rows, indent=2))
(folder / "summary.json").write_text(json.dumps(summary, indent=2))
print(json.dumps(summary, indent=2))
print(f"Evidence saved in: {folder}")
EOF
Run the test. The script reads load-key.json. Create that file by issuing a separate test key with enough
allowance for this run, then start the test:
curl --silent --show-error --fail-with-body --max-time 30 \
http://127.0.0.1:4000/key/generate \
--header "Authorization: Bearer $LITELLM_MASTER_KEY" \
--header "Content-Type: application/json" \
--data '{"key_alias":"llm-gateway-load-test","models":["verda-qwen"],"rpm_limit":1000,"tpm_limit":100000,"duration":"24h"}' \
--output load-key.json
python3 load-test.py baseline 100 10
The printed summary reports success counts, latency, token totals, and the number of requests served by each deployment. The timestamped directory holds both the summary and the individual results. Check for:
successful: 100andfailed: 0.- Both
verda-aandverda-binsuccessful_requests_by_endpoint, adding up to 100. - No
unknowndeployment IDs, which would mean the routing evidence is incomplete.
The split will not be exactly 50/50 and changes from run to run; what matters is that both
deployments receive requests. Latency depends on your model, GPU and prompt length, so use
p50_seconds and p95_seconds as your own baseline for later comparison. This short test
shows that traffic is distributed; it is not a throughput benchmark.
The load key also shows how to configure tpm_limit, a tokens-per-minute limit. Rejection at
that threshold was not part of the reference test.
Step 10: Stop one endpoint and verify recovery¶
The next check answers a different question: can the same client-facing model keep serving after one endpoint is taken out of service?
First, distinguish the three health checks. A healthy gateway process and a healthy model endpoint are different things:
| Check | What it establishes |
|---|---|
| Verda container readiness | A replica is ready for work within its deployment |
LiteLLM /health/readiness |
The gateway and its database are ready |
| LiteLLM background model checks | Each configured model endpoint can complete a small inference request |
Here, the model probes run every 30 seconds with a ten-second timeout and an eight-token
output cap. Health-based routing removes a failing deployment from eligibility and admits it
again after a successful check. Both background_health_checks and
enable_health_check_routing must be enabled. See
health-check-driven routing.
Take A out of service¶
- Confirm both endpoints are healthy and keep deployment B running.
- In the Verda console, pause deployment A.
- Run
check_gateway_health a-stopped. Repeat after another health cycle until A is unhealthy and B is healthy. -
Run the same load test:
After health detection, expect 100 successful requests served by verda-b, with no successful
requests assigned to verda-a.
Restore A¶
- Resume deployment A in the Verda console and wait for its replica to be ready.
- Run
check_gateway_health a-restoreduntil both endpoints are healthy again. -
Run:
Expect successful traffic on both deployment IDs again. Keep all three summaries, baseline, A stopped and recovered, for comparison.
This sequence verifies routing after the gateway has detected an outage. It does not prove that requests already in flight survive, nor that the interval between failure and detection is error-free. Retries can extend latency and repeat upstream computation, so test that transition under your application's workload.
If every deployment is unhealthy, LiteLLM documents that it can bypass the health filter and
keep attempting requests. A successful /health HTTP status therefore does not guarantee
usable model capacity. Monitor the endpoint counts and real completion failures.
Info
The model probes send real inference traffic, so they may keep a deployment from scaling to zero. Check this against your scaling settings. For a latency-sensitive standby, keep capacity ready and account for its running cost.
Congratulations! You now have one model alias, per-application access control, persistent token records and a way to verify which endpoint handled each request.
Conclusion¶
You now have a LiteLLM gateway on a CPU instance that authenticates applications with their own keys, enforces request limits, records token usage in PostgreSQL and spreads traffic across two Serverless Container deployments with health-based failover.
To clean up. The test keys expire after 24 hours. Use the virtual-key management API to revoke them earlier, and keep any saved key files private. To stop the local services while keeping the database volume:
Stopping these services does not stop the instance or the two GPU deployments, which keep billing while they exist. When you are finished, discontinue the instance and delete both container deployments. Keep the database and its secrets if you intend to resume the gateway.
Related guides¶
- Deploy with vLLM: create the model endpoints used by the gateway.
- Serverless Container tutorials: other inference deployment examples.
- Scaling and health checks: replica readiness, queueing and scaling behaviour.
- API credentials: create the Inference API Keys for the upstream endpoints.
- Serverless Container secrets: manage secrets used by GPU containers.
- LiteLLM production deployment: expand the reference gateway for your availability requirements.