Skip to content

Tutorial: LLM API Gateway with LiteLLM

What you'll build

In this guide we will build one API endpoint that applications can use to talk to a language model running on two Verda Serverless Container deployments. The application sends a prompt to the gateway and the gateway chooses which deployment handles it.

An application calling a model directly needs the deployment's Endpoint Address and its inference key. As you add applications and deployments you also need a way to control who can call the model, how much traffic each application can send and what happens when an endpoint becomes unavailable. The gateway puts those controls in one place.

By the end of the guide:

  • Applications call one model name verda-qwen without choosing a GPU deployment.
  • Each application can have its own API key and request limit.
  • You can compare the tokens returned to the client with the usage stored in PostgreSQL.
  • Requests are distributed across two model endpoints and you test routing when one endpoint is stopped and restored.

We use LiteLLM because it understands OpenAI-compatible model requests and provides virtual keys, token accounting and model-aware routing in the same service. The GPU deployments generate the text. LiteLLM runs on a CPU instance and manages access to them.


Architecture

The request travels through two layers. LiteLLM checks the application's key and limits, then sends the request to one of the two Verda endpoints. PostgreSQL keeps the gateway's key records and usage history. In this walkthrough the gateway listens only on the CPU instance itself. The application and the gateway share that machine.


                         CPU instance
+--------------------------------------------------------------+
|                                                              |
|  Application --- localhost:4000 ---> LiteLLM ---> PostgreSQL |
|                                         |       keys + usage |
+-----------------------------------------|--------------------+
                                          |
                           +--------------+--------------+
             HTTPS + key A |                             | HTTPS + key B
                           v                             v
                 +------------------+          +------------------+
                 |    Endpoint A    |          |    Endpoint B    |
                 |    Serverless    |          |    Serverless    |
                 | Container (GPU)  |          | Container (GPU)  |
                 +------------------+          +------------------+

Both deployments serve Qwen/Qwen2.5-1.5B-Instruct. The client-facing name verda-qwen is an alias for that pool.

Component Runs on Purpose
LiteLLM CPU instance Client authentication, rate limits and model routing
PostgreSQL The same CPU instance in this walkthrough Persistent virtual-key records and usage logs
Model A and model B Two Serverless Container deployments Generate model responses

The gateway is a single point of failure. Two deployments protect against one endpoint going down but if the gateway instance stops every request fails.


Prerequisites

  • Have a Hugging Face User Access Token with READ permission. The deployments use it to download the model weights.
  • Have ready an Inference API Key from Credentials → Inference API Keys in the Verda console.

Step 1: Create the two model deployments

The gateway needs two deployments that serve the same model. The configuration in this guide uses Qwen/Qwen2.5-1.5B-Instruct, so create both deployments with the settings below. They differ only in their name.

In the Verda cloud console, go to Containers → New deployment and use these settings:

Setting Value
Deployment name llm-gateway-a for the first deployment, llm-gateway-b for the second
Compute type L40S, which this guide was tested on. General Compute (24 GB VRAM) also works: the vLLM quickstart runs the same model and tag on it
Container image docker.io/vllm/vllm-openai, with Public toggled on
Tag v0.29.0
Exposed HTTP port 8000
Healthcheck port 8000
Healthcheck path /health
Start Command On
CMD --model Qwen/Qwen2.5-1.5B-Instruct --gpu-memory-utilization 0.9 --model-loader-extra-config '{"enable_multithread_load": true}'
Environment variables HF_TOKEN, set to your Hugging Face User Access Token
Scaling Minimum number of replicas 1, so each deployment stays ready during the test

Wait until both deployments are healthy, then copy each one's Endpoint Address from its deployment page. You paste them in Step 3.

For an explanation of each setting and a test request, see the vLLM quickstart. To serve the model with a different framework, see the other Serverless Container tutorials, and keep the model the same on both deployments.


Step 2: Prepare the gateway instance

Log in to the Verda cloud console and create a CPU instance with these settings:

Setting Value Notes
Instance type CPU.4V.16G 4 vCPU and 16 GB RAM.
Image Ubuntu 26.04 CUDA 13.1 + Docker Docker and Docker Compose are preinstalled
SSH key Your registered key

This size handled the short verification workload in this guide. Measure your own request volume and response sizes before treating it as a production sizing recommendation.

Connect over SSH.

Create the working directory with a folder for test results and make your user its owner:

## New files are readable by you only
umask 077
## /opt belongs to root so creating the directory needs sudo
sudo mkdir -p /opt/llm-gateway/evidence
## Hand the directory to your user
sudo chown -R "$USER": /opt/llm-gateway
cd /opt/llm-gateway

This directory will hold the gateway credentials and test keys, which is why umask 077 keeps new files private. evidence/ collects the responses and summaries saved by the tests.

Check that the tools the guide uses are installed:

docker --version
docker compose version
python3 --version
curl --version

Each command should print a version number. If Docker or Compose is missing, complete Docker's Ubuntu installation before continuing.

Pull the gateway and database images

Download the LiteLLM and PostgreSQL 16 images:

sudo docker pull docker.litellm.ai/berriai/litellm:main-stable
sudo docker pull postgres:16

main-stable is LiteLLM's stable release channel, as used in its official quickstart Compose file. Both tags move to newer builds over time, so the image you pull may be newer than the one this guide was tested with. postgres:16 keeps PostgreSQL on major version 16, which matters because a database created by one major version cannot be opened by another.


Step 3: Set the Endpoint Addresses and credentials

LiteLLM needs to know where each model endpoint is and which inference key to use. It also needs a master key for administration, a salt key for encryption and a password for PostgreSQL. Docker Compose reads those values from a private .env file.

The setup script below prompts for the two Endpoint Addresses and Inference API Keys. It then generates the local credentials. It checks that the images were downloaded and refuses to replace an existing .env, so running it again cannot silently change the secrets used by your database.

Create configure.py in /opt/llm-gateway by copying the whole block below into the terminal. cat > ... <<'EOF' writes every line up to the closing EOF into the file and the quotes around 'EOF' keep the shell from changing any $ signs in the code. The other files in this guide are created the same way.

configure.py
# configure.py
cat > /opt/llm-gateway/configure.py <<'EOF'
"""Create private Compose settings for a fresh gateway. Never overwrite secrets."""
import getpass
import os
import re
import secrets
import subprocess
from pathlib import Path
from urllib.parse import urlsplit

IMAGES = {
    "LITELLM_IMAGE": "docker.litellm.ai/berriai/litellm:main-stable",
    "POSTGRES_IMAGE": "postgres:16",
}


def read_base(label):
    value = input(f"Endpoint {label}: paste its Endpoint Address followed by /v1: ").strip().rstrip("/")
    parsed = urlsplit(value)
    if (
        parsed.scheme != "https" or not parsed.hostname
        or parsed.username or parsed.password or parsed.query or parsed.fragment
        or not parsed.path.endswith("/v1")
        or any(c.isspace() or c in "'\"\\$" for c in value)
    ):
        raise SystemExit("Use the HTTPS Endpoint Address followed by /v1, without credentials, query, or fragment.")
    return value


def read_key(label):
    value = getpass.getpass(f"Endpoint {label}: paste its Inference API Key (hidden): ").strip()
    if not re.fullmatch(r"[A-Za-z0-9._~+/\-]+=*", value):
        raise SystemExit("Expected a non-empty Bearer token; paste the token only, without 'Bearer '.")
    return value


def main():
    os.umask(0o077)
    if Path(".env").exists():
        raise SystemExit(".env already exists. Keep it; this script does not rotate existing secrets.")
    for reference in IMAGES.values():
        subprocess.run(["docker", "image", "inspect", reference], check=True, stdout=subprocess.DEVNULL)
    values = dict(IMAGES)
    values.update(
        POSTGRES_PASSWORD=secrets.token_hex(24),
        LITELLM_MASTER_KEY="sk-" + secrets.token_hex(32),
        LITELLM_SALT_KEY="sk-" + secrets.token_hex(32),
    )
    for label in ("A", "B"):
        values[f"VERDA_{label}_API_BASE"] = read_base(label)
        values[f"VERDA_{label}_KEY"] = read_key(label)
    if values["VERDA_A_API_BASE"] == values["VERDA_B_API_BASE"]:
        raise SystemExit("A and B must be different Endpoint Addresses.")
    # Validated endpoint values are single-quoted for literal Compose dotenv parsing.
    lines = [f"{k}='{v}'" if k.startswith("VERDA_") else f"{k}={v}" for k, v in values.items()]
    with Path(".env").open("x") as handle:
        handle.write("\n".join(lines) + "\n")
    print("Created .env with private permissions. No keys were printed.")


if __name__ == "__main__":
    main()
EOF

Run it:

cd /opt/llm-gateway
sudo python3 configure.py

Example

Example: what the terminal asks

Endpoint A: paste its Endpoint Address followed by /v1: <endpoint-address-a>/v1
Endpoint A: paste its Inference API Key (hidden):
Endpoint B: paste its Endpoint Address followed by /v1: <endpoint-address-b>/v1
Endpoint B: paste its Inference API Key (hidden):
Created .env with private permissions. No keys were printed.

For security reasons, the key prompts stay blank as you paste.

ls -l .env

Check that .env was created before continuing. You should see one file with permissions -rw------- owned by root. If you see No such file or directory the script stopped before writing it. Scroll up to the message it printed, fix that problem, and run the script again.

These credentials have different jobs:

Credential Who uses it Why it exists
VERDA_A_KEY, VERDA_B_KEY LiteLLM The Inference API Keys you pasted for endpoints A and B. LiteLLM sends them to authenticate to each deployment
LITELLM_MASTER_KEY The operator Create and manage gateway keys
Virtual key (created later) An application Call the allowed model through the gateway
LITELLM_SALT_KEY LiteLLM Encrypt stored secrets
POSTGRES_PASSWORD LiteLLM and PostgreSQL Authenticate the database connection

Applications receive virtual keys. They do not need the Verda inference keys or the master key. Keep .env private, preserve the salt for the lifetime of the database and provision these values through your secret-management system for ongoing operation.

Warning

Editing the password or salt in .env is not a complete rotation procedure. See LiteLLM's database-backed setup.

Info

The gateway only needs Inference API Keys. Cloud API credentials used to manage Verda resources are not needed here. If a GPU container itself needs a secret such as a model-download token, set it on the deployment by following the Serverless Container secrets guide. Those secrets stay with the container and are not available on the gateway instance.


Step 4: Define the model pool and services

We give both deployments the same client-facing model name. That tells LiteLLM they are two places to serve the same request. Their deployment IDs remain different so that we can see where traffic went.

Create config.yaml in /opt/llm-gateway:

config.yaml
# config.yaml
cat > /opt/llm-gateway/config.yaml <<'EOF'
model_list:
  - model_name: verda-qwen
    litellm_params:
      model: openai/Qwen/Qwen2.5-1.5B-Instruct
      api_base: os.environ/VERDA_A_API_BASE
      api_key: os.environ/VERDA_A_KEY
    model_info:
      id: verda-a
      mode: chat
      input_cost_per_token: 0
      output_cost_per_token: 0
      health_check_max_tokens: 8
      health_check_timeout: 10

  - model_name: verda-qwen
    litellm_params:
      model: openai/Qwen/Qwen2.5-1.5B-Instruct
      api_base: os.environ/VERDA_B_API_BASE
      api_key: os.environ/VERDA_B_KEY
    model_info:
      id: verda-b
      mode: chat
      input_cost_per_token: 0
      output_cost_per_token: 0
      health_check_max_tokens: 8
      health_check_timeout: 10

router_settings:
  routing_strategy: simple-shuffle
  num_retries: 2
  timeout: 30

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL
  background_health_checks: true
  health_check_interval: 30
  enable_health_check_routing: true
EOF

model_name is the client-facing alias. The openai/ prefix selects LiteLLM's OpenAI-compatible adapter; the remainder must match the upstream model's served name. If you chose another model, change both upstream model entries. See the configuration reference.

simple-shuffle picks one of the eligible deployments at random for each request. Because the choice is random, a handful of requests can land unevenly on A and B; the split evens out over many requests. The model_info.id values identify which deployment served a request. See load balancing and response headers.

The input_cost_per_token: 0 and output_cost_per_token: 0 lines deliberately make this a token-accounting example: LiteLLM counts tokens but assigns them no price. They do not make GPU use free, calculate your Verda bill, or enforce monetary budgets. LiteLLM documents that setting both prices to zero bypasses its monetary budget enforcement. Configure an explicit pricing policy if you need monetary accounting. See custom pricing.

Define the containers

PostgreSQL makes virtual keys and usage records persistent. LiteLLM uses the database connection supplied by Compose and reads the model pool from config.yaml.

Create compose.yaml in /opt/llm-gateway:

compose.yaml
# compose.yaml
cat > /opt/llm-gateway/compose.yaml <<'EOF'
name: llm-gateway

services:
  db:
    image: ${POSTGRES_IMAGE:?Missing POSTGRES_IMAGE}
    restart: unless-stopped
    environment:
      POSTGRES_USER: litellm
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:?Missing database password}
      POSTGRES_DB: litellm
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U litellm -d litellm"]
      interval: 5s
      timeout: 5s
      retries: 20

  litellm:
    image: ${LITELLM_IMAGE:?Missing LITELLM_IMAGE}
    restart: unless-stopped
    env_file:
      - .env
    environment:
      DATABASE_URL: postgresql://litellm:${POSTGRES_PASSWORD}@db:5432/litellm
    volumes:
      - ./config.yaml:/app/config.yaml:ro
    command:
      - "--config"
      - "/app/config.yaml"
      - "--host"
      - "0.0.0.0"
      - "--port"
      - "4000"
      - "--num_workers"
      - "1"
    ports:
      - "127.0.0.1:4000:4000"
    depends_on:
      db:
        condition: service_healthy
    healthcheck:
      test:
        - CMD
        - python
        - -c
        - "import urllib.request; urllib.request.urlopen('http://127.0.0.1:4000/health/readiness', timeout=5)"
      interval: 10s
      timeout: 10s
      retries: 18
      start_period: 60s

volumes:
  postgres_data:
EOF

The database uses a named Docker volume, so replacing its container does not discard its data. That volume is local to this host; it is not a backup or a second copy on another machine.

The gateway publishes port 4000 on 127.0.0.1, so it is reachable only from the instance itself during this walkthrough. PostgreSQL has no published host port. The single worker keeps the first rate-limit test straightforward; shared state becomes necessary when adding workers or replicas.


Step 5: Start the gateway and check model health

Start the two services from /opt/llm-gateway. Compose waits for PostgreSQL to be healthy before starting LiteLLM:

sudo docker compose config --quiet
sudo docker compose up -d --wait --wait-timeout 240
sudo docker compose ps
curl --silent --show-error --fail-with-body \
  --max-time 15 http://127.0.0.1:4000/health/readiness

Both services should become healthy and readiness should return HTTP 200. config --quiet validates the files without printing resolved secrets.

Readiness checks the gateway and its database; it does not establish that either model endpoint is healthy. The authenticated /health route reports model health. With background checks enabled, it returns the latest cached probe results. See LiteLLM health checks.

Load the master key for the administrative commands that follow, without printing it:

LITELLM_MASTER_KEY=$(sudo sed -n 's/^LITELLM_MASTER_KEY=//p' .env)
: "${LITELLM_MASTER_KEY:?Master key is missing}"

Define this helper in the same Bash session:

check_gateway_health() {
  local label="$1"
  curl --silent --show-error --fail-with-body --max-time 60 \
    http://127.0.0.1:4000/health \
    --header "Authorization: Bearer ${LITELLM_MASTER_KEY:?Master key missing}" \
    --output "evidence/health-${label}.json" || return

  python3 - "evidence/health-${label}.json" <<'PY'
import json
import sys
with open(sys.argv[1]) as handle:
    data = json.load(handle)
for name in ("healthy_endpoints", "unhealthy_endpoints"):
    endpoints = data.get(name, [])
    print(f"{name}: {len(endpoints)}")
    for endpoint in endpoints:
        print(" ", endpoint.get("api_base", "(URL not reported)"))
PY
}

check_gateway_health baseline

Wait for two healthy endpoints and zero unhealthy endpoints before continuing. If the first background cycle has not completed, wait 30 seconds and call the helper again. A 200 status from /health alone is not enough; inspect its endpoint lists.

If you reconnect

LITELLM_MASTER_KEY, CLIENT_KEY (from Step 7) and check_gateway_health exist only in the current terminal session. After logging out or opening a new terminal, commands fail with HTTP 401 Malformed API Key. Reload the keys from their files:

cd /opt/llm-gateway
LITELLM_MASTER_KEY=$(sudo sed -n 's/^LITELLM_MASTER_KEY=//p' .env)
[ -f client-key.json ] && CLIENT_KEY=$(python3 -c 'import json; print(json.load(open("client-key.json"))["key"])')

Then paste the check_gateway_health block above again before running Step 10.


Step 6: Send the first request

We can now send a request through the gateway. The payload uses verda-qwen, so LiteLLM chooses the underlying deployment.

Create gateway-request.json in /opt/llm-gateway:

gateway-request.json
# gateway-request.json
cat > /opt/llm-gateway/gateway-request.json <<'EOF'
{
  "model": "verda-qwen",
  "messages": [{"role": "user", "content": "Reply with one short greeting."}],
  "max_tokens": 32,
  "stream": false
}
EOF

Send a request. For this first check, use the master key already loaded in the shell. We create an application key in the next step:

curl --silent --show-error --fail-with-body \
  --connect-timeout 5 --max-time 120 \
  http://127.0.0.1:4000/v1/chat/completions \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --header "Content-Type: application/json" \
  --data-binary @gateway-request.json \
  --dump-header evidence/gateway-response.headers \
  --output evidence/gateway-response.json \
  --write-out 'Gateway request: HTTP %{http_code}\n'

python3 -m json.tool evidence/gateway-response.json

Expect HTTP 200, an assistant message, and a usage object containing input and output token counts.

Check that unauthenticated requests are rejected:

curl --silent --show-error --max-time 30 \
  http://127.0.0.1:4000/v1/chat/completions \
  --header "Content-Type: application/json" \
  --data-binary @gateway-request.json \
  --output evidence/auth-missing-key.json \
  --write-out 'No API key: HTTP %{http_code}\n'

curl --silent --show-error --max-time 30 \
  http://127.0.0.1:4000/v1/chat/completions \
  --header "Authorization: Bearer sk-deliberately-invalid" \
  --header "Content-Type: application/json" \
  --data-binary @gateway-request.json \
  --output evidence/auth-invalid-key.json \
  --write-out 'Invalid API key: HTTP %{http_code}\n'

Both requests should return HTTP 401.


Step 7: Give an application its own key

A virtual key lets you control one application without sharing an upstream inference credential. Here we allow only verda-qwen and set a deliberately low limit of three requests per minute, so we can see the rejection happen.

Create the test key once. It expires after 24 hours, and its returned secret is saved in client-key.json. Running key generation again creates another key:

curl --silent --show-error --fail-with-body --max-time 30 \
  http://127.0.0.1:4000/key/generate \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --header "Content-Type: application/json" \
  --data '{"key_alias":"llm-gateway-rpm-test","models":["verda-qwen"],"rpm_limit":3,"duration":"24h"}' \
  --output client-key.json \
  --write-out 'Create client key: HTTP %{http_code}\n'

CLIENT_KEY=$(python3 -c 'import json; print(json.load(open("client-key.json"))["key"])')

Use virtual keys in applications. Administrative keys can bypass limits, so using the master key would not test the client policy. See virtual keys and rate limits.

Send six requests immediately with the new client key:

for i in 1 2 3 4 5 6; do
  curl --silent --show-error --connect-timeout 5 --max-time 30 \
    http://127.0.0.1:4000/v1/chat/completions \
    --header "Authorization: Bearer ${CLIENT_KEY:?Client key missing}" \
    --header "Content-Type: application/json" \
    --data-binary @gateway-request.json \
    --dump-header "evidence/rpm-${i}.headers" \
    --output "evidence/rpm-${i}.json" \
    --write-out "Request ${i}: HTTP %{http_code}, duration %{time_total}s\n"
done | tee evidence/rpm-summary.txt

Expect three HTTP 200 responses followed by HTTP 429 responses. The gateway accepted the first three calls and rejected the rest at the configured limit. Run the loop only once. If you need to repeat it, for example because the requests spanned a one-minute window boundary, first create a new key with a new key_alias such as llm-gateway-rpm-test-2, and use that alias in Step 8's query.

The reference stack uses one worker so its in-memory counters have one owner, and a restart resets that state. Add shared Redis coordination before adding workers or gateway replicas; PostgreSQL alone does not synchronize the rate limiter. See LiteLLM Redis requirements.


Step 8: Check the tokens used

Each successful response includes usage: tokens in the prompt, tokens in the completion, and their total. LiteLLM also records that usage in PostgreSQL, where it is associated with the application key.

We check the count from both sides: first what the client received, then what the gateway stored.

What the client received. Step 7 saved each response as evidence/rpm-1.json to evidence/rpm-6.json. The following script opens those six files, skips the rejected ones, and adds up the usage of the rest:

python3 - <<'PY'
import json
from pathlib import Path
names = ("prompt_tokens", "completion_tokens", "total_tokens")
totals = dict.fromkeys(names, 0)
successful = 0
for i in range(1, 7):
    data = json.loads(Path(f"evidence/rpm-{i}.json").read_text())
    if data.get("choices") and data.get("usage"):
        successful += 1
        for name in names:
            totals[name] += data["usage"].get(name, 0)
print("Responses with usage:", successful)
print(json.dumps(totals, indent=2))
PY

range(1, 7) means files 1 to 6, one per request sent in Step 7. If you sent a different number of requests, change 7 to that number plus one.

In the reference test, it printed:

Responses with usage: 3
{
  "prompt_tokens": 105,
  "completion_tokens": 30,
  "total_tokens": 135
}

What the gateway stored. Now query PostgreSQL for the usage LiteLLM recorded against this test key. Usage writes can be asynchronous, so give them a few seconds to finish first. The query matches requests to the key by its alias, without displaying the key itself:

sudo docker compose exec -T db psql -U litellm -d litellm -v ON_ERROR_STOP=1 <<'SQL'
SELECT
  COUNT(*) FILTER (WHERE s.total_tokens > 0) AS requests_with_usage,
  COALESCE(SUM(s.prompt_tokens), 0) AS prompt_tokens,
  COALESCE(SUM(s.completion_tokens), 0) AS completion_tokens,
  COALESCE(SUM(s.total_tokens), 0) AS total_tokens
FROM "LiteLLM_SpendLogs" AS s
JOIN "LiteLLM_VerificationToken" AS k ON s.api_key = k.token
WHERE k.key_alias = 'llm-gateway-rpm-test';
SQL

In the reference test, it returned:

 requests_with_usage | prompt_tokens | completion_tokens | total_tokens
---------------------+---------------+-------------------+--------------
                   3 |           105 |                30 |          135
(1 row)

The two totals match: three successful requests, 105 prompt tokens plus 30 completion tokens, 135 in total. The three rejected requests are stored as zero-token failure records, so they do not count towards requests_with_usage. Your numbers depend on the model, prompt and output, so compare your own two totals rather than expecting 135.

If you ran the Step 7 loop more than once with the same key, the database total includes every run while the client total covers only the last one, so the two won't match. This read-only query matches the tested schema; recheck it when upgrading LiteLLM.

Token counts describe model usage. They are separate from time-based infrastructure charges. For logging configuration and retention, see LiteLLM spend tracking. This walkthrough verifies non-streaming accounting; validate streaming separately if your application uses it.


Step 9: Send traffic to both deployments

One successful response shows that the gateway can reach a model. To see whether both endpoints serve traffic, we send 100 short requests with ten concurrent clients and count the deployment IDs in the response headers.

Create load-test.py in /opt/llm-gateway. It reads the same gateway-request.json and writes its results to evidence/:

load-test.py
# load-test.py
cat > /opt/llm-gateway/load-test.py <<'EOF'
import json
import math
import statistics
import sys
import time
from collections import Counter
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from urllib.error import HTTPError
from urllib.request import Request, urlopen

label = sys.argv[1] if len(sys.argv) > 1 else "baseline"
count = int(sys.argv[2]) if len(sys.argv) > 2 else 100
workers = int(sys.argv[3]) if len(sys.argv) > 3 else 10

key = json.loads(Path("load-key.json").read_text())["key"]
payload = Path("gateway-request.json").read_bytes()
stamp = time.strftime("%Y%m%dT%H%M%SZ", time.gmtime())
folder = Path("evidence") / f"{label}-{stamp}"
folder.mkdir(parents=True, exist_ok=True)

def send_request(number):
    started = time.perf_counter()
    row = {"number": number, "status": 0}
    request = Request(
        "http://127.0.0.1:4000/v1/chat/completions",
        data=payload,
        headers={
            "Authorization": f"Bearer {key}",
            "Content-Type": "application/json",
        },
    )
    try:
        with urlopen(request, timeout=120) as response:
            row["status"] = response.status
            row["endpoint"] = response.headers.get("x-litellm-model-id")
            row["call_id"] = response.headers.get("x-litellm-call-id")
            row["retries"] = response.headers.get("x-litellm-attempted-retries")
            body = json.loads(response.read())
            row["response_id"] = body.get("id")
            row["usage"] = body.get("usage", {})
            if not body.get("choices") or not row["usage"]:
                row["error"] = "Missing completion or usage"
    except HTTPError as error:
        row["status"] = error.code
        row["error"] = "HTTPError"
    except Exception as error:
        row["error"] = type(error).__name__
    row["seconds"] = round(time.perf_counter() - started, 6)
    return row

started = time.perf_counter()
with ThreadPoolExecutor(max_workers=workers) as pool:
    rows = list(pool.map(send_request, range(1, count + 1)))
elapsed = time.perf_counter() - started

good = [r for r in rows if r["status"] == 200 and "error" not in r]
latencies = sorted(r["seconds"] for r in good)
summary = {
    "label": label,
    "requests": count,
    "concurrency": workers,
    "successful": len(good),
    "failed": count - len(good),
    "http_statuses": dict(Counter(str(r["status"]) for r in rows)),
    "successful_requests_by_endpoint": dict(
        Counter(r.get("endpoint") or "unknown" for r in good)
    ),
    "elapsed_seconds": round(elapsed, 3),
    "successful_requests_per_second": round(len(good) / elapsed, 2),
    "p50_seconds": round(statistics.median(latencies), 4) if good else None,
    "p95_seconds": round(latencies[math.ceil(0.95 * len(latencies)) - 1], 4) if good else None,
    "tokens": {
        name: sum(r["usage"].get(name, 0) for r in good)
        for name in ("prompt_tokens", "completion_tokens", "total_tokens")
    },
}
(folder / "requests.json").write_text(json.dumps(rows, indent=2))
(folder / "summary.json").write_text(json.dumps(summary, indent=2))
print(json.dumps(summary, indent=2))
print(f"Evidence saved in: {folder}")
EOF

Run the test. The script reads load-key.json. Create that file by issuing a separate test key with enough allowance for this run, then start the test:

curl --silent --show-error --fail-with-body --max-time 30 \
  http://127.0.0.1:4000/key/generate \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --header "Content-Type: application/json" \
  --data '{"key_alias":"llm-gateway-load-test","models":["verda-qwen"],"rpm_limit":1000,"tpm_limit":100000,"duration":"24h"}' \
  --output load-key.json

python3 load-test.py baseline 100 10

The printed summary reports success counts, latency, token totals, and the number of requests served by each deployment. The timestamped directory holds both the summary and the individual results. Check for:

  • successful: 100 and failed: 0.
  • Both verda-a and verda-b in successful_requests_by_endpoint, adding up to 100.
  • No unknown deployment IDs, which would mean the routing evidence is incomplete.

The split will not be exactly 50/50 and changes from run to run; what matters is that both deployments receive requests. Latency depends on your model, GPU and prompt length, so use p50_seconds and p95_seconds as your own baseline for later comparison. This short test shows that traffic is distributed; it is not a throughput benchmark.

The load key also shows how to configure tpm_limit, a tokens-per-minute limit. Rejection at that threshold was not part of the reference test.


Step 10: Stop one endpoint and verify recovery

The next check answers a different question: can the same client-facing model keep serving after one endpoint is taken out of service?

First, distinguish the three health checks. A healthy gateway process and a healthy model endpoint are different things:

Check What it establishes
Verda container readiness A replica is ready for work within its deployment
LiteLLM /health/readiness The gateway and its database are ready
LiteLLM background model checks Each configured model endpoint can complete a small inference request

Here, the model probes run every 30 seconds with a ten-second timeout and an eight-token output cap. Health-based routing removes a failing deployment from eligibility and admits it again after a successful check. Both background_health_checks and enable_health_check_routing must be enabled. See health-check-driven routing.

Take A out of service

  1. Confirm both endpoints are healthy and keep deployment B running.
  2. In the Verda console, pause deployment A.
  3. Run check_gateway_health a-stopped. Repeat after another health cycle until A is unhealthy and B is healthy.
  4. Run the same load test:

    python3 load-test.py a-stopped 100 10
    

After health detection, expect 100 successful requests served by verda-b, with no successful requests assigned to verda-a.

Restore A

  1. Resume deployment A in the Verda console and wait for its replica to be ready.
  2. Run check_gateway_health a-restored until both endpoints are healthy again.
  3. Run:

    python3 load-test.py recovered 100 10
    

Expect successful traffic on both deployment IDs again. Keep all three summaries, baseline, A stopped and recovered, for comparison.

This sequence verifies routing after the gateway has detected an outage. It does not prove that requests already in flight survive, nor that the interval between failure and detection is error-free. Retries can extend latency and repeat upstream computation, so test that transition under your application's workload.

If every deployment is unhealthy, LiteLLM documents that it can bypass the health filter and keep attempting requests. A successful /health HTTP status therefore does not guarantee usable model capacity. Monitor the endpoint counts and real completion failures.

Info

The model probes send real inference traffic, so they may keep a deployment from scaling to zero. Check this against your scaling settings. For a latency-sensitive standby, keep capacity ready and account for its running cost.

Congratulations! You now have one model alias, per-application access control, persistent token records and a way to verify which endpoint handled each request.


Conclusion

You now have a LiteLLM gateway on a CPU instance that authenticates applications with their own keys, enforces request limits, records token usage in PostgreSQL and spreads traffic across two Serverless Container deployments with health-based failover.

To clean up. The test keys expire after 24 hours. Use the virtual-key management API to revoke them earlier, and keep any saved key files private. To stop the local services while keeping the database volume:

cd /opt/llm-gateway
sudo docker compose stop

Stopping these services does not stop the instance or the two GPU deployments, which keep billing while they exist. When you are finished, discontinue the instance and delete both container deployments. Keep the database and its secrets if you intend to resume the gateway.