Tutorial: Retrieval-Augmented Generation (RAG) with pgvector¶
What you'll build¶
In this guide we will build a retrieval-augmented generation (RAG) service: a language model that answers questions using only passages retrieved from your own documents, with a citation on every claim.
- You ask "Am I charged for a volume when the instance is shut down?"
- You get a written answer with [1], [2] after each claim and a Sources list naming the page and section each one came from, so a reader can check it.
Where the Vector DB guide returned the matching rows, this guide reads them and answers.
Any self-hosted vector database works for this: Qdrant, Weaviate, Milvus or pgvector. This guide uses pgvector, and the database and embedding deployment are the ones from the Vector DB guide.
Architecture¶
- Your documents live in Object Storage and are uploaded with the Verda CLI.
- Passages are turned into vectors by an embedding model on a Serverless Container.
- Vectors are stored and searched in Postgres + pgvector on a CPU instance with a Block Volume.
- Answers come from an LLM on a second Serverless Container, served by vLLM.
Two flows connect them. They share the database and the embedding deployment, but run at different times.
Ingestion runs whenever your documents change:
+--------------------+
| Object Storage | your documents, as Markdown
+---------+----------+
| 1. read
v
+--------------------+ 2. embed +----------------------+
| ingest.py | ------------> | embedding model |
| chunk > embed | <------------ | |
+---------+----------+ vectors +----------------------+
| 3. store GPU, Serverless Container
v
+----------------------------------+
| Postgres + pgvector | CPU instance + Block Volume
+----------------------------------+
Query runs once per question:
question
|
v
+--------------------+ 1. embed +----------------------+
| | ------------> | embedding model |
| | <------------ | |
| | vector +----------------------+
| |
| query.py | 2. search +----------------------+
| | ------------> | Postgres + pgvector |
| | <------------ | |
| | top-k chunks +----------------------+
| |
| | 3. prompt +----------------------+
| | ------------> | LLM (vLLM) |
| | <------------ | |
+---------+----------+ answer +----------------------+
| GPU, Serverless Container
v
answer + citations
ingest.py and query.py are the two scripts you write, and both run on the CPU instance.
The LLM is only called at query time.
| Component | Tier | Runs when |
|---|---|---|
| Database | CPU.4V.16G instance | Continuously, while the instance is on |
| Data volume | NVMe Block Volume | Always, whether or not the instance runs |
| Embeddings | Serverless Container, L40S | On demand, scales to zero when idle |
| LLM | Serverless Container, L40S | On demand, scales to zero when idle |
| Documents | Object Storage | Always, billed on stored data |
Ingestion and querying must use the same embedding deployment. Vectors from different models cannot be compared, and nothing reports the mismatch. Both scripts import one embedding function for that reason.
Prerequisites¶
- Complete Steps 1, 2 and 4 of the Vector DB guide, so you have
the CPU instance with its Block Volume, Postgres with pgvector 0.8 or newer, and the
embedderdeployment running. Its schema (Step 3) is not needed; this guide creates its own. - Create an Object Storage access key.
- Have ready your Inference API Key and Hugging Face token from the Vector DB guide.
- Install the Verda CLI on the CPU instance from the Vector DB
guide, and authenticate it with your Cloud API Client ID and Secret. Test it with
verda locations; if it lists datacenters, the install and login are working.
Everything here is done in the Verda console, with the Verda CLI and over SSH.
No documents are needed to start: Step 1 creates a small sample corpus. Swap in your own Markdown files later; the pipeline does not change.
Step 1: Upload your documents¶
SSH into the CPU instance from the Vector DB guide.
Create the sample corpus¶
The corpus is a folder of Markdown files. To have something to index, create three short pages about Verda storage, billing and containers. In your home directory, create the folder:
Then create the three files. Each cat command writes one file; paste it whole, including
the closing EOF line:
cat > corpus/block-volumes.md <<'EOF'
# Block Volumes
A Block Volume is NVMe storage that exists independently of any instance, like an external
drive. You attach it to an instance to use it, detach it, and attach it to another instance.
## Attaching and resizing
A volume can be added when an instance is created or attached to an existing instance. It can
be expanded later without interrupting the instance. Volumes cannot be shrunk.
## Cloning a volume
A volume can be cloned within a datacenter or to another one, for backups or to seed a new
instance. The volume must be detached, or its instance shut down, while the clone is made.
## Deleting a volume
A deleted volume goes to the trash for 96 hours and can be restored during that time. To free
the storage quota and stop charges, permanently delete it from the Deleted volumes section.
Permanent deletion cannot be undone.
EOF
cat > corpus/pricing-and-billing.md <<'EOF'
# Pricing and billing
Verda bills Pay As You Go in pre-paid 10-minute increments. If a resource is terminated before
the billed period ends, the unused portion is refunded in the next billing period.
## Storage charges
Storage is billed for as long as it exists, per GiB per month. A Block Volume is charged whether
or not it is attached to an instance, and whether or not that instance is running. Charges stop
only when the volume is permanently deleted. Restoring a volume from the trash costs the Pay As
You Go price for the time it was deleted.
## Shutting down and deleting instances
Shutting down an instance pauses it so that volumes can be attached or detached. A shut-down
instance continues to charge your account. Deleting an instance removes it together with the
volumes you choose to delete; any volume you keep continues to be billed and can be attached to
a new instance.
EOF
cat > corpus/serverless-containers.md <<'EOF'
# Serverless Containers
Serverless Containers run any container image on GPU or CPU tiers and are billed only while
replicas are running.
## Scaling
Autoscaling applies whenever the maximum number of replicas is set higher than the minimum. By
default it scales on queue load, the queue length divided by the number of replicas. Scale-up
and scale-down delays control how long a threshold must hold before replicas change. With a
minimum of zero replicas, a deployment scales to zero when idle and costs nothing until the next
request, which then waits for a cold start.
## Health checks
Traffic is routed only to replicas that report healthy: an HTTP 200 from the healthcheck path.
During scale-down a replica receives SIGTERM and has 30 seconds to finish before SIGKILL.
EOF
Check that all three files are there:
Configure the CLI for Object Storage¶
Verda Object Storage is S3-compatible. We will manage it with the
Verda CLI command verda object-storage, or its short form
verda oss. Object storage uses its own key, separate from the Cloud API credentials the CLI
already has. Create one and add it to the same profile with the following command:
Name the profile default, then enter the access key and the secret. Accept the endpoint
(https://objects.fin-03.verda.storage) and region (us-east-1) with Enter.
Check that the key works:
An empty list is the expected result, since you have no buckets yet. AUTH_ERROR means the
key was not accepted or is not in the active profile; verda auth show tells you which
profile that is.
Create the bucket and upload¶
Bucket names are shared by everyone using the storage service, so pick one that is yours:
rag-corpus- followed by your name or team, in lowercase. Create the bucket:
You can also create it in the console instead; the upload below still uses the CLI.
Sync the folder into the bucket under a raw/ prefix, then list it:
verda oss sync ./corpus/ s3://rag-corpus-<name>/raw/
verda oss ls s3://rag-corpus-<name> --recursive
The listing shows the three files under raw/:
✓ Listing objects...
3 object(s) found
LAST MODIFIED SIZE KEY
------------- ---- ---
YYYY-MM-DD HH:MM:SS 852 raw/block-volumes.md
YYYY-MM-DD HH:MM:SS 893 raw/pricing-and-billing.md
YYYY-MM-DD HH:MM:SS 778 raw/serverless-containers.md
Re-running sync after you edit a document uploads only what changed. To use your own
documents, put the Markdown files in corpus/ instead.
Step 2: Deploy the LLM¶
The answers come from a chat model on a second Serverless Container. In the console, create a deployment with these settings, which follow the vLLM tutorial:
| Field | Value |
|---|---|
| Deployment name | rag-llm |
| GPU type | L40S 48GB, or RTX PRO 6000 if it has no capacity |
| Container image | docker.io/vllm/vllm-openai:v0.27.1 |
| Exposed HTTP port and healthcheck port | 8000, the port vLLM listens on |
| Healthcheck path | /health, the path vLLM exposes |
| Start command | Toggle on, which reveals the Entrypoint and CMD fields |
| Entrypoint | Leave empty, so the image's own entrypoint is used |
| CMD | Qwen/Qwen3-8B --served-model-name rag-llm --max-model-len 16384 |
| Environment variables | Add one with the name HF_TOKEN and <your-hugging-face-token> as the value. Leave HF_HOME as the form sets it |
Any model that vLLM serves with a chat endpoint works here; the retrieval side does not depend on it.
The first start takes longer than the embedder's, because the weights are about 16 GB. While
you wait, the system log shows the normal sequence: Scheduled, Pulling, Pulled,
Created, Started, then several Startup probe failed: connection refused entries while the
weights load. Those failures are expected, because nothing is listening on the port yet.
Once the deployment reports healthy, its page shows an Endpoint Address. Copy it, remove
the last /, and test it:
curl -sL -X POST <rag-llm-endpoint-address>/v1/chat/completions \
--header "Authorization: Bearer <your-inference-api-key>" \
--header 'Content-Type: application/json' \
--data '{"model": "rag-llm", "messages": [{"role": "user", "content": "Say hello."}], "chat_template_kwargs": {"enable_thinking": false}}'
enable_thinking: false is there because Qwen3 is a reasoning model: without it, every reply
starts with a <think>…</think> section of the model's reasoning notes, before the actual
answer, which costs about a hundred extra tokens per reply. query.py in Step 5 sends its
questions to this endpoint, with the same setting.
A working response looks like this, trimmed:
{"id":"chatcmpl-ba255b7f270fa7eb","object":"chat.completion","model":"rag-llm",
"choices":[{"index":0,"message":{"role":"assistant","content":"Hello! How can I assist you today?"},
"finish_reason":"stop"}],
"usage":{"prompt_tokens":15,"total_tokens":27,"completion_tokens":12}}
choices is the list of answers; there is one unless you ask for more, so the text is at
choices[0].message.content, which is where query.py reads it.
Step 3: Create the schema and roles¶
Back on the CPU instance. RAG needs two tables, one for the source files and one for their
chunks. Create a new database, ragdb, to hold them:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE documents (
doc_id TEXT PRIMARY KEY, -- object key without raw/ and .md
title TEXT NOT NULL,
checksum TEXT NOT NULL, -- sha256 of the file, to skip unchanged docs
ingested_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE chunks (
chunk_id TEXT PRIMARY KEY,
doc_id TEXT NOT NULL REFERENCES documents(doc_id) ON DELETE CASCADE,
chunk_index INT NOT NULL,
heading TEXT, -- the section the chunk came from
content TEXT NOT NULL,
embedding vector(1024) NOT NULL -- Qwen3-Embedding-0.6B, same as the Vector DB guide
);
documents holds one row per source file with a checksum, so a re-run can skip files that have
not changed. chunks holds the passages that are searched, each pointing back to its document.
Create two database roles, so the ingester can write while the query side only reads. Pick a password for each and substitute it for the placeholders, brackets included.
Still in psql:
CREATE ROLE rag_ingest LOGIN PASSWORD '<ingest-password>';
GRANT INSERT, UPDATE, DELETE, SELECT ON documents, chunks TO rag_ingest;
CREATE ROLE rag_query LOGIN PASSWORD '<query-password>';
GRANT SELECT ON documents, chunks TO rag_query;
Leave psql with \q.
Set the environment variables¶
The scripts read everything they need from the environment. Use the same two passwords as
above; HF_TOKEN is the Hugging Face token; the two endpoints are the Endpoint Address
shown on the embedder and rag-llm deployment pages, with or without the trailing /.
If a password contains @, : or /, percent-encode it:
export VERDA_INFERENCE_KEY='<your-inference-api-key>'
export HF_TOKEN='<your-hugging-face-token>'
export EMBEDDER_ENDPOINT='<embedder-endpoint-address>'
export LLM_ENDPOINT='<rag-llm-endpoint-address>'
export VERDA_S3_ACCESS_KEY='<object-storage-access-key>'
export VERDA_S3_SECRET_KEY='<object-storage-secret-key>'
export RAG_BUCKET='rag-corpus-<name>'
export PG_INGEST_DSN='postgresql://rag_ingest:<ingest-password>@localhost/ragdb'
export PG_QUERY_DSN='postgresql://rag_query:<query-password>@localhost/ragdb'
These variables last only for the current shell. If you disconnect from the instance, run this block again after you reconnect.
Check the database connection before going further:
The expected output is a one-row table containing 1:
That means the rag_ingest role and password match. If you get password authentication
failed, the DSN password differs from the one you set; reset it in psql as postgres on
ragdb with ALTER ROLE rag_ingest PASSWORD '<ingest-password>'; and run the check again.
Step 4: Chunk and ingest¶
Install the Python packages in a virtual environment, in the same shell where you set the variables:
sudo apt-get install -y python3-venv
python3 -m venv ~/vecenv
source ~/vecenv/bin/activate
pip install requests "psycopg[binary]" numpy boto3 tokenizers
The ingester¶
The ingester splits each document into chunks, short passages that are embedded and searched individually.
Save the following script as ingest.py in your home directory:
# ingest.py
import hashlib, os, re, requests
import boto3, numpy as np, psycopg
from tokenizers import Tokenizer
BUCKET = os.environ["RAG_BUCKET"]
ENDPOINT = os.environ["EMBEDDER_ENDPOINT"].rstrip("/") + "/v1/embeddings"
KEY = "".join(os.environ["VERDA_INFERENCE_KEY"].split()) # strips stray whitespace
TARGET_TOKENS = 450 # chunk size; see "How chunking works"
BATCH = 64
tokenizer = Tokenizer.from_pretrained("Qwen/Qwen3-Embedding-0.6B") # counts tokens like the model
FENCE = re.compile(r"^\s*(```|~~~)")
HEADING = re.compile(r"^#{1,6}\s+(.*)$")
def embed(texts: list[str]) -> np.ndarray:
"""The single source of truth for turning text into vectors. query.py imports it."""
resp = requests.post(
ENDPOINT,
headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
json={"model": "embedder", "input": texts},
timeout=300, # generous: covers cold start
)
resp.raise_for_status()
rows = [d["embedding"] for d in sorted(resp.json()["data"], key=lambda d: d["index"])]
vecs = np.asarray(rows, dtype=np.float32)
vecs /= np.linalg.norm(vecs, axis=1, keepdims=True) # unit length, for cosine
return vecs
def blocks(md: str):
"""Yield (heading, text) units: paragraphs, and fenced code blocks kept whole."""
heading, buf, fence = "", [], None
for line in md.split("\n") + [""]:
if fence is None and FENCE.match(line):
if buf:
yield heading, "\n".join(buf); buf = []
fence = FENCE.match(line).group(1); buf = [line]
elif fence and line.strip().startswith(fence):
buf.append(line); yield heading, "\n".join(buf); buf, fence = [], None
elif fence:
buf.append(line)
elif HEADING.match(line):
if buf:
yield heading, "\n".join(buf); buf = []
heading = HEADING.match(line).group(1).strip()
elif line.strip() == "":
if buf:
yield heading, "\n".join(buf); buf = []
else:
buf.append(line)
def chunk(md: str):
"""Pack consecutive units under one heading into chunks of about TARGET_TOKENS."""
chunks, cur, cur_heading, cur_tokens = [], [], None, 0
for heading, text in blocks(md):
n = len(tokenizer.encode(text).ids)
if cur and (heading != cur_heading or cur_tokens + n > TARGET_TOKENS):
chunks.append((cur_heading, "\n\n".join(cur))); cur, cur_tokens = [], 0
cur.append(text); cur_heading = heading; cur_tokens += n
if cur:
chunks.append((cur_heading, "\n\n".join(cur)))
return chunks
def ingest():
s3 = boto3.client(
"s3", endpoint_url="https://objects.fin-03.verda.storage", region_name="us-east-1",
aws_access_key_id=os.environ["VERDA_S3_ACCESS_KEY"],
aws_secret_access_key=os.environ["VERDA_S3_SECRET_KEY"],
)
conn = psycopg.connect(os.environ["PG_INGEST_DSN"])
pages = s3.get_paginator("list_objects_v2").paginate(Bucket=BUCKET, Prefix="raw/")
for obj in (o for p in pages for o in p.get("Contents", []) if o["Key"].endswith(".md")):
key = obj["Key"]
text = s3.get_object(Bucket=BUCKET, Key=key)["Body"].read().decode("utf-8")
doc_id = key[len("raw/"):-len(".md")]
checksum = hashlib.sha256(text.encode()).hexdigest()
with conn.cursor() as cur:
cur.execute("SELECT checksum FROM documents WHERE doc_id = %s", (doc_id,))
if (row := cur.fetchone()) and row[0] == checksum:
print(f"{doc_id}: unchanged"); continue
title = next((m.group(1) for m in map(HEADING.match, text.split("\n")) if m), doc_id)
cur.execute(
"""INSERT INTO documents (doc_id, title, checksum) VALUES (%s, %s, %s)
ON CONFLICT (doc_id) DO UPDATE
SET title = EXCLUDED.title, checksum = EXCLUDED.checksum, ingested_at = now()""",
(doc_id, title, checksum),
)
cur.execute("DELETE FROM chunks WHERE doc_id = %s", (doc_id,)) # replace wholesale
parts = chunk(text)
for i in range(0, len(parts), BATCH):
batch = parts[i:i + BATCH]
vecs = embed([f"{h}\n\n{c}" if h else c for h, c in batch])
cur.executemany(
"""INSERT INTO chunks (chunk_id, doc_id, chunk_index, heading, content, embedding)
VALUES (%s, %s, %s, %s, %s, %s)""",
[(f"{doc_id}#{i + j}", doc_id, i + j, h, c, v.tolist())
for j, ((h, c), v) in enumerate(zip(batch, vecs))],
)
conn.commit()
print(f"{doc_id}: {len(parts)} chunks")
if __name__ == "__main__":
ingest()
Four functions do the work:
embed()turns text into vectors. It is the function from the Vector DB guide.blocks()andchunk()split a Markdown file into chunks, one section at a time, keeping code blocks whole. How chunking works explains the rules and how to tune the size.ingest()reads every.mdfile underraw/in the bucket, skips files that have not changed since the last run, and replaces the chunks of any file that has.
Run the script:
The output shows one line per file:
The numbers match the sections in each file: block-volumes.md has an intro and three
headings, so four chunks. The sections are short, so none needed splitting further.
Build the index¶
Now that the rows are loaded, we will create the index. With this index, Postgres finds the nearest vectors without comparing the question against every chunk. Building it after the load, rather than before, means the inserts did not have to update it as they went.
Run this as the postgres superuser:
SET maintenance_work_mem = '2GB';
CREATE INDEX chunks_embedding_idx
ON chunks USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
Leave psql with \q.
Step 5: Ask a question¶
Save this as query.py in the same directory as ingest.py, since it imports embed()
from it. Keep the venv active and the environment variables from Step 3 set:
# query.py
import os, requests, psycopg
from ingest import embed
LLM = os.environ["LLM_ENDPOINT"].rstrip("/") + "/v1/chat/completions"
KEY = "".join(os.environ["VERDA_INFERENCE_KEY"].split())
conn = psycopg.connect(os.environ["PG_QUERY_DSN"])
SQL = """
SELECT c.heading, c.content, d.title, d.doc_id,
1 - (c.embedding <=> %(q)s::vector) AS similarity
FROM chunks c JOIN documents d ON d.doc_id = c.doc_id
ORDER BY c.embedding <=> %(q)s::vector
LIMIT %(k)s
"""
SYSTEM = """Answer the question using only the numbered sources below.
Put the source number in square brackets after each claim, like [2].
If the sources do not contain the answer, say so. Do not invent source numbers."""
def retrieve(question: str, k: int = 6):
qvec = embed([question])[0].tolist()
with conn.cursor() as cur:
cur.execute("SET LOCAL hnsw.ef_search = 100")
cur.execute(SQL, {"q": qvec, "k": k})
cols = [d.name for d in cur.description]
return [dict(zip(cols, r)) for r in cur.fetchall()]
def answer(question: str):
hits = retrieve(question)
sources = "\n\n".join(f"[{i}] {h['title']} — {h['heading']}\n{h['content']}"
for i, h in enumerate(hits, 1))
resp = requests.post(
LLM,
headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
json={"model": "rag-llm", "temperature": 0.2,
"chat_template_kwargs": {"enable_thinking": False}, # no <think> preamble
"messages": [{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"Sources:\n\n{sources}\n\nQuestion: {question}"}]},
timeout=300,
)
resp.raise_for_status()
return resp.json()["choices"][0]["message"]["content"], hits
if __name__ == "__main__":
import sys
text, hits = answer(" ".join(sys.argv[1:]) or "Am I charged for a volume when the instance is shut down?")
print(text, "\nSources")
for i, h in enumerate(hits, 1):
print(f" [{i}] {h['title']} — {h['heading']} ({h['doc_id']}.md, {h['similarity']:.2f})")
retrieve() is the search from the Vector DB guide, joined to documents for the title.
answer() numbers the retrieved chunks, tells the model to cite by number, and prints the
numbered list itself, so a citation always resolves to a real chunk.
Run the script with a question:
On the sample corpus, the output is:
Yes, you are charged for a volume when the instance is shut down. A Block Volume is charged
whether or not it is attached to an instance, and whether or not that instance is running [2].
Sources
[1] Pricing and billing — Shutting down and deleting instances (pricing-and-billing.md, 0.75)
[2] Pricing and billing — Storage charges (pricing-and-billing.md, 0.72)
[3] Pricing and billing — Pricing and billing (pricing-and-billing.md, 0.61)
[4] Block Volumes — Attaching and resizing (block-volumes.md, 0.61)
[5] Block Volumes — Deleting a volume (block-volumes.md, 0.59)
[6] Block Volumes — Cloning a volume (block-volumes.md, 0.56)
The claim points at a section of one of the three files you created. All six retrieved chunks are listed with their similarity to the question; the model used only the one it needed.
Now ask something the documents do not cover:
None of the provided sources address the question "Am I a cat?" Therefore, the answer cannot
be determined from the given sources.
Sources
[1] Serverless Containers — Serverless Containers (serverless-containers.md, 0.24)
...
The model says so instead of guessing, because the system prompt tells it to answer from the sources only. The similarity scores, around 0.2 instead of 0.7, show that nothing relevant was retrieved.
The score after each source runs from 0 to 1:
- 1 means the chunk has the same meaning as the question.
- 0 means the chunk is unrelated to it.
Congratulations! You now have a working RAG service.
After it works¶
When an answer is wrong¶
A citation tells you which source to check; it does not guarantee the model read that source correctly. Look at the Sources list first. If the passage that answers the question is not among the six, the model never saw it, and the fix is in chunking or retrieval rather than in the prompt.
How chunking works¶
The embedding model turns a piece of text into one vector that stands for its meaning as a whole, and search returns the vectors closest to the question's. A vector for a whole page captures what the page is about, not the details inside it; a vector for one section does capture that section's details, so it matches a question about them. Documents are therefore split into chunks of a few hundred tokens, each embedded on its own, and it is chunks that are searched and handed to the LLM.
Two rules matter when your documents are Markdown files with headings and text blocks, as most
product docs are, and ingest.py applies both:
- Split on headings, so every chunk carries its section heading. The heading helps retrieval and is what the citation shows.
- Never cut a command or code block in two. It is added to a chunk whole or not at all, because half a command looks like a complete answer but fails when run.
The chunk size is TARGET_TOKENS at the top of ingest.py. 450 suits documentation pages:
smaller chunks are more precise but lose surrounding context, larger ones keep context but
match specific questions less well. Go lower for FAQ-style content, higher for long narrative
text, and re-run ingest.py on the whole corpus after any change.
Updating the corpus¶
Edit or add Markdown files in corpus/, then run verda oss sync and python3 ingest.py
again. Files that have not changed (their checksum matches the one stored in documents) are
skipped, so only what changed is embedded.
Conclusion¶
You now have documents in Object Storage, their passages indexed in pgvector, an embedding model and an LLM on Serverless Containers, and a script that turns a question into a cited answer.
To clean up. The instance, the volume and the bucket keep billing while they exist. When you are finished, discontinue the instance, permanently delete the volume, delete both container deployments, and remove the bucket together with its files:
--force deletes everything in the bucket before removing it, and asks for confirmation first.
Related guides¶
- Vector DB guide: the database and embedding deployment this guide builds on, with sizing, thresholds and backups.
- Object Storage CLI: the full command reference.
- Serverless Containers storage: scratch disk and shared filesystems, for caching model weights across cold starts.