diff --git a/docs.json b/docs.json
index e0a2e738..ccec450d 100644
--- a/docs.json
+++ b/docs.json
@@ -127,6 +127,15 @@
"serverless/development/dual-mode-worker"
]
},
+ {
+ "group": "Storage",
+ "pages": [
+ "serverless/storage/modelrepo/overview",
+ "serverless/storage/modelrepo/security",
+ "serverless/storage/modelrepo/testing"
+ ]
+ },
+ "serverless/modelrepotest",
{
"group": "Manage endpoints",
"pages": [
diff --git a/serverless/modelrepotest.mdx b/serverless/modelrepotest.mdx
new file mode 100644
index 00000000..630f14a6
--- /dev/null
+++ b/serverless/modelrepotest.mdx
@@ -0,0 +1,171 @@
+---
+title: "Model Repo testing"
+description: "Upload a model to Model Repo and deploy it to a Serverless endpoint."
+---
+
+
+Model Repo is currently in alpha and is available on Mac and Linux only. Windows support is coming soon.
+
+
+## Why use Model Repo
+
+Model Repo lets you upload your own models to private storage on Runpod and attach them directly to Serverless endpoints. Key benefits:
+
+- **Faster cold starts**: Models are pre-cached on the worker host rather than downloaded at runtime.
+- **No HuggingFace dependency**: Your models are stored in Runpod's infrastructure, so endpoints don't require an outbound download on every cold start.
+- **Private storage**: Models are stored in your account and are not accessible to other users.
+
+---
+
+## Manual testing
+
+### Prerequisites
+
+- Your email is feature-flagged for Model Repo access.
+- `jq` is installed for parsing JSON output.
+
+### Set environment variables
+
+Export the following before running any commands. **Make sure to set your actual API key. Missing this is the most common source of auth errors later.**
+
+```bash
+export RUNPOD_API_URL="https://rest.runpod.io/v1"
+export RUNPOD_GRAPHQL_URL="https://api.runpod.io/graphql"
+export RUNPOD_API_KEY="your-api-key" # replace with your actual API key
+
+export MODEL_NAME="model_name" # unique name per test run — reusing the same name uploads a new version, not a new model
+export MODEL_PATH="/path/to/model" # local path to the model files you want to upload
+```
+
+
+`MODEL_NAME` must be unique for each test run. If you reuse the same name, the upload creates a new version of the existing model rather than a new model.
+
+
+---
+
+### Step 1: Install runpodctl
+
+**Option A: Install via Homebrew (recommended)**
+
+```bash
+brew install runpod/runpodctl/runpodctl
+```
+
+**Option B: Build from source**
+
+```bash
+brew install go # install Go, required to build runpodctl
+git clone git@github.com:runpod/runpodctl.git
+cd runpodctl
+make # builds the binary to ./bin/runpodctl
+```
+
+
+If you build from source, the binary is at `./bin/runpodctl`. Either run it with that path, or add `./bin` to your `PATH`. The steps below use `runpodctl`. Adjust accordingly.
+
+
+---
+
+### Step 2: Upload the model
+
+```bash
+# --name: the name to register the model under in your repo
+# --model-path: local path to the model files
+# --create-upload: creates the upload session and transfers files
+runpodctl model add \
+ --name "$MODEL_NAME" \
+ --model-path "$MODEL_PATH" \
+ --create-upload
+```
+
+This outputs a JSON string listing all uploaded files.
+
+---
+
+### Step 3: Wait for the model to be hashed
+
+After upload, the model must be hashed by an asynchronous background process. This typically completes in a few minutes but can take up to 10–15 minutes.
+
+Poll until the `hash` field is non-null:
+
+```bash
+runpodctl model list --name "$MODEL_NAME" | jq -r '.[0].versions[0].hash'
+```
+
+While hashing is in progress, the command returns `null`:
+
+```
+% runpodctl model list --name "$MODEL_NAME" | jq -r '.[0].versions[0].hash'
+null
+```
+
+Once hashing is complete, it returns the hash value:
+
+```
+% runpodctl model list --name "$MODEL_NAME" | jq -r '.[0].versions[0].hash'
+71a311bdf0ca44119ed74dbef8cf573bc89b58cbc48a10fe508f756ebb1922dc
+```
+
+---
+
+### Step 4: Get your user ID and model hash
+
+```bash
+export USER_ID="$(runpodctl user | jq -r '.id')" # your Runpod user ID
+export MODEL_HASH="$(runpodctl model list --name "$MODEL_NAME" | jq -r '.[0].versions[0].hash')" # the hash from step 3
+```
+
+---
+
+### Step 5: Deploy a Serverless endpoint with the model attached
+
+```bash
+# --name: name for the endpoint
+# --hub-id: the Hub template to deploy
+# --gpu-id: GPU type
+# --workers-max: maximum number of active workers
+# --workers-min: minimum number of workers kept warm
+# --model-reference: attaches your model to the endpoint
+# --env: sets the model path on the worker
+# --min-cuda-version: works around a bug in runpodctl
+runpodctl serverless create \
+ --name "my_worker" \
+ --hub-id "cm8h09d9n000008jvh2rqdsmb" \
+ --gpu-id "AMPERE_24" \
+ --workers-max 3 \
+ --workers-min 1 \
+ --model-reference "https://local/$USER_ID/$MODEL_NAME:$MODEL_HASH" \
+ --env MODEL_NAME="/runpod/model-store/modelrepo-local/models/$USER_ID/$MODEL_NAME/$MODEL_HASH" \
+ --min-cuda-version "13.0"
+```
+
+
+`--model-reference` is only supported with `--hub-id` and GPU endpoints. It is repeatable if you need to attach multiple models to the same endpoint.
+
+
+---
+
+### Step 6: Verify the model is working
+
+Send a test request to confirm the endpoint is live and the model is accessible. Replace `ENDPOINT_ID` with the ID returned in the previous step:
+
+```bash
+curl -s -X POST "https://api.runpod.ai/v2/${ENDPOINT_ID}/runsync" \
+ -H "Authorization: Bearer $RUNPOD_API_KEY" \
+ -H "Content-Type: application/json" \
+ -d '{"input": {"prompt": "hello"}}' | jq
+```
+
+A successful response confirms the endpoint is running and the model is attached. If the request fails with an auth error, verify that `RUNPOD_API_KEY` is set correctly.
+
+If you prefer a graphical interface to curl, you can also send requests to the worker from the web UI.
+
+---
+
+### Step 7: Clean up
+
+Delete the endpoint after testing to stop accruing spend. Use the web UI or:
+
+```bash
+runpodctl serverless delete
+```
diff --git a/serverless/storage/modelrepo/overview.mdx b/serverless/storage/modelrepo/overview.mdx
new file mode 100644
index 00000000..26db375a
--- /dev/null
+++ b/serverless/storage/modelrepo/overview.mdx
@@ -0,0 +1,64 @@
+---
+title: "Overview"
+sidebarTitle: "Overview"
+description: "Store, version, and pre-cache your model files on Runpod infrastructure."
+---
+
+
+ Model Repo is currently in beta and is available on Mac and Linux only. Windows support is coming soon.
+
+
+Model Repo is a private model storage service built into Runpod. Upload your model files once, and Runpod caches them directly on your Serverless worker hosts so they are ready before the worker starts. No external download is required at cold start time.
+
+## Key benefits
+
+- **Faster cold starts**: Models are pre-cached on worker hosts when available, so workers start faster without waiting for an external download.
+- **No HuggingFace dependency**: Your models are stored in Runpod's infrastructure, so endpoints don't require an outbound download on every cold start.
+- **Private storage**: Models are stored in your account and are not accessible to other users.
+- **Version control**: Each upload is content-addressed, so you can pin an endpoint to an exact model version and roll back at any time.
+
+## How it works
+
+```mermaid
+flowchart LR
+ A["Upload model\nrunpodctl model add"] --> B["Runpod hashes\n& stores files"]
+ B --> C["Configure endpoint\n--model-reference URL"]
+ C --> D["Worker host\npre-caches files"]
+ D --> E["Worker starts\nfiles ready on disk"]
+```
+
+1. Upload your model files using `runpodctl model add`.
+2. Runpod computes a content hash and stores the files in private, secure storage.
+3. Reference the model using the URL `https://local/{user-id}/{model-name}:{hash}` when configuring your endpoint.
+4. When a worker starts, Runpod pre-caches the model files on the host before the container boots.
+5. Your handler reads the model from a local path inside the container.
+
+## What you can upload
+
+Model Repo accepts any file format — PyTorch checkpoints, GGUF, safetensors, ONNX, or any format your worker needs. There is no type checking.
+
+- **Max file size:** 5TB per file
+- **Retention:** Files are retained indefinitely
+- **Total storage limit:** Still being finalized
+
+## Model Repo vs network volumes
+
+**Use Model Repo when:**
+- You have fixed, versioned model weights and want faster cold starts without external downloads
+- You need version control over exactly which checkpoint is deployed
+- Your endpoint spans multiple datacenters — Model Repo works across all DCs, while a network volume is tied to one
+
+**Use a network volume when:**
+- Workers need shared read/write access during a run (e.g., checkpoints, fine-tuning outputs)
+- Your files change frequently and all workers need to see updates immediately
+- You need many concurrent workers accessing the same files simultaneously
+
+## Your data
+
+Your models are stored in private, secure storage and are isolated to your Runpod account — no other users can access or list your models. Runpod does not access, analyze, or use your model files for any purpose.
+
+For encryption details, certifications, and compliance information, see [Security](/serverless/storage/modelrepo/security).
+
+## Getting started
+
+See [Model Repo testing](/serverless/storage/modelrepo/testing) for a step-by-step guide to uploading a model and deploying it to a Serverless endpoint.
diff --git a/serverless/storage/modelrepo/security.mdx b/serverless/storage/modelrepo/security.mdx
new file mode 100644
index 00000000..97e7a2de
--- /dev/null
+++ b/serverless/storage/modelrepo/security.mdx
@@ -0,0 +1,40 @@
+---
+title: "Security"
+sidebarTitle: "Security"
+description: "How Model Repo stores and protects your model data."
+---
+
+## Storage
+
+Model Repo stores your model files in Runpod's infrastructure, backed by [Cloudflare R2](https://developers.cloudflare.com/r2/). Your data is isolated to your Runpod account and is not accessible to other users.
+
+## Encryption
+
+All model data is encrypted at every stage:
+
+- **At rest**: AES-256-GCM encryption, applied by default to all objects stored in Cloudflare R2.
+- **In transit**: TLS encryption for all transfers, both between your machine and Runpod when uploading and between Runpod's systems when caching models on worker hosts.
+
+## Your data is yours
+
+Runpod does not access, analyze, or use your model files for any purpose. Models stored in Model Repo are not used for training, evaluation, or any other internal purpose. Only you can access your models through your account credentials.
+
+## Compliance
+
+/* [ENGINEERING: Ben Rosenberg mentioned linking to SOC1/SOC2 certs. What certifications does Runpod currently hold, and where are they published? Confirm the correct link for Runpod's security/compliance page.] -->
+
+Runpod maintains industry-standard security certifications. For a full overview of Runpod's security posture, certifications, and policies, see [runpod.io/security](https://runpod.io/security).
+
+## Frequently asked questions
+
+**Does Runpod use my models to train other models?**
+
+No. Your model files are stored privately and are never used by Runpod for any purpose.
+
+**Can other Runpod users access my models?**
+
+No. Models are scoped to your Runpod account. Other users cannot access or list your models.
+
+**Who manages the storage infrastructure?**
+
+Cloudflare R2 is the underlying object storage backend. Cloudflare applies encryption-at-rest using AES-256-GCM by default. Runpod manages the access layer and authentication. Only your Runpod API key can retrieve your models.
diff --git a/serverless/storage/modelrepo/testing.mdx b/serverless/storage/modelrepo/testing.mdx
new file mode 100644
index 00000000..43d3701a
--- /dev/null
+++ b/serverless/storage/modelrepo/testing.mdx
@@ -0,0 +1,147 @@
+---
+title: "Model Repo testing"
+description: "Upload a model to Model Repo and deploy it to a Serverless endpoint."
+---
+
+
+ Model Repo is currently in beta and is available on Mac and Linux only. Windows support is coming soon.
+
+
+## Why use Model Repo
+
+Model Repo lets you upload your own models to private storage on Runpod and attach them directly to Serverless endpoints. Key benefits:
+
+- **Faster cold starts**: Models are pre-cached on worker hosts when available, so cold starts are typically faster. In some cases a model may still need to be fetched from storage.
+- **No HuggingFace dependency**: Your models are stored in Runpod's infrastructure, so endpoints don't require an outbound download on every cold start.
+- **Private storage**: Models are stored in your account and are not accessible to other users.
+
+---
+
+## Prerequisites
+
+- Your email is feature-flagged for Model Repo access.
+- `jq` is installed for parsing JSON output.
+
+---
+
+## Step 1: Set environment variables
+
+Export the following before running any commands. **Make sure to set your actual API key. Missing this is the most common source of auth errors later.**
+
+```bash
+export RUNPOD_API_URL="https://rest.runpod.io/v1"
+export RUNPOD_GRAPHQL_URL="https://api.runpod.io/graphql"
+export RUNPOD_API_KEY="your-api-key" # replace with your actual API key
+
+export MODEL_NAME="my-model" # name to register your model under
+export MODEL_PATH="/path/to/model" # local path to the model files you want to upload
+```
+
+
+ Each upload under the same `MODEL_NAME` creates a new version of that model, not a new model. Use a different name if you want a separate model entry.
+
+
+---
+
+## Step 2: Install runpodctl
+
+**Option A: Install via Homebrew (recommended)**
+
+```bash
+brew install runpod/runpodctl/runpodctl
+```
+
+**Option B: Build from source**
+
+```bash
+brew install go # install Go, required to build runpodctl
+git clone git@github.com:runpod/runpodctl.git
+cd runpodctl
+make # builds the binary to ./bin/runpodctl
+```
+
+
+ If you build from source, the binary is at `./bin/runpodctl`. Either run it with that path, or add `./bin` to your `PATH`. The steps below use `runpodctl`. Adjust accordingly.
+
+
+---
+
+## Step 3: Upload the model
+
+```bash
+# --name: the name to register the model under in your repo
+# --model-path: local path to the model files
+# --create-upload: creates the upload session and transfers files
+# --wait-for-hash: waits for the background hashing process to complete
+runpodctl model add \
+ --name "$MODEL_NAME" \
+ --model-path "$MODEL_PATH" \
+ --create-upload \
+ --wait-for-hash
+```
+
+This outputs a JSON string listing all uploaded files.
+
+---
+
+## Step 4: Get your user ID and model hash
+
+```bash
+export USER_ID="$(runpodctl user | jq -r '.id')" # your Runpod user ID
+export MODEL_HASH="$(runpodctl model list --name "$MODEL_NAME" | jq -r '.[0].versions[0].hash')" # the model version hash
+```
+
+---
+
+## Step 5: Deploy a Serverless endpoint with the model attached
+
+```bash
+# --name: name for the endpoint
+# --hub-id: Hub template to deploy
+# --gpu-id: GPU type
+# --workers-max: maximum number of active workers
+# --workers-min: minimum number of workers kept warm
+# --model-reference: attaches your model to the endpoint
+# --env MODEL_NAME: sets the model path on the worker
+# --min-cuda-version: works around a bug in runpodctl
+runpodctl serverless create \
+ --name "my_worker" \
+ --hub-id "cm8h09d9n000008jvh2rqdsmb" \
+ --gpu-id "AMPERE_24" \
+ --workers-max 3 \
+ --workers-min 1 \
+ --model-reference "https://local/$USER_ID/$MODEL_NAME:$MODEL_HASH" \
+ --env MODEL_NAME="/runpod/model-store/modelrepo-local/models/$USER_ID/$MODEL_NAME/$MODEL_HASH" \
+ --min-cuda-version "13.0"
+```
+
+
+ `--model-reference` is only supported with `--hub-id` and GPU endpoints. It is repeatable if you need to attach multiple models to the same endpoint. The `--env MODEL_NAME` flag passes the full local path to the worker so your handler knows where to find the model files.
+
+
+---
+
+## Step 6: Verify the model is working
+
+Send a test request to confirm the endpoint is live and the model is accessible. Replace `ENDPOINT_ID` with the ID returned in the previous step:
+
+```bash
+curl -s -X POST "https://api.runpod.ai/v2/${ENDPOINT_ID}/runsync" \
+ -H "Authorization: Bearer $RUNPOD_API_KEY" \
+ -H "Content-Type: application/json" \
+ -d '{"input": {"prompt": "hello"}}' | jq
+```
+
+A successful response confirms the endpoint is running and the model is attached. If the request fails with an auth error, verify that `RUNPOD_API_KEY` is set correctly.
+
+You can also send requests from the web UI if you prefer a graphical interface.
+
+---
+
+## Step 7: Clean up
+
+Delete the endpoint after testing to stop accruing spend. Use the web UI or:
+
+```bash
+runpodctl serverless delete
+```