vielhuber / runpodhelper
Automates self-hosted llm inference
Requires
- php: >=8.1
- monolog/monolog: ^3.10.0
- php-mcp/server: ^3.3
- vielhuber/aihelper: ^4.2.1
- vlucas/phpdotenv: ^5.6.4
This package is auto-updated.
Last update: 2026-07-18 06:33:12 UTC
README
⛈ runpodhelper ⛈
runpodhelper automates the full lifecycle of self-hosted llm inference on runpod gpu cloud. it provisions pods via the runpod graphql api, installs lm studio or llama.cpp, downloads gguf models from huggingface, and serves them behind a cloudflare tunnel.
usage
inference
./vendor/bin/runpod.sh create --config pods.yaml
./vendor/bin/runpod.sh delete --all
./vendor/bin/runpod.sh status
./vendor/bin/runpod.sh create \
--gpu "4x RTX PRO 6000" \
--hdd 250 \
--model "unsloth/MiniMax-M2.7-GGUF-UD-Q4_K_XL" \
--image "runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404" \
--type "llamacpp" \
--api-key "your-static-api-key" \
--context-length 262144 \
--parallel 2 \
--datacenter "EUR-IS-2" \
--auto-destroy 3600
./vendor/bin/runpod.sh delete --id 001
./vendor/bin/runpod.sh test quality --runs 5
./vendor/bin/runpod.sh test quantity --runs 80
./vendor/bin/runpod.sh scale --start \
--gpu "RTX 5090" \
--hdd 50 \
--model "unsloth/Qwen3.5-27B-GGUF-UD-Q4_K_XL" \
--image "runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404" \
--type "lmstudio" \
--api-key "your-static-api-key" \
--context-length 65536 \
--parallel 2 \
--datacenter "EUR-IS-2" \
--auto-destroy 3600 \
--pod-count 3
./vendor/bin/runpod.sh scale --start \
--gpu "L40S" \
--hdd 60 \
--model "unsloth/Qwen3.5-35B-A3B-GGUF" \
--image "runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404" \
--type "llamacpp" \
--api-key "your-static-api-key" \
--context-length 131072 \
--parallel 1 \
--pod-count 1
./vendor/bin/runpod.sh scale --start \
--gpu "2x RTX PRO 6000" \
--hdd 250 \
--model "unsloth/MiniMax-M2.7-GGUF-UD-Q4_K_XL" \
--image "runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404" \
--type "llamacpp" \
--api-key "your-static-api-key" \
--context-length 131072 \
--parallel 1 \
--pod-count 1
./vendor/bin/runpod.sh scale --start \
--gpu "4x RTX PRO 6000" \
--hdd 250 \
--model "unsloth/MiniMax-M2.7-GGUF-UD-Q4_K_XL" \
--image "runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404" \
--type "llamacpp" \
--api-key "your-static-api-key" \
--context-length 262144 \
--parallel 2 \
--pod-count 1
./vendor/bin/runpod.sh scale --stop
./vendor/bin/runpod.sh scale --pod-count 20
./vendor/bin/runpod.sh scale --refresh --context-length 65536 --parallel 2
./vendor/bin/runpod.sh scale --refresh
post-training
post-training runs entirely on runpod with the standard pytorch image and a local pod volume. no local unsloth studio installation or custom image is required.
./vendor/bin/runpod.sh studio up ./vendor/bin/runpod.sh studio deploy ./vendor/bin/runpod.sh studio down
studio up tests each new host against the pytorch cdn and replaces hosts below 20 mb/s before installation. it then installs studio, downloads the configured transformers/safetensors model, opens a local ssh tunnel and prints the login data. its complete console output is mirrored to logs/studio/latest-up.log. upload, training and gguf/lora/safetensors export happen in studio. an optional HF_TOKEN from .env is available to the backend but is deliberately not exposed in the browser field.
dataset workflow:
- edit
posttraining/data.xlsxin excel and keep the first row with the column namesinputandoutput. - add one training example per row:
inputcontains the complete question or task including its context, andoutputcontains the desired answer. - export the active worksheet as utf-8 csv with a comma delimiter and verify that the first line is
"input","output"; german excel often exports semicolons, which studio does not accept. - upload the exported csv as a local dataset in studio and keep the target format on automatic so studio maps both columns to the conversation roles.
studio deploy serves the newest gguf through the configured api. download all required artifacts before studio down: it permanently deletes the pod, its local volume and all studio data so no pod or volume costs remain.
studio.yaml:
studios: - gpu: 'RTX 4090' hdd: 40 volume: 200 image: runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404 model: unsloth/Qwen3.6-27B type: llamacpp api_key: your-static-api-key context_length: 8192
rules
gpu-vram ≈ model-size + context-length * model-factortoken-budget-per-session ≈ parallel * context-lengthworkers-per-pod ≈ workers-count / pod-countrunning-workers-per-pod ≈ parallelconcurrent-workers ≈ parallel * pod-countconcurrent-workers ≈ 0.2 * parallel * workers-countpod-count ≈ 0.2 * workers-count
RTX 5090 + Qwen3.5-27B
gpu-vram ≈ 32 GBmodel-size ≈ 17.6 GBmodel-factor ≈ 0.00022=> max-context-length ≈ 65536=> max-parallel ≈ 2(at context-length 65536)
L40S + Qwen3.5-27B
gpu-vram ≈ 48 GBmodel-size ≈ 17.6 GBmodel-factor ≈ 0.00022=> max-context-length ≈ 138240=> max-parallel ≈ 4(at context-length 65536)
RTX PRO 6000 + Qwen3.5-122B-A10B (MoE)
gpu-vram ≈ 96 GBmodel-size ≈ 66 GBmodel-factor ≈ 0.00013(MoE, 10B active params)=> max-context-length ≈ 131072(128K model limit)=> max-parallel ≈ 1(at context-length 131072, ~83 GB total)=> max-parallel ≈ 2(at context-length 98304, ~92 GB total)
RTX PRO 6000 + Qwen3.5-27B
gpu-vram ≈ 96 GBmodel-size ≈ 17.6 GBmodel-factor ≈ 0.00022=> max-context-length ≈ 356352=> max-parallel ≈ 10(at context-length 65536)
RTX PRO 6000 + gemma-4-26B-A4B (MoE)
gpu-vram ≈ 96 GBmodel-size ≈ 27.9 GBmodel-factor ≈ 0.00013(MoE, 30 layers)=> max-context-length ≈ 256K (model limit)=> max-parallel ≈ 8(at context-length 65536)
RTX PRO 6000 + gemma-4-31B
gpu-vram ≈ 96 GBmodel-size ≈ 27.5 GBmodel-factor ≈ 0.00025(dense, 60 layers)=> max-context-length ≈ 256K (model limit)=> max-parallel ≈ 4(at context-length 65536)
RTX PRO 6000 + Qwen3.6-35B-A3B (MoE)
gpu-vram ≈ 96 GBmodel-size ≈ 22 GB(UD-Q4_K_XL)model-factor ≈ 0.00013(MoE, 3B active params)=> max-context-length ≈ 256K (model limit)=> max-parallel ≈ 2(at context-length 131072, ~56 GB total)=> max-parallel ≈ 4(at context-length 65536, ~56 GB total)
installation
- install library
composer require vielhuber/runpodhelper./vendor/bin/runpod.sh init
- setup cloudflare
- Create a domain
custom.xyz - Profile > API Tokens > Create Token
- Permissions:
Zone / DNS / EditZone / Single Redirect / EditAccount / Cloudflare Tunnel / Edit
- Account Resources
Include > Your account
- Zone Resource
- Include / Specific zone /
custom.xyz
- Include / Specific zone /
- Permissions:
- Set
CLOUDFLARE_DOMAIN/CLOUDFLARE_API_KEYin.env - Each pod gets a subdomain based on its config ID:
001.custom.xyz002.custom.xyz- …
- Create a domain
- edit config
vi ./.envvi ./models.yaml
mcp server
{
"mcpServers": {
"runpodhelper": {
"command": "/usr/bin/php",
"args": ["/path/to/project/runpodhelper/bin/mcp-server.php"]
}
}
}
recommended models
| Name | HDD | Model | Context length | Parallel | tok/s | Notes |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | 50 GB | Qwen3.5-27B-GGUF-UD-Q4_K_XL | 65536 | 2 | ~43 | best current MCP/tool-use baseline |
| NVIDIA L40S | 50 GB | Qwen3.5-27B-GGUF-UD-Q4_K_XL | 65536 | 4 | ~25 | 2x parallel slots vs. RTX 5090 |
| NVIDIA RTX PRO 6000 | 50 GB | Qwen3.5-27B-GGUF-UD-Q4_K_XL | 65536 | 10 | ~20 | max parallel slots, single pod |
| NVIDIA RTX PRO 6000 | 50 GB | gemma-4-26B-A4B-it-GGUF-UD-Q8_K_XL | 65536 | 8 | ~65 | MoE: 3.8B active params, best parallelism on 96 GB |
| NVIDIA RTX PRO 6000 | 50 GB | gemma-4-31B-it-GGUF-UD-Q6_K_XL | 65536 | 4 | ~18 | dense, best reliability on 96 GB |
| NVIDIA RTX PRO 6000 | 80 GB | Qwen3.5-122B-A10B-GGUF-UD-Q4_K_XL | 131072 | 1 | ~? | MoE: 10B active params, #1 intelligence index |
| NVIDIA RTX PRO 6000 | 50 GB | Qwen3.6-35B-A3B-GGUF-UD-Q4_K_XL | 131072 | 2 | ~? | MoE: 3B active params, high throughput at 128K ctx |
| NVIDIA A40 | 50 GB | Qwen3.5-27B-GGUF-UD-Q4_K_XL | 65536 | 2 | ~20 | discontinued/unavailable as of 2026-03 |
manual deployment
- https://www.runpod.io > Pods > Deploy
- Pod template > Edit
- Expose HTTP ports (comma separated):
1234 - Container Disk:
100 GB - Copy: SSH over exposed TCP
ssh root@xxxxxxxxxx -p xxxxx
curl -fsSL https://lmstudio.ai/install.sh | bash export PATH="/root/.lmstudio/bin:$PATH" # this is unreliable #lms get -y qwen/qwen3-coder-next mkdir -p ~/.lmstudio/models/unsloth/MiniMax-M2.1-GGUF cd ~/.lmstudio/models/unsloth/MiniMax-M2.1-GGUF wget -c https://huggingface.co/unsloth/MiniMax-M2.1-GGUF/resolve/main/MiniMax-M2.1-UD-TQ1_0.gguf mkdir -p ~/.lmstudio/models/lmstudio-community/Qwen3.5-35B-A3B-GGUF cd ~/.lmstudio/models/lmstudio-community/Qwen3.5-35B-A3B-GGUF wget -c https://huggingface.co/lmstudio-community/Qwen3.5-35B-A3B-GGUF/resolve/main/Qwen3.5-35B-A3B-Q4_K_M.gguf lms server start --port 1234 --bind 0.0.0.0
alternative: use runpodctl
ssh-keygen -t ed25519 -C "name@tld.com"wget https://github.com/Run-Pod/runpodctl/releases/download/v1.14.3/runpodctl-linux-amd64 -O runpodctlchmod +x runpodctlmv runpodctl /usr/bin/runpodctlrunpodctl config --apiKey <RUNPOD_API_KEY>runpodctl version
more commands
curl http://localhost:1234/v1/modelslms --helplms statuslms server stop- Copy: HTTP services > URL
curl https://xxxxxxxxx-1234.proxy.runpod.net/v1/responses \ -X POST \ -H "Content-Type: application/json" \ -H "Authorization: Bearer your-static-api-key" \ -d '{ "model": "xxxxxxxxxxxxx", "messages": [ {"role": "user", "content": [{"type": "input_text", "text": "hi"}]} ], "temperature": 1.0, "stream": true }'