Deploy any open LLM as your own API
Paste a HuggingFace GGUF link, pick a GPU, and get a dedicated OpenAI- and Anthropic-compatible endpoint. Flat monthly price. No DevOps.
Build your endpoint
Configure your deployment
Pick a trending model — or paste any HuggingFace GGUF link — then choose a GPU and deploy. We pre-select the cheapest card that fits and re-check it before you pay.
Pick a trending model
Top GGUF models on HuggingFace, updated daily.
Pick a GPU
One flat monthly price — no per-token or idle charges.
You're deploying
Xing4.0 29B A4B · 256K ctx on RTX 5090 · $749/mo
From link to live endpoint in three steps
No Dockerfiles, no vLLM configs, no GPU shopping. Paste, pick, deploy.
Paste a GGUF link
Drop in a HuggingFace GGUF repo or a direct .gguf URL — or pick one of the trending models. Choose your quant: Q4_K_M, Q5_K_M or Q8_0.
We recommend a GPU
GGUF Cloud reads the model size and pre-selects the cheapest dedicated GPU that fits, showing the flat monthly price before you commit.
Get your endpoint + key
One click spins up a dedicated llama.cpp container. You get an OpenAI- and Anthropic-compatible endpoint and your ModelsLab API key.
Why GGUF Cloud
Built for the models everyone else turns away
GGUF & llama.cpp native
Bring your own quantization — Q4, Q6 or Q8. The managed LLM hosts force vLLM and safetensors and reject GGUF outright. We run it as a first-class primitive.
Flat monthly per GPU
One predictable price per dedicated GPU. No per-token billing, no idle meter ticking at zero requests — the single most-complained-about thing in the category.
Dedicated single-tenant
Your model runs on its own GPU. No noisy neighbours, no shared-pool throttling, no preemptible boxes vanishing mid-run. Guaranteed throughput.
OpenAI and Anthropic
Every endpoint speaks both protocols natively. Point the OpenAI SDK, the Anthropic SDK, or Claude Code straight at your own model — no translation layer.
Automatic GPU sizing
We compute the VRAM your model needs from its parameters, quant and context, then recommend the cheapest card that fits. Nobody else does this.
Transparent pricing
Every GPU tier and its monthly price is on this page. No "contact sales" wall to find out what a dedicated endpoint actually costs.
Point any SDK at your own model
Your endpoint speaks both the OpenAI and Anthropic protocols natively — no shim, no rewrites. Swap the base URL and keep your existing code.
OpenAI SDK (Python)
Python1from openai import OpenAI23client = OpenAI(4 base_url="https://modelslab.com/api/gguf/{id}/v1",5 api_key="$MODELSLAB_API_KEY",6)78resp = client.chat.completions.create(9 model="local",10 messages=[{"role": "user", "content": "Explain GGUF in one line."}],11 stream=True,12)13for chunk in resp:14 print(chunk.choices[0].delta.content or "", end="")
Anthropic SDK (Python)
Python1from anthropic import Anthropic23# Point Claude Code or the Anthropic SDK at your own model.4client = Anthropic(5 base_url="https://modelslab.com/api/gguf/{id}",6 api_key="$MODELSLAB_API_KEY",7)89msg = client.messages.create(10 model="local",11 max_tokens=512,12 messages=[{"role": "user", "content": "Explain GGUF in one line."}],13)14print(msg.content[0].text)
cURL
bash1curl https://modelslab.com/api/gguf/{id}/v1/chat/completions \2 -H "Authorization: Bearer $MODELSLAB_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "local",6 "messages": [{"role": "user", "content": "Hello!"}]7 }'
Deploy a trending model
Llama, Qwen, Mistral, Mixtral, Phi, DeepSeek and more — pre-verified GGUF builds, or paste any HuggingFace repo of your own.
Xing4.0 29B A4B
TrendingNew31.22B · Q4_K_M · 18.95 GB · 256K ctx
Fits RTX 5090 · $749/mo
Humanizer
TrendingNew11.91B · Q4_K_M · 7.63 GB · 256K ctx
Fits RTX 4090 · $499/mo
Deepseek V4 Flash 0731
TrendingNew19.85B · BF16 · 11.31 GB · 1024K ctx
Fits RTX A6000 · $899/mo
Ternary Bonsai 2 27B
TrendingNew26.9B · BF16 · 0.93 GB · 256K ctx
Fits A100 80GB · $1,499/mo
Qwen3.8 27B
Trending27.32B · Q4_K_M · 16.46 GB · 256K ctx
Fits RTX 3090 · $249/mo
Qwen3.8 27B Gsq Rco
TrendingNew26.9B · BF16 · 0.93 GB · 256K ctx
Fits A100 80GB · $1,499/mo
Swift 1.5 Qwen3.8 27B Gsq Rco
TrendingNew26.9B · IQ2_S · 18.87 GB · 256K ctx
Fits RTX 3090 · $249/mo
Qwen3.8 9B
TrendingNew9.2B · Q4_K_M · 5.78 GB · 256K ctx
Fits RTX 4090 · $499/mo
Qwen3.8 27B
TrendingNew27.32B · Q4_K_M · 17.12 GB · 256K ctx
Fits RTX 3090 · $249/mo
Qwen3.8 27B Humanlike Chat
TrendingNew26.9B · Q4_K_M · 16.56 GB · 256K ctx
Fits RTX 3090 · $249/mo
Deepseek V4
TrendingMoE19.85B · IQ2XXS · 1189.42 GB
Fits RTX 5090 · $749/mo
Xing4.0 29B A4B
TrendingNew31.22B · IQ4_NL · 20.1 GB · 256K ctx
Fits RTX 3090 · $249/mo
Twil Lm3
TrendingNew3.08B · Q4_K_M · 1.92 GB · 64K ctx
Fits RTX 4090 · $499/mo
Redcell 26B A4B Osint Cyber Apex
TrendingMoE25.23B · BF16 · 1.19 GB · 256K ctx
Fits A100 80GB · $1,499/mo
Qwen Image 2.1 Text Encoder Heretic
TrendingNew8.19B · Q4_K_M · 5.03 GB · 256K ctx
Fits RTX 4090 · $499/mo
Qwen3 Coder 30B A3B Instruct
TrendingMoE30.53B · Q4_K_M · 18.56 GB · 256K ctx
Fits RTX 5090 · $749/mo
Dolphin3 Cyber 8B
Trending8.03B · Q4_K_M · 4.92 GB · 128K ctx
Fits RTX 4090 · $499/mo
Qwopus3.6 27B Coder Mtp
TrendingNew0.46B · Q4_K_M · 16.81 GB
Fits RTX 4090 · $499/mo
Gemma 4 12B Coder Fable5 Composer2.5 V1
Trending11.91B · Q4_K_M · 7.38 GB · 256K ctx
Fits RTX 4090 · $499/mo
Kwaipilot Kat Coder V2.5 Dev
TrendingNewMoE34.66B · Q4_K_M · 21.39 GB · 256K ctx
Fits RTX 5090 · $749/mo
Huihui Qwythos 9B Claude Mythos 5 1M Abliterated
TrendingNew9.2B · BF16 · 19.33 GB · 1024K ctx
Fits RTX 4090 · $499/mo
Qwen27B Abliterated Fable Mtp
NewMoE27.32B · Q4_K_M · 16.81 GB · 256K ctx
Fits RTX 3090 · $249/mo
One flat price per dedicated GPU
Pick the card your model fits on. Billed monthly, cancel anytime — no per-token charges, no idle meter, no surprises.
- Flat monthly billing
- Cancel anytime
- No per-token charges
- Pause to stop compute billing
Why teams pick GGUF Cloud
The managed LLM hosts reject GGUF. The dedicated GPU hosts bill hourly and punish idle time. GGUF Cloud is the synthesis.
| Capability | GGUF Cloud | HF Endpoints | Together / Fireworks | Featherless |
|---|---|---|---|---|
| Runs GGUF / llama.cpp | Yes | Yes | Rejects GGUF | FP8 only |
| Bring your own quant | Q4 / Q6 / Q8 | Limited | No | No |
| Dedicated single-tenant | Yes | Yes | Shared | Shared |
| Flat monthly price | Per GPU | Hourly | Per token | Yes |
| No idle / per-token charges | Yes | Idle billed | Per token | Yes |
| OpenAI + Anthropic API | Both | OpenAI only | Varies | OpenAI only |
| Automatic GPU recommendation | Yes | No | No | No |
Comparison reflects each provider's standard offering as of June 2026. Based on publicly documented pricing and engine support.
GGUF Cloud FAQ
Deploy your first model in minutes
Bring your own GGUF or pick a trending one. Get a dedicated, single-tenant endpoint that speaks OpenAI and Anthropic — for a price you can predict.
Explore Our Other Solutions
Unlock your creative potential and scale your business with ModelsLab's comprehensive suite of AI-powered solutions.
AI Image Generation & Tools
Generate, edit, upscale, and transform images with state-of-the-art AI models.
AI Audio Generation
Text-to-speech, voice cloning, music generation, and audio processing APIs.
AI Video Generation & Tools
Create, edit, and enhance videos with AI-powered generation and transformation tools.
Create Stunning 3D Models
Transform images and text into 3D models with advanced AI-powered generation.