A bar drawn to scale from 0 to 262,144 tokens, the max_position_embeddings in the Clef-flash config.json. Marks show where each route stops: SGLang caps prompts at 16,384 tokens, Workers AI Clef-flash at 24,576 and Workers AI Clef at 65,536.

Self-hosting Cloudflare Clef-flash with SGLang: context, GPU memory and cost

The open Clef-flash weights were trained for a 256k window, but the route you serve them on sets a smaller one. Here is what each route gives you, what it needs, and when it pays.

What does self-hosting Cloudflare Clef-flash with SGLang get you?#

Self-hosting Cloudflare Clef-flash with SGLang runs the Apache-2.0 weights on one GPU, but the SGLang cookbook, read 11 October 2026, caps each prompt at 16,384 tokens. First, Cloudflare put the weights on Hugging Face under the Apache-2.0 license. The Clef-flash model card lists the files: the Qwen3.5-9B backbone, a small joint schema head and the reference Python code. So you can run the model on your own hardware and keep each state inside your network. Then you pay for a GPU by the hour instead of for tokens.

However, the window is the catch. Cloudflare's 9 October 2026 launch post says the weights were "trained to support a 256k context window should you choose to self-host it". Yet the SGLang Clef cookbook is blunt about its own route: "Prompts, including images, are capped at 16,384 tokens". As the cookbook explains, that cap is the default max_length in Cloudflare's reference code.

So there are two routes. SGLang, an open source serving engine, gives you an HTTP server with a /v1/systemone endpoint and that cap. Meanwhile, the reference code in the Hugging Face repo runs on transformers, Hugging Face's Python model library. But no official page reports answer quality past the default max_length, though that code lets you raise it.

Three routesWhere the Clef-flash weights can run and the prompt limit each route sets: SGLang 16,384 tokens, transformers max_length you raise, Workers AI 24,576 (SGLang cookbook and Cloudflare model page, read 11 October 2026)

In short, self-hosting buys control and a fixed bill. On the SGLang route, it does not buy a longer window than Workers AI gives you. If you came for 256k tokens, read the transformers section before you rent a card.

When does self-hosting Clef-flash cost less than Workers AI?#

In this worked example, 1,000,000 decisions a month at about 3,400 tokens each cost $129.20 on hosted Clef-flash and $816 on hosted Clef at Cloudflare's 9 October 2026 prices. The 3,400-token size is the middle row of the latency table in Cloudflare's 9 October 2026 launch post. So it is a size Cloudflare itself tests. The sum is simple: tokens per decision, times decisions, divided by a million, times the price. On the Workers AI pricing page, updated 9 October 2026, Clef-flash costs $0.038 per M input tokens. On the same page, Clef costs $0.24 per M input tokens.

Also, the model writes nothing, so there are no output tokens to pay for. The SGLang cookbook's sample run, read 11 October 2026, prints usage with output_tokens at 0.

Show data table
Monthly hosted bill in US dollars at 3,400 tokens a decision, for 100,000, 1,000,000 and 10,000,000 decisions, arithmetic from Cloudflare Workers AI prices of 9 October 2026
Dimension Clef-flash hosted Clef hosted
100,000 decisions 12.92 81.6
1,000,000 decisions 129.2 816
10,000,000 decisions 1,292 8,160

Hosted Clef costs more than hosted Clef-flash at every volume, and the Clef-flash bill stays small.

Hosted bill by volume Monthly hosted bill in US dollars at 3,400 tokens a decision, for 100,000, 1,000,000 and 10,000,000 decisions, arithmetic from Cloudflare Workers AI prices of 9 October 2026 Arithmetic from Cloudflare Workers AI pricing, 9 October 2026, and the 3,400-token row of Cloudflare's 9 October 2026 launch post. Modelled, not measured

So a self-hosted GPU wins on cost only when its monthly price is below the hosted bill at your volume. For Clef-flash, that bill is small. If you send 10,000,000 decisions a month, Cloudflare's 9 October 2026 prices give $1,292 on Clef-flash and $8,160 on Clef.

Your hosted bill beside your GPU quote

Put in your monthly volume, your state size and the monthly price of the GPU you would rent; the two prices are Cloudflare's of 9 October 2026, and you can change them.

Your traffic and your GPU quote

measured on this projectan assumption, change ityours to enter

Prices per million input tokens, 9 October 2026

Only input tokens are billed, since Clef-flash writes no output. Idle GPU hours, drivers and on-call time are not counted.

Monthly bill on hosted Clef-flash

$129.20

Monthly bill on hosted Clef
$816.00
Your GPU quote minus the hosted Clef-flash bill (below zero, the GPU costs less)
Enter your GPU quote

Arithmetic from Cloudflare Workers AI prices, 9 October 2026. Modelled, not measured.

Then count what the hosted bill hides. A rented GPU bills while it sits idle, while Workers AI bills per token used. Also, someone has to patch drivers, move off a nightly build and answer the pager. In practice, cost alone is a weak reason to move Clef-flash. Instead, the strong reasons are data that must stay on your hardware, or a volume where the hosted bill passes your GPU quote.

Where do the 262,144, 24,576 and 16,384 token limits come from?#

The Clef-flash config.json on Hugging Face prints max_position_embeddings 262,144, Workers AI serves 24,576 tokens, and SGLang trims prompts to 16,384, all read 11 October 2026. The 256k figure is the backbone's position limit and Cloudflare's training claim, while each serving route sets its own smaller cap.

First, the config.json file on Hugging Face, read 11 October 2026, sets max_position_embeddings to 262,144, the 256k in the launch post. That is the most positions the model can encode. But it is not a promise that any server will accept a prompt that long.

Second, Cloudflare cut the hosted window. Its 9 October 2026 launch post says "The hosted version of Clef-flash now has a context window of 24k". That was down from 64k, a trade made to cut the price. The Clef-flash model page, read 11 October 2026, prints the exact figure of 24,576 tokens. Meanwhile, the Clef model page, read the same day, lists 65,536 tokens for the larger model.

Third, SGLang trims at 16,384 tokens, per its Clef cookbook read 11 October 2026. The cookbook adds that the real limit can be lower. In particular, it depends on the server context, the KV token pool and the --max-prefill-tokens flag. When a state runs long, the server cuts the state to leave room for the questions and images.

Show data table
Four context limits a self-hoster meets, and only one is 256k, in tokens, from Cloudflare model pages, the Clef-flash config.json and the SGLang Clef cookbook, read 11 October 2026
Item Value
SGLang /v1/systemone prompt cap 16,384
Workers AI Clef-flash 24,576
Workers AI Clef 65,536
Clef-flash config.json max_position_embeddings 262,144

Only the config.json maximum reaches 256k; every serving route stops at 65,536 tokens or fewer.

Context limit by route Four context limits a self-hoster meets, and only one is 256k, in tokens, from Cloudflare model pages, the Clef-flash config.json and the SGLang Clef cookbook, read 11 October 2026 Cloudflare Workers AI model pages, the Clef-flash config.json on Hugging Face and the SGLang Clef cookbook, read 11 October 2026

So the limits climb from SGLang, to hosted Clef-flash, to hosted Clef, to the config maximum. Only the last one is 256k. So the Clef-flash 256k context window lives in config.json, and no serving route on an official page enforces it.

How much GPU memory does Clef-flash need?#

By arithmetic from config.json, Clef-flash's 8 full-attention layers store 32,768 bytes of KV cache per token, so a 262,144-token window adds 8 GiB beside the weights. The hybrid attention layout keeps the KV cache small, so the weights and the working memory set the GPU size, not the cache.

Here is the sum, all of it arithmetic from config.json, read 11 October 2026. The model has 32 layers, and every fourth one uses full attention, so 8 layers keep a KV cache. Per that config.json, each of those stores a key and a value for 4 KV heads of 256 numbers. Each number takes 2 bytes in bfloat16. That gives 8 times 2 times 4 times 256 times 2, or 32,768 bytes per token. The other 24 layers use linear attention and keep a fixed-size state, so they sit outside the per-token sum.

8 x 2 x 4 x 2562 B32,768 B

Then add the weights. The Hugging Face model card, read 11 October 2026, lists 9B parameters in BF16. At 2 bytes each, that is 18 GB before anything runs. For scale, NVIDIA's H200 page, read the same day, lists 141 GB of HBM3e memory.

Show data table
Weights and full-window cache by arithmetic, beside one H200, in decimal GB; activations and engine pools not included (Clef-flash model card and config.json, NVIDIA H200 page, read 11 October 2026)
Item Value
Clef-flash BF16 weights (9B x 2 bytes) 18
KV cache at 262,144 tokens 8.59
NVIDIA H200 memory 141

The weights and a full-window cache together use a small share of one H200.

Memory by arithmetic Weights and full-window cache by arithmetic, beside one H200, in decimal GB; activations and engine pools not included (Clef-flash model card and config.json, NVIDIA H200 page, read 11 October 2026) Arithmetic from the Clef-flash model card and config.json on Hugging Face, and NVIDIA's H200 page, read 11 October 2026. Modelled, not measured

However, that arithmetic leaves out activations, the vision encoder's working memory and the serving engine's own pools. Instead, Cloudflare's own figure includes them. Michelle Chen is a product manager in Cloudflare's AI Platform group. She told The Register in October 2026 that Clef-flash "will run on any GPU with at least 41 GB of VRAM". Still, she added that this assumes one request at a time and a 64k window.

So for Clef-flash GPU memory requirements, plan on 41 GB and treat the 18 GB of weights as the floor. By the same config.json arithmetic, a 64k window holds only 2 GiB of KV cache. So most of the gap is the engine and its working memory, not the cache. For reference, every accuracy run in the SGLang cookbook used a single H200, B200 or B300.

Which SGLang version serves Clef-flash, and how do you launch it?#

Clef support merged into SGLang on 9 October 2026, after the v0.5.21 release, so the SGLang cookbook pins a 0.5.22 nightly wheel or the lmsysorg/sglang:dev-clef image. Until SGLang 0.5.22 ships, the supported path is that nightly or the Docker image, launched with only the model path, host and port at TP=1. TP=1 means one GPU, with no tensor parallel split.

The SGLang releases page, read 11 October 2026, lists v0.5.21 as the newest release. Cloudflare's 9 October 2026 launch post says pull request #42721 "will be released in SGLang 0.5.22". To run Clef-flash on SGLang until then, build from these pages in this order.

  1. SGLang Clef cookbook: the pinned nightly wheel, the lmsysorg/sglang:dev-clef image, TP=1, and the prompt cap.
  2. SGLang install guide: the other install methods, if pip and Docker do not fit your host.
  3. Clef-flash model card: the docker run launch command and the /v1/systemone route.
  4. Transformers installation: the Python 3.10 and PyTorch 2.5 floor for the reference-code route.
  5. SGLang releases: where 0.5.22 will appear, so you can move off the nightly.
Build orderThe official pages in build order (SGLang docs, Hugging Face and GitHub, read 11 October 2026)

The model card on Hugging Face prints this command for one H200. It runs the dev-clef Docker image, a packaged copy of the server, and it mounts your Hugging Face cache:

bash
docker run --gpus all \
  --shm-size 32g \
  -p 30000:30000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --ipc=host \
  lmsysorg/sglang:dev-clef \
  sglang serve \
  --model-path Cloudflare/clef-flash \
  --host 0.0.0.0 \
  --port 30000

No access token is needed, because the cookbook says both checkpoints are public. Also, there is nothing to tune. Per the cookbook, SGLang finds joint_head_config.json, loads the head and turns on embedding mode by itself. It also turns off radix caching, chunked prefill and CUDA graphs, and it needs no chat template. For the pip route instead, the cookbook's wheel targets Python 3.10 on Linux x86_64 with glibc 2.39 or newer.

How do you know the server scores the way Cloudflare's does? The SGLang cookbook, read 11 October 2026, publishes its own check, GSM8K accuracy through /v1/systemone on three GPUs. In those runs, Clef Flash lands within 0.06 points of the reported 67.3%.

The served models match their reported scores on three GPUs, all BF16 on one GPU with no tensor parallel split: GSM8K accuracy in percent through /v1/systemone. Source: SGLang Clef cookbook, read 11 October 2026; runs by the SGLang team, not ours

ModelReportedH200B200B300
Clef80.880.7180.5980.59
Clef Flash67.367.3267.3667.36

Those are the SGLang team's runs, not ours. Still, they show that the served model matches the reported score on H200, B200 and B300.

What does a systemone() request to self-hosted Clef-flash look like?#

A self-hosted Clef-flash answers POST /v1/systemone with a model, a state and up to 64 typed questions, the same three fields the Workers AI Clef-flash page documents. One JSON body with model, state and questions works on both routes once the model value and base URL change.

The name systemone() comes from the reference function in Cloudflare's joint_schema_model.py. The model card says it takes "a Jev/SystemOne POST /v1/systemone request body" and returns the same response body. Per the SGLang cookbook, System One is the format TypeSafe SDK clients already send, so those clients can point at your server.

Below, step 1 is the launch command above. Then step 2 sends a body to your server, and step 3 sends the same body to Workers AI. The question block comes from the Workers AI Clef-flash page:

python
import os
import requests

# Step 1: start the server with the docker run command above.
SELF_HOSTED = "http://127.0.0.1:30000/v1/systemone"  # the server from step 1
WORKERS_AI = (
    "https://api.cloudflare.com/client/v4/accounts/"
    f"{os.environ['CLOUDFLARE_ACCOUNT_ID']}/ai/run/@cf/cloudflare/clef-flash"
)

# Step 2: one body with model, state and questions, sent to your server.
body = {
    "model": "Cloudflare/clef-flash",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
        "urgent": {"type": "noul", "instructions": "Is this support request urgent?"},
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "billing": "Payments, invoices, and refunds",
                "technical": "Outages, errors, and configuration",
                "sales": "Plans and upgrades",
            },
        },
    },
}
local = requests.post(SELF_HOSTED, json=body, timeout=60)
local.raise_for_status()
print("self-hosted:", local.json()["answers"])

# Step 3: the same body on Workers AI; only the model value and the URL change.
token = os.environ["CLOUDFLARE_AUTH_TOKEN"]
hosted = requests.post(
    WORKERS_AI,
    json={**body, "model": "clef-flash"},
    headers={"Authorization": " ".join(("Bearer", token))},
    timeout=60,
)
print("Workers AI:", hosted.json())

Two things differ, and both come from the official pages. First, the Workers AI page allows only "clef" or "clef-flash" in the model field. The cookbook's example, by contrast, uses the checkpoint name "Cloudflare/clef-flash". Second, the URL and the API token change. Yet the questions, their types and the answer shape stay the same.

Also, every question rides in one pass. SGLang pull request #42721 puts it plainly: "Nothing is generated, so a request with 64 questions is a single prefill". So 64 questions cost one read of the state, not 64 reads.

Images differ a little too. The Workers AI page takes at most 4 embedded images and refuses remote URLs. However, the cookbook's server also accepts an HTTP(S) image URL. Confidence can differ as well, since the cookbook says SGLang computes it with the System One adapter formulas rather than the reference systemone().

Can the transformers route reach the full 256k window?#

The reference transformers code bounds input with encode_record max_length, 16,384 tokens by default, and raising it is the only documented path past the SGLang cap. The long window is reachable that way, but no official page states answer quality above 16,384 tokens.

The model card's line is short: encode_record "accepts max_length (default 16,384 tokens) and max_state_tokens to bound the input". So you load the model with load_release_model, pass a larger max_length to encode_record, and call the model yourself. The card, read 11 October 2026, says its code was tested with torch 2.11 and transformers 5.10.2 on a single H200.

The memory cost of a bigger window is small. By the arithmetic from config.json, read 11 October 2026, each extra token adds 32,768 bytes of KV cache.

Show data table
KV cache by window in GiB for the 8 full-attention layers in BF16, arithmetic from config.json; activations not included (Clef-flash config.json on Hugging Face, read 11 October 2026)
Item Value
16,384 tokens (SGLang cap) 0.5
24,576 tokens (hosted Clef-flash) 0.75
65,536 tokens (hosted Clef) 2
262,144 tokens (config.json maximum) 8

Each step up in window adds cache in proportion, and the full 262,144 tokens needs 8 GiB.

KV cache by window KV cache by window in GiB for the 8 full-attention layers in BF16, arithmetic from config.json; activations not included (Clef-flash config.json on Hugging Face, read 11 October 2026) Arithmetic from the Clef-flash config.json on Hugging Face, read 11 October 2026. Modelled, not measured

So the full 262,144 tokens adds 8 GiB of cache, and 65,536 tokens adds 2 GiB. Activations grow with the prompt too, and no official page gives that figure.

The open question is quality. Cloudflare's 9 October 2026 launch post says the weights were "trained to support a 256k context window should you choose to self-host it". Yet its published latency table stops at about 16,000 tokens. Also, the model card's benchmark tables give no window length. Therefore, test answer quality on your own long states before real decisions depend on them.

There is also no server on this route. The reference code is a library, so you write the HTTP layer around systemone() yourself.

When is self-hosting Clef-flash the wrong tool?#

Self-hosting is the wrong tool when states fit in 24,576 tokens and volume is modest, because hosted Clef-flash then needs no GPU; states up to 65,536 tokens fit hosted Clef. Stay on Workers AI unless data must stay on your hardware, volume is high enough that a GPU costs less than the hosted bill, or states exceed 65,536 tokens.

First, say your states fit in the 24,576 tokens the Workers AI Clef-flash page listed on 11 October 2026. Then hosted Clef-flash already reads more per prompt than the SGLang route does. At the 9 October 2026 price Cloudflare lists, a million 3,400-token decisions cost $129.20, with no card to patch. Cloudflare says in its 9 October 2026 launch post that only 0.24% of its requests exceed 24k input tokens.

Second, if your states run past 24,576 tokens but stay under the 65,536 on the Workers AI Clef page, use hosted Clef. The same post says Cloudflare would "encourage folks with larger context needs to switch to Clef instead of Clef-flash". At its 9 October 2026 price of $0.24 per million input tokens, that costs more per token but needs no hardware.

Third, if you need written text, such as a reply or a summary, no Clef model fits. The model card says "There is no free-form text generation and no output parsing", so a text model is the better pick. Our post on Jev and large language models covers that split.

In short, self-hosting earns its place in three cases. A state must not leave your network. Or a long state needs the transformers route, which you will test yourself. Or your GPU quote sits below the hosted bill.

When to stay on Workers AI or pick another tool, and what to use instead. Source: Cloudflare Workers AI model pages and launch post, 9 and 11 October 2026; Clef-flash model card

Your caseUse insteadWhy
Your states fit in 24,576 tokens and volume is modestHosted Clef-flash$0.038 per M input tokens, no card to patch
Your states run past 24,576 tokens but stay under 65,536Hosted Clef$0.24 per million input tokens, no hardware
You need written text, such as a reply or a summaryA text modelThere is no free-form text generation

Where should you go next with self-hosting Cloudflare Clef-flash with SGLang?#

Pick the hosted Clef model first, then self-host only the decisions that need it, using the comparison of Clef, Clef-flash and Clef-omni as the starting point. The hosted comparison, the System One background and two service pages are the next steps.

Our comparison of Cloudflare Clef, Clef-flash and Clef-omni picks a hosted model by input type, state size and price. For the idea behind typed decisions, read why a typed decision model beats a large language model on cost. For the business case, read why small AI calls on the page pay off.

If you would like a hand, our teams do Cloudflare development and DevOps for running your own GPU servers. That said, the SGLang cookbook and the Clef-flash model card are enough to run this on your own.

Questions this post answers

What context window does self-hosting Cloudflare Clef-flash with SGLang give you?
Per the SGLang Clef cookbook, read 11 October 2026, prompts on that route stop at 16,384 tokens. That is below the 24,576 tokens of hosted Clef-flash on Cloudflare's model page, read the same day.
How much GPU memory does Clef-flash need?
Michelle Chen is a product manager in Cloudflare's AI Platform group. She told The Register in October 2026 that Clef-flash "will run on any GPU with at least 41 GB of VRAM". Still, she added that this assumes one request at a time and a 64k window. So for Clef-flash GPU memory requirements, plan on 41 GB and treat the 18 GB of weights as the floor.
When does self-hosting Clef-flash cost less than Workers AI?
So a self-hosted GPU wins on cost only when its monthly price is below the hosted bill at your volume. In this worked example, 1,000,000 decisions a month at about 3,400 tokens each cost $129.20 on hosted Clef-flash and $816 on hosted Clef at Cloudflare's 9 October 2026 prices.

Keep reading