AI news, with the context that matters.

Technology & Research · ·

Strata runs a large MoE model by sharing computation between the CPU and GPU

Strata v0.1.39 speeds up local generation and adds concurrent requests and Codex CLI support. We explain the memory layout behind a large MoE on a 12 GB GPU, RAM and SSD requirements, long-prompt and concurrency trade-offs, and the limits of its Responses API.

A 125-billion-parameter AI model on a graphics card with 12 GB of memory: that is the use case Strata targets. It does not fit the entire model into VRAM. It keeps frequently used expert weights in VRAM for GPU computation and handles much of the rest on the CPU using weights in system RAM. It also stores the model’s PLE table on SSD and retrieves the rows needed to turn short token sequences into numerical inputs for computation. Capacity and speed across the whole PC determine the experience.

Released on October 4, 2026, v0.1.39 improves generation and long-prompt processing and adds concurrent requests plus a Responses API for Codex CLI. It makes a concrete workflow possible: have a local model write code, receive tool results, and continue working. Concurrency does not always make it faster, and the API has limits.

This explainer is based on the v0.1.39 release and documentation and code pinned to its tag. Performance figures are reports by the developer or contributors. LATENT has not downloaded or run the models or tested the Codex connection.

Key takeaways

  1. Strata is an inference runtime for the Qwen3.8-Flash-Next family. It caches frequently used expert weights in VRAM for GPU computation; the CPU handles much of the uncached expert computation using weights in system RAM. Running on a 12 GB GPU does not mean the entire model fits into 12 GB.
  2. In the developer’s RTX 5070 comparison with the previous release, v0.1.39 improved decode speed by 2.5–6%. For four concurrent requests, the last response began after 1.8 seconds instead of 11.2, while aggregate throughput fell by about 11%. Quantization, context length, and cache capacity affect the result.
  3. Codex CLI can connect to the local model through the Responses API and return tool results. The developer tested CLI 0.160.0, but continuation by stored response ID and OpenAI-hosted tools are unsupported. This does not make OpenAI’s Codex models local.

A 125-billion-parameter model uses the whole PC

Strata is not a newly trained AI model. It is an inference runtime that loads learned weights and computes outputs from inputs. Version 0.1.39 focuses on Qwen3.8-Flash-Next and its quantized and derivative versions. It is not a universal engine for arbitrary models.

MoE stands for Mixture of Experts. It routes part of the computation through multiple networks called experts. These are not literal job roles or neatly separated fields of knowledge. A router selects experts based on the input, so each computation can use a fraction of the total stored weights.

The full model has 48 layers and 512 experts per layer, for 24,576 experts in total. Each token uses 10 experts at each layer. A token is a processing unit such as part of a word or a symbol, not necessarily one Japanese character. Reading this as “only 10 experts for the entire model” misses the repeated processing across layers. The model geometry and router implementation show this structure.

Decode and prompt processing use experts differently

During decode, which generates text incrementally, Strata keeps frequently selected expert weights in VRAM. The GPU uses those weights for expert computation and also computes attention, the output head, and other components. Normal mode also retains all expert weights in system RAM. The CPU handles much of the uncached expert computation by reading weights in place from RAM. CPU and GPU work overlap, so every weight missing from VRAM does not have to be copied to the GPU for each step.

The CPU/GPU split is not fixed. Strata changes which expert weights it caches in VRAM as the conversation evolves. It can also send a share of uncached weights to the GPU for computation across PCIe, the connection between the host and GPU. CPU speed, RAM bandwidth, PCIe traffic, and the cache hit rate all affect generation speed.

During decode, GPU and CPU share expert computation

The router selects 10 experts at each layer

GPU and VRAM

Keep frequently used expert weights in VRAM

Compute experts, attention, and other components on the GPU

CPU and system RAM

Normally hold all expert weights in RAM

Use uncached weights for CPU computation

Combine both results and continue processing

Store the PLE table on SSD: about 28.8 GB in IQ4_NL format

Read rows absent from the dedicated cache from SSD and reuse some of the rows already read

A simplified conceptual diagram of the Qwen3.8-Flash-Next family running in a normal single-GPU configuration. Box sizes do not represent capacity. Cache adaptation and the optional PCIe share are described in the text. Low-RAM mode can also read expert weights from SSD.

The Qwen3.8-Flash-Next family has a mechanism called PLE that retrieves numerical vectors for model computation from short token sequences. Its learned embedding table is per_layer_token_embd.weight, which Strata accesses from model files on SSD. It does not store prose or example answers, and PLE is not a requirement of MoE models in general.

The approximately 28.8 GB figure here is for the IQ4_NL PLE table shared by the original model’s Q2_0, IQ2_XS, IQ3_XXS, and IQ3_S versions and Coder. Its 320,001,536 rows occupy 90 bytes each, totaling 28,800,138,240 bytes, or about 28.8 GB in decimal units. Expert quantization names and the PLE table’s storage format are separate; this size does not apply to every table format.

A lookup uses the current token and its two predecessors. The current token and its immediate predecessor form a 2-gram; adding the earlier predecessor forms a 3-gram. Each uses eight heads to calculate row indices, giving 16 row references per token. This retrieves multiple numerical vectors from the same short token sequence.

Strata looks for the required rows in its dedicated row cache and reads cache misses from SSD. It dequantizes the retrieved values into numerical vectors and sends them to the GPU for model computation. Some rows remain in the bounded cache for reuse, so a lookup does not necessarily require an SSD read.

To limit RAM use, the SSD-oriented default, --ple-io direct, does not keep the entire table resident in RAM and bypasses the OS file cache. The dedicated row cache and read buffers still use RAM. This PLE-table behavior also applies in normal mode, which retains all expert weights in RAM. A shorter overview still describes reads through the OS cache; this article follows the explicit setting in DETAILS, the table definition, and the row-index implementation.

Prefill, which processes an input prompt, works differently. Normally using chunks of up to 8,192 tokens, Strata streams the needed expert weights to the GPU for larger matrix operations. Transfers for the next layer overlap processing of the current one. Weights used by the CPU for in-place computation in RAM during decode can be worth sending to the GPU for long prompts because many tokens can reuse the transferred weights.

The KV cache, which reuses past context, consumes memory separately from model weights. At 64K context and above, when supported, Strata can keep much of it in RAM and retain the frequently accessed 32K attention window in VRAM, leaving more space for expert weights. This streaming mode is unavailable under WSL and for some KV formats. See KV-streaming conditions, GPU copies, and resident weight access.

A 12 GB GPU still needs tens of gigabytes of system RAM

The 12 GB figure describes VRAM capacity, not total PC memory. Normal mode also retains all expert weights in RAM, including those cached in VRAM. Adding VRAM therefore does not automatically reduce RAM requirements by the same amount. Quantization represents weights with fewer bits, reducing size while changing numerical representation.

Model and format Expert weights in RAM System RAM guidance
Q2_0 About 34 GB 48 GB or more
IQ2_XS About 35.5 GB 48 GB or more
IQ3_XXS About 43 GB 64GB
IQ3_S About 50 GB 64 GB with little else open
Coder / IQ1_M About 23 GB 32GB

The table follows the v0.1.39 capacity guide. The “Expert weights in RAM” column excludes other weights, the OS, applications, and context state. The guide allows roughly 10 GB beyond the expert weights. Setup’s long-context estimate is more conservative and normally recommends up to 128K for IQ3_XXS and IQ3_S on a 64 GB PC.

Low-RAM mode is available for PCs with little system RAM and a large GPU. It memory-maps model files and chiefly retains expert weights not held in VRAM in system RAM. The documentation gives examples of Q2_0 and IQ2_XS on 32 GB RAM with a 24 GB GPU. With less VRAM, more expert weights come from SSD and performance falls. Simply adding RAM and VRAM capacities is not a sufficient fit test.

Coder is a derivative that retains 256 of each layer’s original 512 experts, not merely a more aggressively compressed copy of the same full model. The documentation warns of weaker performance outside code and in languages other than English, including CJK text. Its ability to fit 32 GB should be weighed separately from suitability for Japanese-language work.

Disk space also needs allowances. The three smaller full-model formats download approximately 66–76 GB, the MTP draft weights add about 6 GB, and image support adds about 1 GB. Q2_0 on an AVX-512 CPU also creates an approximately 40 GB pack for its fast CPU kernels. The README’s “about 80 GB” is not an upper limit for every configuration. See installation requirements.

The larger Unsloth versions have different requirements: UD-IQ4_XS downloads 94 GB and has 59.5 GB of expert weights; experimental UD-Q4_K_XL downloads 111 GB and has about 77 GB of expert weights. Weights beyond the RAM budget are read from SSD during generation. For the latter on 64 GB RAM and a 12 GB GPU, the project reports 7–8.5 tokens/s. Results in the tens of tokens/s for smaller formats cannot be applied to that configuration. See large-format requirements.

The new release gains a few percent in decode speed

Strata uses the model’s MTP module to predict a few upcoming tokens and verifies them together with the main model. Accepted predictions let one expensive main-model pass advance several tokens. Version 0.1.39 reduces GPU launches and host round trips in verification and batches related operations.

Format and output v0.1.38 v0.1.39 Reported gain
Q2_0 story 68.4 tok/s 72.7 tok/s +6%
Q2_0 code 75.8 tok/s 80.0 tok/s +6%
IQ3_XXS story 46.1 tok/s 48.8 tok/s +6%
IQ3_XXS code 50.7 tok/s 51.9 tok/s +2.5%

Source: the release notes. These are medians from 10 interleaved old/new pairs on an RTX 5070. Tok/s measures generated tokens per second; the gain column preserves the release’s rounded percentages. Decode after 4K and 32K prompts reportedly improved by 5–7%. These are not reductions in total job time including prompt processing and queueing.

Not every optimization in the release was measured on a 12 GB GPU. The verification graph that further removes host notifications runs only when all expert weights of a layer are in VRAM. The developer’s 12 GB card never met that condition, so the additional benefit on larger GPUs was not measured.

The README’s figures, such as 94 tokens/s for Q2_0, come from different measurements. Its NVIDIA Q2_0 row uses engine 0.1.36, while the others mainly use 0.1.26; the README speed table was not remeasured for v0.1.39. Comparing 94 with the 80 above across different prompts and settings does not establish a regression.

Long-prompt gains depend on free VRAM and configuration

Long prompts require a balance between temporary space for streamed expert weights and the expert cache that remains in VRAM. Version 0.1.39 sizes the streaming ring by the quantized pack’s bytes and adjusts prefill chunks. The ring is temporary storage repeatedly reused for incoming expert weights.

For an RTX 5070 and a 32K prompt, the release reports 18.5% faster prefill with IQ3_XXS and a 1,500-slot expert cache, but no change with setup’s default configuration. Coder processed a 30K prompt about 8% slower with a 32K context limit and about 6% faster with a 64K limit. These slots store expert weights; they are different from the conversation slots discussed next.

For long prompts, the mix of cached and streamed expert weights changes, so numerical results need not be bit-identical to the previous version. The release includes a teacher-forced comparison against an FP16 prompt path. For IQ3_XXS after a 32K prompt, top-1 agreement over the following 2,001 tokens was 95.3%, compared with 95.0% before. This measures agreement on the most likely next token, not question-answer accuracy.

Version labels also differ within the sources. DETAILS at the same tag calls the ring change “0.1.39b” and describes comparisons against 0.1.39, while the published release presents it as v0.1.39 versus 0.1.38. This article uses the release notes for old/new comparisons and does not combine the additional DETAILS conditions as if they were the same experiment.

The release also fixes unnecessary SSD reads when RAM usage is capped by a budget. Version 0.1.38 chose its read mode too early and compared against all model-file bytes. On a reported 96 GB RAM system using UD-Q4_K_XL with a 72 GiB budget, this mistake made prompts 15–40% slower than 0.1.34. The new version counts expert-weight bytes outside the RAM copy and rechecks after that copy is built. The release says either read mode produces the same tokens.

Concurrency trades waiting time against total throughput

By default, Strata serves one request at a time. A model setting such as "parallel": 4 lets several conversations advance together. A waiting short prompt can also take priority at a chunk boundary during a long prompt. Multiple users therefore need not all wait for the preceding answer to finish.

Four requests sent together Serial processing Four slots
First token of the last request 11.2 s 1.8 s
Aggregate throughput 70.7 tok/s 63.1 tok/s
Per-request decode rate 79.6 tok/s 16.9 tok/s

Conditions come from the BATCHING measurement table: RTX 5070 12 GB, Ryzen 5 7600, 64 GB DDR5, Q2_0, 32K context, and distinct simultaneous HTTP requests asking for an 800-word essay. Each answer was capped at 256 tokens, using greedy decoding with thinking off; results are medians of three rounds. Aggregate throughput and per-request decode rate use different measurement boundaries, so multiplying the latter by four is not a valid comparison.

The last request started much sooner, but aggregate throughput fell by about 11%. Each conversation’s state consumes VRAM and shrinks the expert cache. At 32K context with 8-bit KV, one slot needs 0.56 GiB. A GiB is 2 to the power of 30 bytes; it should not be added to the capacity guide’s GB figures as if the units were identical.

The cost remains when only one request runs with slots reserved. The release reports an 11% single-request slowdown with two slots and 22% with four. BATCHING’s table shows aggregate-throughput losses of 9% and 22%, while its per-request decode figures imply approximately 11% and 24%. Its prose also differs, so the metric and source need to remain explicit.

A batch verification window carries one token from each conversation without MTP drafts. Shared weights can be read together, but conversations selecting different experts increase CPU work. When one request remains, it can return to the normal speculative path, yet reserved session memory still reduces the expert cache. This explains the serial default on a 12 GB GPU.

A larger configuration can behave differently. A contributor’s test used four 16 GB GPUs, PCIe Gen3, IQ3_S, eight slots, and four pipeline groups. Aggregate throughput rose from 120 tokens/s for one request to 360 for eight, with 400 tokens per answer and temperature 0.7. This configuration keeps many expert weights in VRAM and splits layers to process different conversations in parallel; it does not generalize to one 12 GB GPU.

Output equality is conditional too. The exactness tests use STRATA_IQ_MT_MIN=1, --pcie-frac 0, and an effectively fixed expert tier, among other settings. In a four-conversation Q2_0 test with defaults, only one conversation matched its solo output; the others diverged at the first or a later token. Batch windows also do not apply repetition, frequency, or presence penalties. Users should check required generation settings instead of assuming that parallel runs always produce the same text.

Codex CLI connects the local model to a tool loop

Codex CLI is a client that converses with a model and invokes tools such as file operations and shell commands. Strata’s Responses API support lets it direct inference to a local Qwen-family model. The model returns a tool call, the CLI executes it within its configured permissions, and the result goes back to the model. That loop distinguishes one-off code generation from an agent that continues a task.

The model and client exchange decisions and tool results

Codex CLI

Send instructions, history, and tool definitions

→

Strata and the local model

Return text or a tool call

↓ Return the tool call to the CLI

The CLI runs the tool and includes its result in the next input

History is resent; Strata reuses computation for matching prefixes

Conceptual diagram based on the developer’s Codex CLI 0.160.0 test. The model does not directly execute commands. Tool permissions and external communication depend on the client and connected tools.

The core provider settings documented by Strata are shown below. They belong in the user-level ~/.codex/config.toml; the context limit should match the value selected in Strata setup. The value 32,768 is the documentation’s example, not a recommendation for every PC. Sources are Strata’s example and the OpenAI configuration reference.

model = "strata"
model_provider = "strata"
model_context_window = 32768
show_raw_agent_reasoning = true

[model_providers.strata]
name = "Strata (local)"
base_url = "http://127.0.0.1:8080/v1"
wire_api = "responses"
stream_idle_timeout_ms = 600000

The model name strata is a client-side identifier, not an instruction to download an OpenAI model. Strata responds with the model it currently has loaded. If the server has an API key, the documented env_key = "STRATA_API_KEY" setting is also needed.

The developer tested a tool loop with Codex CLI 0.160.0, Q2_0, and an RTX 5070. Initial instructions and tool descriptions totaled 9,443 tokens and took 10 seconds to read. Later turns reused about 96% of the prompt from cache and read the new portion in 1–2 seconds. This does not mean the entire answer finishes in 1–2 seconds or that every workload achieves 96% reuse.

The Responses API is stateless. The client sends the whole conversation in input on every request; Strata does not continue a conversation by looking up a stored response ID. It can still reuse cached prompt computation. Not storing API responses is different from recomputing the entire prompt each time.

Feature v0.1.39 behavior
Functions and tool results Handles function and custom calls and their results
Reasoning effort and streaming Maps effort and returns Responses events
JSON formats json_schema and json_object are validated after generation
previous_response_id Unsupported; returns an error
Hosted tools Leaves out web_search, file_search, and similar tools
Reasoning summaries No summaries are generated; summary is empty

JSON validation is not grammar-constrained generation. Apart from none, tool_choice leaves the decision to the model and cannot force a particular tool. The API also excludes conversation, background execution, and retrieval, deletion, or cancellation of stored responses. The adapter code defines accepted inputs and rejection conditions.

Strata’s encrypted_content needs care when replaying reasoning on later turns. Despite its name, the implementation is a Base64-encoded replay string, not encryption. Local inference also does not automatically make external tools used by the CLI, such as Web or MCP services, operate offline.

Standard hardware paths and experimental ports differ

The standard installation targets Windows 10/11 or Linux with a supported NVIDIA or AMD GPU of at least 12 GB. The normal NVIDIA engine uses CUDA 13 and requires driver 580 or newer. The standard CPU path is x86-64 with AVX2. AMD support depends on the GPU; image input is unsupported on Windows and uses the CPU on Linux.

Version 0.1.39 also provides an experimental CUDA 12.9 engine for Pascal and Volta. It requires driver 528 or newer on Windows or 525 or newer on Linux. The maintainer does not own the older target GPUs and distinguishes build/output checks on other hardware from contributors’ reports on actual devices.

The Intel Arc SYCL port is an experimental Linux source-build path with no prebuilt engine. The release explicitly leaves its current version on Arc hardware, Windows, WSL2, A-series and integrated Arc GPUs, and images unverified by the maintainer. Older CPUs without AVX2 also gain a path, but only for i-quant models and with slow CPU computation. Appearing in a support list is different from being validated on a particular PC.

Before installation, check the GPU, RAM, free disk space, and driver against INSTALL. The documented entry points are START-HERE.bat on Windows and ./setup.sh on Linux. First startup downloads models and dependencies and loads tens of gigabytes into RAM. Testing one short request before the intended context length and concurrency makes capacity problems easier to distinguish from execution trade-offs.

The server listens on 127.0.0.1 by default. Exposing it to other devices changes requirements such as API-key configuration; consult the description of exposed surfaces. The runtime is MIT-licensed, while models and some bundled components have their own terms. Local execution does not give everything the same license.

TensorFold differs in model scope and memory strategy

The previously covered TensorFold v0.6.1 also verifies speculative drafts to accelerate local generation, but the selection criteria differ. TensorFold serves several supported model families through MLX on Macs and CUDA on supported NVIDIA GPUs. Strata focuses on the Qwen3.8-Flash-Next family, keeping expert weights that do not fit in VRAM mainly in system RAM and also performing expert computation on the CPU.

Aspect Strata v0.1.39 TensorFold v0.6.1
Model scope Qwen3.8-Flash-Next and supported derivatives Multiple supported model families
Main backends NVIDIA CUDA and AMD HIP Apple MLX and NVIDIA CUDA
Focus of this comparison Weight caching in VRAM plus CPU computation Model-specific execution including draft verification
What is not established No matched head-to-head speed test No matched head-to-head speed test

This is an architectural comparison based on Strata’s scope and TensorFold’s README. Speeds measured on different models, GPUs, and quantizations cannot determine a winner. In either case, agreement in speculative decoding does not guarantee agreement with unquantized weights or another runtime.

To evaluate Strata, first choose the derivative and quantization and check whether its expert weights fit in RAM. Then consider how context length changes cache capacity and whether the priority is one user’s response rate or less waiting for several users. Codex CLI integration is a workflow built on top; Japanese-language ability, code correctness, and tool selection still require separate evaluation.

Sources and verification scope

GitHub’s API records the v0.1.39 release at 12:32:47 UTC on October 4, 2026, or 21:32:47 that day in Japan. Code and documents are pinned to commit 6f32ec070f23ced9f50e704d854d775da52591ab. Sources were checked on October 5, 2026; this does not mean the Strata project itself first appeared on October 4.

  1. v0.1.39 release and GitHub API: publication time, old/new comparisons, and the maintainer’s verification scope.
  2. README, MODELS, and INSTALL: models, memory and disk guidance, and standard versus experimental environments.
  3. DETAILS and BATCHING: memory tiers, the Codex test, and concurrency measurements and limits.
  4. Responses adapter, API tests, expert sources, expert cache, and prefill implementation: read for verification; LATENT did not execute Strata’s tests or inference.
  5. LICENSE and model and component notices: distinguish the runtime’s terms from those of model weights.

Remaining unknowns include speed and stability on a reader’s PC, quality in Japanese and real tasks, and behavior on untested hardware. Differences in version labels and metrics within the sources are identified in the article. None of the reported figures is treated as a guarantee across all environments.