AI news, with the context that matters.

Technology & Research · ·

TensorFold speeds up local AI generation by drafting and verifying tokens

TensorFold runs models on Macs and supported NVIDIA GPUs. We explain draft verification, quantization, memory, and model support. Project measurements show faster single-request generation but lower throughput than vLLM at eight concurrent requests, so the workload changes the verdict.

The same AI model can take different amounts of time and memory depending on the software that runs it. TensorFold is an inference runtime that predicts upcoming text, checks those predictions with the main model, and accepts them only after verification. It lets applications call local models on supported Macs and NVIDIA GPUs.

The speed benefit depends on the workload. In the project's CUDA measurements, TensorFold generated faster than vLLM for one request, but produced fewer tokens in total with eight concurrent requests. Its design requirement of “exact” decoding also does not promise the same answers as the model before quantization. The mechanism, hardware, and comparison conditions need separate attention.

This technical explainer examines TensorFold v0.6.1, checked on October 2, 2026. It covers ashhart/TensorFold, a separate project from tensorfold.ai. LATENT has not installed the models or measured their performance.

What the runtime does / Why drafting can be faster / What exactness covers / Hardware and models / Memory requirements / Performance under concurrency / Installation prerequisites / Licenses / Sources

TensorFold runs models that have already been trained

An AI model stores the numerical values learned during training as “weights.” An inference runtime loads those weights, processes the input, and computes the next output. TensorFold does not itself contain a newly trained model with new knowledge.

TensorFold uses MLX on Apple silicon and CUDA on supported NVIDIA GPUs. These are software foundations for GPU computation. TensorFold implements computation and draft verification tailored to each model architecture.

Applications connect through an OpenAI-compatible API. It provides chat completions, text completions, and a Responses API endpoint, with a default client base URL of http://127.0.0.1:8080/v1. Compatible clients can call a local model by setting the endpoint and model ID. Compatibility describes the API format; it does not mean reproducing OpenAI's models or every service feature. API specification

Correct drafts let the main model verify several tokens at once

Language models work with units called tokens. A token can represent a word, part of a word, punctuation, or another symbol. It does not always correspond to one Japanese character or one English word. Ordinary generation selects the next token from the text so far, then repeats that process.

Speculative decoding uses a cheaper process to propose upcoming tokens and checks several candidates together with the main model. TensorFold accepts only candidates that match the token the main model would select under those execution conditions. It discards the path beyond a wrong prediction and continues from the confirmed position.

MTP stands for Multi-Token Prediction, a mechanism for predicting multiple upcoming tokens. TensorFold can use a model's included MTP prediction module or a separately supplied drafter such as DFlash2. It can also propose continuations by copying from the existing context. The available method depends on the model.

This is faster when checking several candidates together costs less than running the main model repeatedly for one token at a time. But drafting and verification also take time. Few accepted candidates mean a smaller benefit, and using the GPU's resources for other requests can sometimes produce more output overall. Design requirements and measurement method

Exactness preserves the token sequence under the same conditions

TensorFold defines exactness within the same weights, runtime, and settings. It compares speculative generation with generation one token at a time, a resumed prompt with the same prompt processed fresh, and a concurrent request with its solo run. Where concurrency is supported, it aims to prevent changes in arithmetic order or rounding from changing the output.

Choosing the most probable token is called greedy decoding; selecting according to probabilities is called sampling. In speculative decoding generally, preserving a distribution means preserving each candidate's probability, which is different from producing the same text on every run. TensorFold's documentation goes further: it keys selection by the seed, absolute position, and token ID so that sampling also produces the same token sequence with and without drafting under the same conditions. If no seed is supplied, it derives one from the input. Evaluation conditions

This equality does not extend across MLX and CUDA, quantization formats, or the number of ranks used to divide computation across GPUs. Matching another inference implementation requires separate validation. Quantization represents weights or other values with fewer bits and can round the original numbers. Agreement between speculative and serial decoding on quantized weights does not establish agreement with the original higher-precision model.

The CUDA document's Qwen3.8-27B NVFP4 measurements illustrate this distinction. With the same stored weights, the default --precision checkpoint mode had 92.9% top-1 agreement with an internal fp32 reference, compared with 98.9% for --precision full. The latter uses bf16 activations, the intermediate values in the computation. Both modes maintain agreement between their own speculative and serial paths.

Top-1 here is the share of 4,095 positions in each of eight sequences where the most probable next token matches the reference. The sequences comprise four wikitext-2 text samples and four CPython code samples, and the fp32 reference uses the same stored weights. The 92.9% figure is not an answer accuracy score, and the difference between modes is not a percentage loss in overall answer quality. Precision measurement conditions

Hardware and weight-format support depend on the model

Python 3.11 or newer is required. On a Mac, the v0.6.1 package specifies MLX at least 0.32.2 and below 0.32.4. CUDA requires compute capability 8.9 or newer, a designation of GPU hardware features, covering RTX 40-series, Hopper, Blackwell, and related hardware. RTX 30-series GPUs are rejected at startup.

Having an implementation is different from having tested it on the target hardware. The NVFP4 path was measured on an RTX PRO 6000 Blackwell Max-Q and a DGX Spark. Builds for SM 8.9, 9.0, and 10.0 were compiled and their arithmetic checked on Blackwell, but the documentation says they have not run on those target cards. CUDA implementation and scope

Representative modelDocumented environmentWhat to check before choosing
Qwen3.8-27BMLX and CUDAOn CUDA, supply DFlash2 or select serial decoding with --no-drafts. NVFP4 uses one rank.
Qwen3.8 Flash NextMLX and CUDACheck the weight format and whether MTP is included. Two-rank CUDA rejects concurrent generation.
Nemotron 3.5 LightningMLX and CUDACUDA uses 4-bit/group-64 weights. Explicitly set --no-drafts when not using MTP.
GLM-5.3-FlashMLX on a 256 GB Mac or two-rank CUDACUDA serves requests one at a time. Its optional DFlash2 has separate usage terms.
Gemma 4 26B-A4BMLXLimited to specified weight layouts, including 4-bit groups of 32 or 64.
DeepSeek-V4-FlashMLX on a 256 GB MacCan use its dedicated DSpark or MTP module. This release has no CUDA engine for this family.

These are examples; models such as Ternary Bonsai 2 have their own support. TensorFold cannot load arbitrary Hugging Face models and does not read GGUF. Check the converted weight format and compatible drafter in the support table, rather than relying on the model name alone.

Native Windows support is experimental. The documentation explicitly says it has not yet served a request on a real Windows PC and has not been run under WSL2 either. Documented Windows instructions should not be read as confirmed working support. Windows validation status

Fitting the weights does not guarantee enough memory for a conversation

A 4-bit value uses one quarter as many bits as a 16-bit value. Quantization can therefore shrink weights, but there are also scale values and components retained at higher precision. TensorFold's Qwen3.8-27B model card lists a weight size of 16.06 GB, or 14.95 GiB. That is not the total memory the computer needs.

Execution also needs the KV cache and other model-specific state used to refer to previous tokens without recomputing them, the drafter, and temporary workspace. Longer conversations and more concurrent requests increase these additional demands. Space must remain for the OS and other applications. With a discrete GPU, budget its dedicated VRAM separately from the computer's system RAM. On a Mac with unified memory, the CPU and GPU share physical memory; their capacity cannot be added as separate pools.

TensorFold reads quantized weights and manages conversation reuse and request admission within a memory budget. Flash Next on CUDA also offers optional int8 or int4 KV caches to reduce storage. Cache quantization changes the computation conditions too; it does not promise the same output as a higher-precision cache.

For a larger model, the DeepSeek-V4-Flash MLX recipe lists about 151 GiB of resident weights. On a 256 GB M3 Ultra using DSpark, measurements with 64k- and 128k-token prompts reached a server peak of 165 GiB. This is a project measurement under particular conditions, not a capacity requirement guaranteed to cover every length or concurrency level. DeepSeek memory and measurement conditions

The README's memory-class table still contains many TBD entries, meaning results remain undetermined. Its 64 GB figures come from older measurements with TensorFold 0.3.5.1 and MLX 0.31.2, not revalidation on v0.6.1. MLX-VLM or oMLX measurements in Hugging Face cards must likewise be read as results for those runtimes. Memory table and qualification gaps

Single-request speed does not guarantee higher concurrent throughput

The v0.6.1 CUDA recipe includes a project-author comparison on an RTX PRO 6000 Blackwell Max-Q capped at 250 W. The model is nvidia/Qwen3.8-27B-NVFP4. It compares TensorFold with DFlash2 against vLLM with MTP=3 using the same weights and a 34,816-token context window. Replies contain 256 tokens, and decode rates are the median of two passes.

The table reports tokens per second. Each cell lists code / chat, and the four- and eight-request rows are aggregate rates. The TensorFold column uses default checkpoint arithmetic: FP4 in NVFP4 layers and FP8 in FP8 layers on Blackwell. The original table also includes bf16 full measurements, but those are not used in the ratios below.

Concurrent requests and selectionTensorFoldvLLMTensorFold / vLLM
1 request greedy272.8 / 182.9151.4 / 130.61.80 / 1.40
1 request sampling274.1 / 162.2134.2 / 115.72.04 / 1.40
4 requests greedy675.8 / 466.5557.3 / 451.71.21 / 1.03
4 requests sampling654.9 / 420.5535.3 / 422.61.22 / 0.99
8 requests greedy858.0 / 656.81,009.5 / 907.50.85 / 0.72
8 requests sampling834.5 / 584.1896.8 / 730.60.93 / 0.80

The ratios are 1.40–2.04x for one request, 0.99–1.22x for four, and 0.72–0.93x for eight. TensorFold's own aggregate output increases with concurrency, but vLLM produces more at eight requests. One person waiting for an answer and a service returning as much output as possible to several users are different performance questions.

Cold prefill, processing an input for the first time, is another separate phase. For 2k-, 8k-, 16k-, and 32k-token prompts, TensorFold reports 6,990, 7,484, 7,100, and 6,296 tokens per second. vLLM reports 7,390, 7,804, 7,350, and 6,491, giving ratios of 0.95–0.97x. These are described as final-build measurements with three prompts of each length. Faster decoding does not imply a uniformly shorter wait for the first token.

vLLM's GPU memory utilization setting was 0.29, increased to 0.32 for eight requests. The comparison table does not specify the vLLM version or exact input token counts for the decode tests. The public benchmark procedure measures generation after the first token, so that rate is distinct from total elapsed time including input processing and queues. All figures above are project reports, not independent LATENT measurements. Original comparison table

Check the model format and free capacity before installation

The Mac instructions install TensorFold in a Python virtual environment. CUDA instructions offer NVIDIA's PyTorch container or, for supported RTX hardware, a virtual environment with PyTorch and a CUDA compiler. The compiler must match PyTorch's CUDA version. There is no tensorfold[cuda] installation extra.

When choosing a model, tensorfold models lists candidates and tensorfold info MODEL checks the configuration. info does not fetch weights. In contrast, pull downloads weights, and serve fetches anything missing. First check the model card's file size, compatible drafter, and usage terms.

After startup, inspect the reported context capacity and check /health and /v1/models for server status and the model ID before connecting a client. Starting with one short conversation separates the ability to load a model from the ability to handle long contexts or concurrent requests. The v0.6.1 runbook provides the settings and procedures.

The runtime and each set of model weights have separate terms

The TensorFold runtime uses Apache-2.0 from v0.6.0; releases through v0.5.0 used MIT. The MIT notice is retained for code written before v0.6.0. The runtime ships no model weights, and each model and drafter retains its own license. Third-party notices

In particular, the optional GLM drafter incoai/GLM-5.3-Flash-DFlash2 lists CC BY-NC-ND 4.0 in its model card. Its noncommercial and no-sharing-of-adaptations conditions mean that reading only the runtime's Apache-2.0 license is not enough to establish permission for commercial use. The relevant model card

TensorFold's practical appeal is serving an API on compatible hardware you control, with the potential to reduce waiting through drafting and conversation reuse. To assess it, first check whether the weight format you want appears in the support table. Then look for measurements close to your input lengths and concurrency, and check memory and license requirements. Peak single-request speed alone cannot settle that decision.

Sources and scope of verification

The GitHub sources are pinned to v0.6.1 commit 17c73e189f5e6a5304cda7ea37f086f9c49b4788. The GitHub API records publication at October 1, 2026, 20:19:24 UTC, or October 2 at 05:19:24 in Japan. This article explains that version; it does not describe the project as newly launched that day.

  1. README and release notes. Supported models, exactness, and the qualification scope of the memory table.
  2. RUNBOOK and pyproject.toml. Installation, Python and MLX requirements, and unverified Windows status.
  3. CUDA recipe and DeepSeek-V4-Flash recipe. Performance figures are cited as project-author measurements.
  4. Recipe design requirements and API specification. Token-sequence equality, sampling, and measurement methods.
  5. LICENSE and third-party notices. Runtime and model terms are considered separately.
  6. TensorFold's Hugging Face catalog. Alongside the Qwen card cited in the body, the Flash Next, GLM, Nemotron, and DeepSeek DSpark and MTP cards were saved and checked. Measurements using other runtimes in those cards are not presented as TensorFold runtime performance.

Remaining uncertainties include untested GPUs and Windows environments, current-release qualification for each memory class, and the comparison's vLLM version. LATENT has not evaluated every model format or context length, or independently reproduced the quality and performance results.