AI news, with the context that matters.

Columns · ·

How practical is a local LLM you can actually rely on?

What can today’s Mac, DGX Spark, and Windows PCs offer someone who wants to own their AI? A practical look at Japanese quality, published performance evidence, VRAM and RAM, Strata, and what ¥1 million or ¥2 million buys.

There is a reason to want AI on your own computer. Drafts and work documents can stay with you instead of being sent out for every question. A model you like can remain available after a service changes its lineup. Owning AI means more than reducing a monthly bill.

The satisfaction of seeing a local model produce a reply wears off quickly. Can it write Japanese without repeated correction? Can it handle a long document within a tolerable wait? When asked to change code, can it carry the task through to verification? For daily use, these differences matter more than the price of the machine.

As of October 7, 2026, moving a bounded set of tasks onto your own hardware is realistic. Replacing everything you ask a frontier cloud service to do, with the same speed and effort, is a much less straightforward purchase. The ¥1 million and ¥2 million budgets below help expose that difference; they are not spending targets.

Key takeaways

  1. For Japanese drafts and bounded code changes, smaller models and the 27B class are useful candidates. A 16GB GPU constrains quantization and context choices; 32GB provides more room to keep a 27B-class model on the GPU.
  2. Mac and Spark offer large shared memory pools; a Windows GPU workstation offers a different balance. Strata expands the options for large MoE models, but its speed depends on memory bandwidth, quantization, and GPU caching as well as RAM capacity.
  3. Read published benchmarks together with their model, quantization, and input conditions. Judge the initial wait and completion time as well as generation speed, while retaining cloud options for demanding reasoning and long autonomous tasks.

From a model that runs to a tool for work

The meaning of usable depends on the job. Standardizing the wording of an internal document puts a premium on preserving meaning and responding quickly. Discussing a design calls for noticing conflicts between requirements. Both tasks produce fluent text, but they do not demand the same capabilities.

Start with work you repeat each week, rather than the largest model a machine can load: extracting decisions from meeting notes, proposing tests from a short specification, or editing your own notes. When the input and expected result are clear, it is easier to check the answer and decide whether local execution is worthwhile.

An open-ended request with little context makes failures harder to diagnose. You may have to distinguish a model limitation from quantization, missing information, or a tool configuration problem. More expensive hardware does not remove that work.

Japanese, code, and long documents need different checks

Japanese quality involves more than grammatical sentences. Register, omitted subjects, terminology, and the preservation of qualifications all matter. A small model may be useful for a fixed formatting task while demanding more checking when it rewrites a long explanation without changing the meaning. Hardware speed alone cannot solve that problem.

For a starting comparison, consider Gemma 4 12B QAT for lighter tasks and Qwen3.8-27B for a broader workload. Gemma 4 also offers the 26B A4B MoE. A4B describes the active parameter scale during generation, not the storage requirement of a 4B model. Parameter counts alone do not establish which is better at Japanese.

01 / Dense and MoE

Weights you store and weights you use at each step

An FFN transforms information for each token. This sketch shows one layer, with the blocks used for this token highlighted.

Dense

Input for one token

After shared operations such as attention, enter this layer’s feed-forward network (FFN).

↓

Use the full FFN

Each token passes through the same complete FFN.

Store this FFN’s weights and use them in the computation.

MoE

Input for one token

After shared operations such as attention, enter this layer’s feed-forward network (FFN).

↓

A router selects experts

The selection can change for the next token.

↓

Expert 1

Expert 2

Expert 3

Expert 4

Expert 5

Expert 6

Expert 7

Expert 8

Use the two green experts now. Keep the other experts’ weights available for later selections.

Reading “26B A4B”

Gemma 4 26B A4B has roughly 26 billion parameters in total and uses about 4 billion per token. Its stored weights do not shrink to those of a 4B model.

Placement depends on the implementation, including GPU and system RAM. Parameter count and architecture alone do not determine Japanese quality.

Selecting two of eight experts is illustrative, not the actual model topology or memory ratio. Operations and weights outside the FFN are also required. Conceptual diagram based on Google’s Gemma 4 documentation.

Google estimates roughly 6.7GB to load the 4-bit 12B model and 14.4GB for 26B A4B, including 20% loading overhead. Those estimates are not bounds for arbitrary context lengths or multiple users. On a 16GB GPU, the 12B choice leaves more room; the 26B A4B choice needs closer attention to actual memory use, including context.

Gemma 4 variants and memory estimates

Qwen3.8-27B emphasizes coding, professional work, and longer agent workflows. Its reported benchmark results depend on evaluation settings, context, and reasoning configuration. They should not be treated as guaranteed results from a compressed local setup.

Qwen3.8-27B model card and evaluation conditions

Quantization reduces storage by representing weights with fewer bits. Unsloth lists Qwen3.8-27B UD-Q4_K_M at 16.5GB and UD-Q6_K at 22GB. A 16.5GB file being close to a GPU’s advertised 16GB is not evidence of a comfortable fit. Unit conventions differ, and the runtime also needs context state such as the KV cache and working memory.

02 / Model size and quantization

The same 27B model can have different file sizes

Distinguish how many parameters a model has from how their values are stored.

Qwen3.8-27B

Choose the storage precision for the same model of roughly 27 billion parameters.

UD-Q4_K_M

16.5GB ≈ 15.4GiB

UD-Q6_K

22GB ≈ 20.5GiB

The bars show published file sizes. Q4 and Q6 identify quantization formats; the files are not simply 27B × 4 bits or 27B × 6 bits.

Changing precision changes storage needs and can affect outputs. Check the answers on your own tasks when choosing a setting.

Converted from rounded sizes in Unsloth’s file listing. 1GB = 10⁹ bytes; 1GiB = 2³⁰ bytes. Bar lengths have a 16.5:22 ratio. These are file sizes, not measured runtime memory or performance; KV cache and working memory are additional.

03 / Inference memory budget

Loading the weights is only part of the memory budget

Budget within each memory pool: GPU VRAM or a unified pool such as a Mac’s shared memory.

① Loaded weights

The allocation made by the runtime can differ from the downloaded file size.

② KV cache

Stores context state. Longer contexts and concurrent requests generally need more space.

③ Working memory

Temporary buffers and other workspace depend on the runtime, execution method, and added features.

④ Other uses and headroom

A unified pool also serves the OS and applications. Display output and other GPU tasks can consume VRAM.

Fit ① + ② + ③ + ④ within the capacity of the same memory pool.

Check free space before extending context

Weights can load successfully while leaving too little space for the subsequent KV cache.

Consider 32GB VRAM for a daily 27B model

The article’s Q6-class, 32K-context configuration is a starting point to verify, not a guarantee for arbitrary extras or concurrency.

Keep the memory pools distinct

Adding system RAM does not expand VRAM. Moving from a 4TB to an 8TB SSD does not expand these working memory pools.

Conceptual planning diagram. Box areas are not capacity ratios, and no unmeasured GB allocations are assigned. With offloading, budget GPU and CPU memory separately. KV requirements also depend on model architecture and cache format.

This gives 32GB VRAM a concrete role if a 27B-class model will be your daily tool. Q6-class weights with around 32K context provide a reasonable starting configuration to verify, not a guarantee after adding vision components, speculative decoding, or concurrent requests. With 16GB, first decide whether to accept a smaller model, stronger compression, or partial CPU offload.

Actual quantized files and sizes

Writing one function and investigating a repository to change several files are different jobs. The latter also requires software for search, editing, testing, and recovery from failures. Connecting a model API does not reproduce a cloud coding agent’s complete working environment. Small changes with existing tests make a clearer starting point.

For long documents, accepted input length and reliably used information are different limits. A supported context window does not eliminate omissions. Selecting relevant chapters or files, or retrieving passages with sources, can reduce prompt-processing time and make the answer easier to verify.

How Strata uses VRAM and system RAM

A Windows PC has two relevant capacities: GPU VRAM and system RAM. A model and its working state that fit on the GPU can be processed there. Offloading parts to RAM introduces CPU work or transfers over PCIe. Expanding system RAM from 64GB to 128GB does not turn a 16GB GPU into a 128GB GPU.

Strata is interesting because it designs that division of work around a large mixture-of-experts model. An MoE selects some experts for each input instead of applying every expert’s weights equally at every step. Strata caches frequently used experts on the GPU and keeps others in RAM for CPU work. That architecture and placement allow it to handle a model larger than GPU memory.

Upstream Strata environments and inference approach

The main target is Qwen3.8-Flash-Next, not arbitrary large models. Its 125B core is accompanied by 51B of per-layer embeddings and a 4B multi-token-prediction component. PLE provides embedding lookups; MTP supports predicting subsequent tokens. Neither the 125B label alone nor a simple sum of these counts determines required RAM.

Qwen3.8-Flash-Next architecture

Strata’s model guide lists a 54.8GB main component for IQ3_S, with about 50GB of experts, plus a separate roughly 28.8GB PLE file kept on SSD. A 64GB system can fit it with little room for other applications. Moving to 128GB provides room for less aggressive quantization and daily applications; UD-IQ4_XS involves roughly a 94GB download and 59.5GB of experts.

Keep two uses of SSD storage separate: sparse PLE lookups and frequent expert-weight reads caused by insufficient RAM. The latter slows inference. Moving from a 4TB to an 8TB SSD primarily increases storage for models and documents; it does not resolve a RAM shortage.

Strata quantization sizes and low-RAM behavior

04 / Where the model lives

Adding capacity does not make all memory equally fast

Distinguish where weights are stored from which processor computes.

1

Conventional GPU inference

Windows PC / NVIDIA GPU

GPU · VRAM

Model weights + KV cache + working memory

Process on the GPU when everything fits

PCIe

CPU · RAM

Some configurations keep offloaded weights here and compute on the CPU

More system RAM is not more VRAM

2

Strata divides work between CPU and GPU

A specialized path for supported large MoE models

GPU · VRAM

Attention and other operations, plus a cache of frequently used experts

Cache hit rate also affects speed

Shared work

CPU · RAM

Store expert weights and compute work outside the GPU cache

Both capacity and memory bandwidth matter

SSD · Look up required PLE entries

Sparse PLE reads differ from repeatedly reading expert weights from SSD because RAM is short. The latter slows inference.

3

CPU and GPU share memory capacity

Mac Studio / DGX Spark・OEM

CPUGPU

Unified memory

Model weights, KV cache, OS, and applications use one memory pool.

Mac uses MLX or Metal; Spark uses CUDA. Their execution software differs.

LATENT’s conceptual placement diagram. Areas are not proportional to capacity or performance. The KV cache stores context state; its size depends on context length and concurrency.

The execution platform is another boundary. Upstream Strata’s usual installation targets are x86 Windows and Linux; GB10 support comes through a separate Arm64 fork. This is not the same installation path on a Mac. Mac users instead consider MLX or llama.cpp with Metal, so Strata speed reports cannot be transferred directly into a Mac purchase argument.

Operations are still evolving. Version 0.1.40.1, released October 6, fixed quoted tool calls being executed and requests hanging during engine restarts. Those are concrete reasons to check more than throughput, including whether displayed examples remain text rather than actions. For file-changing tasks, use a recoverable working area and explicit review points.

Strata v0.1.40.1 fixes

What published benchmarks tell us about waiting time

When reading a benchmark, separate the wait for the answer to start, the rate at which text appears, and the time to finish. Fast generation can still follow a long wait for input processing or hidden reasoning.

05 / Response timing

Starting quickly and finishing the answer are different

From submission to completion. Box lengths are not proportional to time.

A

Read the input

Model loading, queueing, and prompt processing. Longer inputs can increase this wait.

No answer is visible yet
B

Generate reasoning

Additional time when reasoning is enabled. Internal tokens may appear before the answer intended for the reader begins.

Depends on settings and the task
C

Display the answer

Distinguish how quickly the text appears from how long it takes to finish.

Visible answer

Time to first visible answer

Submission → start of C

Includes the waits in A and B.

After the answer starts

How quickly text appears

The rate of visible answer text, separate from internal reasoning.

Time to completion

Submission → end of C

Judge the wait on answers representative of your work.

LATENT’s conceptual diagram of response timing, not measured device performance. The first API token need not be the first readable answer.

A published Strata RTX 5090 test illustrates the distinction. It used Linux, 64GB RAM, v0.1.29, and IQ2_XS on September 30, 2026. Reasoning was disabled, with synthetic code-explanation inputs and 256-token outputs. The following values are medians of three runs per condition.

Input tokens Decode rate First API text delta (TTFT) Time to 256-token completion
4,096 179.4tok/s 0.974 s 2.392 s
32,768 175.7tok/s 5.950 s 7.390 s
128,000 165.0tok/s 22.262 s 23.850 s

Decode rates differ relatively little, but TTFT exceeds 22 seconds on the longest input. Here TTFT runs from the HTTP request to the first nonempty text delta, rather than the first screen display. Model loading is excluded, and GPU expert-cache hit rates were 97.8–99.7%. These speeds cannot be assumed for the 128GB Windows configuration below or higher-precision quantization.

RTX 5090 test conditions and complete results

On GB10, an October 4 test used an ASUS GX10 and the Arm64 fork, with IQ3_XXS, MTP enabled, reasoning off, synthetic inputs, and a 128-token output cap. Each condition ran once. Inputs of roughly 1.6K–75K tokens produced 37–42 tok/s; the short 431-token input produced 15.3 tok/s.

The same test recorded total wall times of 16.1 seconds for 431 input tokens, 8.6 seconds for 6,190, and 60.0 seconds for 75,047, each with 128 output tokens. Shorter input was not always faster, making prompts representative of everyday work particularly relevant.

GB10 fork measurements, including the short-input result

The two reports differ in quantization, software, prompts, output lengths, and cache conditions. They are neither a controlled 5090-versus-Spark comparison nor an evaluation of answer quality. They establish a fast 5090 configuration and a working GB10 execution path with measured rates. Read tok/s with the understanding that token boundaries vary by model and text.

Before buying, try the intended model with your own documents and questions, and judge answer quality alongside waiting time. If a smaller model handles the work, there is no need to spend the full ¥1 million. More VRAM or shared memory becomes valuable when the work calls for larger models or longer documents.

What Mac, Spark, and Windows PCs are good at

Mac Studio’s clear advantage is a large memory pool shared by CPU and GPU. A 64GB or 256GB configuration is not split into GPU-only memory and separate system RAM. The OS and applications share that capacity too, so the entire installed amount is not a model-weight budget.

The M5 Mac Studio was announced August 25 and released September 22, while the 512GB configuration is scheduled for late October. The 64GB Max and 256GB Ultra examples are released products, but that does not mean immediate delivery: the retailer lists incoming-stock estimates of five to six and nine to ten weeks respectively.

The quoted 40-core GPU M5 Max has a stated memory bandwidth of 614GB/s; M5 Ultra is rated at 1.2TB/s. Bandwidth matters when repeatedly reading large models. Dividing these figures by Spark’s 273GB/s does not yield a Japanese or English generation-speed ratio: computation, quantization, libraries, and context differ.

Apple’s M5 Mac Studio announcement

A Mac is a natural choice for someone who already works on macOS and wants to explore models supported by MLX or Metal. The value of 256GB is room for models and higher-precision weights that are difficult to place in 32GB VRAM. CUDA-only implementations and the upstream Strata installation call for a different platform.

MLX LM and llama.cpp provide starting points for checking supported models and execution methods. Sufficient capacity does not establish software support for a newly released model.

DGX Spark and related OEM systems are best understood as compact CUDA machines with GB10 and 128GB shared memory. They use an Arm CPU and Linux-based DGX OS. They make sense as a separate inference host accessed from a laptop or as a CUDA development environment. An advertised maximum parameter count does not establish acceptable latency or concurrency.

DGX Spark specifications

NVIDIA’s October 2 announcement was not the initial launch of the Spark family. The new 64GB configuration is OEM-only, scheduled for October 23 from $4,999. This research did not establish Japanese pricing, so a currency conversion is not presented as a domestic sub-¥1-million purchase option.

Availability and OEM scope of the new 64GB configuration

A Windows PC provides a practical route to GPU-resident models with Strata as an additional option. GPUs and SSDs are easier to change later, and the machine can also serve gaming or creative work. The trade-offs include a large GPU’s power and heat, drivers, and software support. Buying faster GPU computation and buying room for larger models are separate decisions.

What ¥1 million and ¥2 million actually buy

These tax-inclusive Japanese configuration examples were checked on October 7, 2026. They are neither a lowest-price list nor a comparison at guaranteed equal answer quality. The baseline SSD is 4TB. An 8TB drive is an option for storing many large models and documents, not a prerequisite for fast inference.

Example configuration Tax-inclusive price and availability Reason to choose Remaining constraint
Windows/RTX 5080 16GB
Core Ultra 7 270K Plus
RAM 64GB・4TB SN850X
¥829,280
Options plus shipping; delivery unconfirmed
Keep smaller models on the GPU; share the PC with creative work or games Little headroom for GPU-resident 27B models, depending on quantization and context
Mac Studio / M5 Max
18-core CPU, 40-core GPU
64GB unified memory, 4TB SSD
¥923,800
Incoming stock estimated 5–6 weeks after order
Combine daily Mac work with experiments in the 27B class and beyond No CUDA; not the upstream Strata execution path
HP ZGX Nano / GB10
128GB unified memory, 4TB SSD
¥1,097,800
Listed in stock during research
Use a compact, separate CUDA inference machine Exceeds ¥1 million; requires Arm support and an appropriate inference stack
Windows/RTX 5090 32GB
Core Ultra 7 270K Plus
RAM 128GB・4TB SSD
¥1,873,280
Options plus shipping; delivery unconfirmed
Switch between GPU-resident 27B models and Strata VRAM remains 32GB; four RAM modules run at DDR5-3600
Mac Studio / M5 Ultra
30-core CPU, 64-core GPU
256GB unified memory, 4TB SSD
¥1,939,800
Incoming stock estimated 9–10 weeks after order
Prioritize capacity for larger models or higher-precision weights Capacity alone does not establish chat speed or answer quality

For a new machine under ¥1 million, first decide whether you need a Mac or a Windows GPU workstation. A 64GB Max has a clear role for exploring several models including the 27B class on macOS. If you already own a GPU with at least 16GB, establish useful tasks before rushing to replace it with another card of the same capacity. A smaller model is not a reason to exhaust the budget.

The 128GB/4TB Spark examples sit just above ¥1 million. NVIDIA’s own DGX Spark was listed at ¥1,078,000, but the retailer marked it sold out with the next shipment expected in mid-October. A lower listed price does not mean immediate availability. Deciding whether you need a compact separate CUDA machine makes comparison with a Mac or tower clearer.

NVIDIA DGX Spark price and expected shipment

At up to ¥2 million, a Windows machine with 32GB VRAM and 128GB RAM can be compared with a Mac’s 256GB unified pool. For coding centered on a GPU-resident 27B model, with Strata used when appropriate, the 5090 path is attractive. If larger-model experiments and less aggressive quantization matter more, the Ultra path has a clearer purpose. Neither purchase guarantees frontier-cloud equivalence.

Pay particular attention to the Windows 128GB configuration. The listed four 32GB modules run at DDR5-3600, versus DDR5-5600 for two 32GB modules in the 64GB option. More capacity comes with lower CPU memory bandwidth. For Strata’s CPU expert work, capacity alone does not establish a superior configuration. If two 64GB modules are available, prioritize checking motherboard support and actual operating speed.

As another reference, ark lists an RTX 5090 system with 128GB DDR5-5600 and a 2TB SSD at ¥1,698,000. That is not a matching 4TB quote, and component availability matters. Treat it as a place to confirm a two-module 128GB layout, the 4TB total, and delivery timing rather than an automatic choice based on the lower listed price.

ark’s 128GB configuration example

The Dospara 5090 total adds ¥1,369,980 for the base machine, ¥308,000 for 128GB RAM, ¥192,000 for a 4TB SN850X, and ¥3,300 shipping. It is not a formal cart quote or reserved inventory. Cases and cooling also contribute to price differences, so the totals are not pure comparisons of LLM performance per yen.

The 5080 example adds ¥534,980 for the base PC, ¥99,000 for 64GB RAM, ¥192,000 for a 4TB SN850X, and ¥3,300 shipping. It includes the standard 1000W power supply and 240mm liquid cooler. This too is a total calculated from listed options, not a confirmed cart quote.

Deciding what your own AI should do

Define the local workload by its properties rather than a model name. Formatting drafts, extracting information from short documents, searching personal material, and making small testable code changes have inspectable inputs and outputs. Batch work with flexible deadlines is another natural local use.

Retain a strong cloud model for difficult design decisions, research in unfamiliar domains, and long autonomous workflows. Local preparation can reduce what needs to be sent out: organize documents or draft an approach locally, then escalate a limited question. If information cannot leave the machine, decide whether it can be separated from the question or whether more human judgment is needed.

The cost extends beyond hardware and electricity. Downloads, checks after updates, and troubleshooting consume your time. Dividing a machine’s price by a cloud subscription fee does not settle which is cheaper. Conversely, control over where data goes and the ability to retain a model version have value that a subscription comparison cannot fully capture.

If AI democratization means everyone owning the strongest possible AI, hardware and operational costs remain substantial. Yet the portion of useful work that can be retained independently of a provider is growing. Choosing the model, update schedule, supplied documents, and delegated authority is meaningful control.

A practical build starts with a small daily job and enough memory for that job. Expand after judging both the wait for a complete Japanese answer and the effort needed to correct it. The useful first step toward owning AI is finding a task you want it to keep doing, before spending the full budget on a large machine.

Research scope and sources

This research article draws on public documentation and Japanese retail listings checked on October 7, 2026, Japan time. LATENT did not benchmark the hardware, independently score Japanese quality, purchase a system, or download large models. The selection advice is editorial judgment based on specifications, published measurements, and software support. Prices, options, stock, and delivery estimates can change.

Model capabilities and formats come from their developers, file sizes from distributors, and speed figures from the original measurement reports. The reported measurements are used only within the conditions stated in the article. Retail prices can be checked through the links in the configuration table.