← All models

lean-reason-heavy

Heavyweight class · v1.0 · Qwen 3.5 base

Near-frontier - deep reasoning, complex analysis, research

Apache 2.0Coming soonPaid

397B total, 17B active per token. Near-frontier reasoning from a massive expert pool, running entirely on your hardware.

Specifications

Total params

397B

Active per token

17B

Base model

Qwen3.5-397B-A17B

Architecture

GDN hybrid MoE

Experts

256 experts, 8 active

Target VRAM / RAM

48 GB / 64 GB

Tool calling

Yes - non-streaming

tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send"stream": false when you need structured calls. See theAPI reference.

Source weights

The exact public GGUF this model is packed from, pinned to a commit. Download it and you can reproduce our numbers with stock llama.cpp.

Quantization

UD-Q4_K_XL

Revision

da33c16fa4440f831149fcf53b98a22bc07785e5

Source size

245.27 GB (6 files)

This .lmpack

226.17 GB - core mostly Q6_K, experts mostly Q4_K

Packing splits the source GGUF into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.

Not yet verified

This model hasn't been benchmarked or verified on our hardware yet. Performance and hardware requirements are targets, not measured results. It won't be available for download or purchase until it clears the same verification the available models passed.

Get the model

Download 226.17 GB - single .lmpack, from UD-Q4_K_XL