lean-reason-heavy
Heavyweight class · v1.0 · Qwen 3.5 base
Near-frontier - deep reasoning, complex analysis, research
397B total, 17B active per token. Near-frontier reasoning from a massive expert pool, running entirely on your hardware.
Specifications
Total params
397B
Active per token
17B
Base model
Qwen3.5-397B-A17B
Architecture
GDN hybrid MoE
Experts
256 experts, 8 active
Target VRAM / RAM
48 GB / 64 GB
Tool calling
Yes - non-streaming
tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send"stream": false when you need structured calls. See theAPI reference.
Source weights
The exact public GGUF this model is packed from, pinned to a commit. Download it and you can reproduce our numbers with stock llama.cpp.
Repository
Quantization
UD-Q4_K_XL
Revision
da33c16fa4440f831149fcf53b98a22bc07785e5
Source size
245.27 GB (6 files)
This .lmpack
226.17 GB - core mostly Q6_K, experts mostly Q4_K
Packing splits the source GGUF into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.
Not yet verified
This model hasn't been benchmarked or verified on our hardware yet. Performance and hardware requirements are targets, not measured results. It won't be available for download or purchase until it clears the same verification the available models passed.