← All models

lean-agent-middle

Middleweight class · v1.1 · Qwen 3.5 base

Advanced agent - complex orchestration, long-context workflows

Verified vs llama.cppApache 2.0 basePaid

122B total, 10B active per token. A 75 GB model that runs on a single 24 GB GPU through expert offloading - massive knowledge with efficient per-token compute.

Specifications

Total params

122B

Active per token

10B

Base model

Qwen3.5-122B-A10B

Architecture

GDN hybrid MoE

Experts

256 experts, 8 active

Min VRAM / RAM

24 GB / 32 GB

Tool calling

Yes - non-streaming

tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send "stream": false when you need structured calls. See theAPI reference.

Source weights

The original model release this pack is built from, pinned to an exact commit. Our own build pipeline quantizes and tunes it for consumer GPUs, and every release is tested against the original model on standard benchmarks before it ships.

Quantization

bf16

Revision

dc4d348443bc740c68e2d77492492c11606384d5

Source size

250.2 GB (39 files)

This .lmpack

75.30 GB - core mostly Q6_K, experts mostly Q4_K

Packing converts the source weights into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.

Measured performance

Real numbers from our reference rig, not projections.

HardwarePrefillDecode
1x RTX 3090 (24 GB)6-12 tok/s5-6 tok/s
2x RTX 30904-8 tok/s6-8 tok/s

Decode figures are sustained rates measured over a 256-token generation. Short generations run slower while the expert cache warms - a 64-token run measures about 6 tok/s on two GPUs. Prefill spans 11- and 73-token prompts.

Get the model

Download 75.30 GB - single .lmpack, from bf16Price $19.99 one-time · launch pricing

Not available to buy or download yet - the model is packaged and verified, checkout is being finalized.

Once checkout opens, a purchase emails you a license key - run lean login once, then pull as usual. Free updates to purchased models.

Quick start

$ curl -sSf https://leanmodels.ai/install.sh | sh
$ lean run lean-agent-middle

Single binary, 15 MB. No Python, no Docker, no cloud dependency.

The pull step needs a license key, which is issued on purchase - checkout for this model is not open yet.