lean-agent-middle
Middleweight class · v1.1 · Qwen 3.5 base
Advanced agent - complex orchestration, long-context workflows
122B total, 10B active per token. A 75 GB model that runs on a single 24 GB GPU through expert offloading - massive knowledge with efficient per-token compute.
Specifications
Total params
122B
Active per token
10B
Base model
Qwen3.5-122B-A10B
Architecture
GDN hybrid MoE
Experts
256 experts, 8 active
Min VRAM / RAM
24 GB / 32 GB
Tool calling
Yes - non-streaming
tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send "stream": false when you need structured calls. See theAPI reference.
Source weights
The original model release this pack is built from, pinned to an exact commit. Our own build pipeline quantizes and tunes it for consumer GPUs, and every release is tested against the original model on standard benchmarks before it ships.
Repository
Quantization
bf16
Revision
dc4d348443bc740c68e2d77492492c11606384d5
Source size
250.2 GB (39 files)
This .lmpack
75.30 GB - core mostly Q6_K, experts mostly Q4_K
Packing converts the source weights into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.
Measured performance
Real numbers from our reference rig, not projections.
| Hardware | Prefill | Decode |
|---|---|---|
| 1x RTX 3090 (24 GB) | 6-12 tok/s | 5-6 tok/s |
| 2x RTX 3090 | 4-8 tok/s | 6-8 tok/s |
Decode figures are sustained rates measured over a 256-token generation. Short generations run slower while the expert cache warms - a 64-token run measures about 6 tok/s on two GPUs. Prefill spans 11- and 73-token prompts.
Get the model
Not available to buy or download yet - the model is packaged and verified, checkout is being finalized.
Once checkout opens, a purchase emails you a license key - run lean login once, then pull as usual. Free updates to purchased models.
Quick start
$ curl -sSf https://leanmodels.ai/install.sh | sh
$ lean run lean-agent-middleSingle binary, 15 MB. No Python, no Docker, no cloud dependency.
The pull step needs a license key, which is issued on purchase - checkout for this model is not open yet.