← All models

lean-agent-light

Lightweight class · v1.1 · Qwen 3.5 base

General-purpose agent - tool calling, structured output, multi-step reasoning

Verified vs llama.cppApache 2.0 baseFree

The entry point. A 22 GB model that runs on 12 GB VRAM - expert offloading handles the rest. Qwen3.5 GDN hybrid architecture surpasses last-gen models many times its size.

Specifications

Total params

35B

Active per token

3B

Base model

Qwen3.5-35B-A3B

Architecture

GDN hybrid MoE

Experts

256 experts, 8 active

Min VRAM / RAM

12 GB / 16 GB

Tool calling

Yes - non-streaming

tools and tool_calls work on non-streaming requests. Streaming does not break tool calls out into deltas yet, so send "stream": false when you need structured calls. See theAPI reference.

Source weights

The original model release this pack is built from, pinned to an exact commit. Our own build pipeline quantizes and tunes it for consumer GPUs, and every release is tested against the original model on standard benchmarks before it ships.

Quantization

bf16

Revision

59d61f3ce65a6d9863b86d2e96597125219dc754

Source size

71.9 GB (14 files)

This .lmpack

21.60 GB - core mostly Q6_K, experts mostly Q4_K

Packing converts the source weights into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.

Measured performance

Real numbers from our reference rig, not projections.

HardwarePrefillDecode
1x RTX 3090 (24 GB)16-41 tok/s30-40 tok/s
2x RTX 309011-27 tok/s24-37 tok/s

Download

Download 21.60 GB - single .lmpack, from bf16
$ lean pull lean-agent-light

Quick start

$ curl -sSf https://leanmodels.ai/install.sh | sh
$ lean pull lean-agent-light
$ lean run lean-agent-light

Single binary, 15 MB. No Python, no Docker, no cloud dependency.