lean-agent-cruiser
Cruiserweight class · v1.0 · DeepSeek V4-Flash base
Million-token context - frontier long-context work on a single GPU
284B total, 13B active per token. DeepSeek's hybrid attention pairs Compressed Sparse Attention with Heavily Compressed Attention to hold a 1,048,576-token context, and Manifold-Constrained Hyper-Connections stabilise signal propagation across its 43 layers.
Specifications
Total params
284B
Active per token
13B
Base model
DeepSeek-V4-Flash
Architecture
Hybrid CSA/HCA attention MoE
Experts
256 experts, 6 active + 1 shared
Target VRAM / RAM
24 GB / 64 GB
Tool calling
Not yet
Source weights
The original model release this pack is built from, pinned to an exact commit. Our own build pipeline quantizes and tunes it for consumer GPUs, and every release is tested against the original model on standard benchmarks before it ships.
Repository
Quantization
fp8
Revision
60d8d70770c6776ff598c94bb586a859a38244f1
Source size
159.6 GB (46 files)
This .lmpack
155.14 GB - core mostly Q8_0, experts mostly MXFP4
Packing converts the source weights into an offloadable layout, so the .lmpack size differs from the source file. K-quantized weights are mixed by design - the dtypes named are the ones holding the most bytes, not the only ones present.
Not yet verified
This model hasn't been benchmarked or verified on our hardware yet. Performance and hardware requirements are targets, not measured results. It won't be available for download or purchase until it clears the same verification the available models passed.