Skip to content

OpenAI-compatible API

lean serve exposes an HTTP API shaped like OpenAI’s, so most clients and SDKs work by changing the base URL.

Terminal window
lean serve lean-agent-light --port 8080
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="dummy")
resp = client.chat.completions.create(
model="lean-agent-light",
messages=[{"role": "user", "content": "Explain expert offloading in two sentences."}],
)
print(resp.choices[0].message.content)

The server holds one model and serves one request at a time. Concurrent requests queue rather than interleave.

Without --api-key (or LEAN_API_KEY) the API is open, which is why the default bind address is loopback-only. With a key set, every /v1 route requires it:

Authorization: Bearer <key>

/health is always reachable, key or not, so process supervisors do not need credentials.

Field Type Default Notes
messages array required {"role": "...", "content": "..."}. Roles: system, user, assistant.
model string server’s model Accepted and echoed back; the server serves whatever it was launched with.
max_tokens integer model limit Rejected with 400 if larger than the model’s context length.
temperature number 0.7 0 is greedy.
top_p number unset Nucleus sampling.
seed integer random Fixes sampling for reproducible output.
stream boolean false Server-sent events instead of one JSON body.
thinking boolean true Extension. Set false to skip the <think> phase. Ignored by models without one.
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1769472000,
"model": "lean-agent-light",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"reasoning_content": "The user wants a short explanation...",
"content": "Expert offloading keeps only the experts a token actually routes to in VRAM..."
},
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 24, "completion_tokens": 96, "total_tokens": 120 }
}

finish_reason is stop when the model emits an end-of-sequence token, or length when it hits max_tokens.

On thinking models, the <think> block is split out into reasoning_content rather than left inline in content. The field is omitted entirely when there is no reasoning to report, so clients that ignore it still see clean prose.

With "stream": true the response is text/event-stream. The first chunk carries the role, subsequent chunks carry content (or reasoning_content) deltas, the last chunk carries finish_reason and usage, and the stream ends with the literal data: [DONE].

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Expert"},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":24,"completion_tokens":96,"total_tokens":120}}
data: [DONE]

The legacy raw-text endpoint. It exists mainly so evaluation harnesses that score prompts (lm-eval-harness and friends) can drive the runtime.

Field Type Default Notes
prompt string or array required
max_tokens integer 16 OpenAI’s legacy default, not the chat default.
temperature number 0.7
top_p number unset
stop string or array unset Stop strings, trimmed from the output.
seed integer random
logprobs integer unset Return per-token logprobs. Only the top-1 alternative is reported.
echo boolean false Include the prompt in the response, with its logprobs when scoring.

Setting max_tokens: 0 with echo and logprobs selects pure scoring mode: the server runs a single chunked all-position prefill and returns per-token prompt logprobs without generating anything.

Unlike the chat endpoint, this one applies no chat template and no repeat penalty. What you send is what the model sees.

Lists the one model this server is hosting, in OpenAI’s list shape.

{
"object": "list",
"data": [
{ "id": "lean-agent-light", "object": "model", "created": 1769472000, "owned_by": "leanmodels" }
]
}

Returns ok with status 200. Unauthenticated.

Errors use OpenAI’s envelope:

{ "error": { "message": "max_tokens (999999) exceeds the model's maximum context length (262144)", "type": "invalid_request_error", "code": null } }
Status When
400 Malformed body, or max_tokens beyond the model’s context length.
401 Missing or wrong bearer key when one is configured.
408 Generation exceeded --request-timeout-secs.
500 Generation failed. The server recovers; the next request is served normally.

tools on the request and tool_calls on the response are supported for the Qwen-family models (lean-agent-light, lean-agent-middle, lean-coder-welter, lean-reason-heavy). A turn that calls a tool returns finish_reason: "tool_calls" with content: null, matching the OpenAI schema.

Tool syntax is not implemented by lean - it comes from the chat template the model authors shipped, which is stored in the model pack. That is why the format is always the one the model was trained on rather than an approximation.

Two limits worth knowing before you wire up a client:

  • Streaming tool calls are not supported yet. With "stream": true the response streams as ordinary text and tool calls are not broken out into deltas. Send "stream": false when you need structured tool_calls. Streaming support is planned.
  • tool_choice is accepted but not enforced. It is forwarded to the chat template, which may act on it. lean does not constrain decoding, so "required" cannot guarantee the model emits a call.

If the model emits a malformed tool call, the block is returned as ordinary message content, delimiters and all, rather than erroring. You see exactly what the model produced instead of the request failing.

These are reasoning models. They spend tokens inside a <think> block before producing anything you see, and max_tokens covers both. Set it too low and the budget is gone before the visible answer starts, so you get an empty content with finish_reason: "length":

// max_tokens: 32
{ "content": "", "finish_reason": "length" }
// max_tokens: 512
{ "content": "hello", "finish_reason": "stop" }

Nothing failed in the first case, the model simply never got to the answer. The same applies to tool calls, which are also emitted after thinking: too small a budget yields no tool_calls rather than a wrong one.

Give tool-calling requests room. A few hundred tokens is enough for a short call, and reasoning on a hard prompt can use far more. If you do not want thinking at all, send "thinking": false and the whole budget goes to the answer.

lean-agent-cruiser (DeepSeek-V4) uses a different tool format and is not covered yet; its runtime is still in progress.

  • Streaming tool calls. See above - use non-streaming requests for tool use.
  • Multiple completions. No n parameter; one choice per request.
  • Concurrency. Requests are serialised behind a single model instance.
  • Multiple models per server. One process serves one model. Run several processes on different ports if you need more.
  • Multimodal content. content is a plain string; no image or audio parts.
  • logit_bias, presence_penalty, frequency_penalty, and response_format are ignored if sent.