Documentation
Everything here was run on real hardware. Where there is a number, it was measured, and the machine it came from is named.
Quick start
One command. It brings its own Python, the llama.cpp binary and the GPU setup — nothing to install first. On Windows, run it inside WSL (Ubuntu), not PowerShell.
curl -fsSL https://github.com/Ga0512/SursumAI/raw/v1.0.18/install.sh | bash
sursumai
The dashboard opens at http://localhost:3000. Create an account
there — it is local, stored on your machine — then New deployment, pick
a model, and Deploy. The first one takes a few minutes while the weights
come down.
Calling the API
It speaks the OpenAI API. Point your existing code at your own machine and change nothing else.
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8001/v1",
api_key="sk-sursum-...", # API keys tab in the dashboard
)
r = client.chat.completions.create(
model="Qwen/Qwen3-8B-GGUF",
messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)
Or with curl:
curl http://localhost:8001/v1/chat/completions \
-H "Authorization: Bearer sk-sursum-..." \
-H "Content-Type: application/json" \
-d '{"model": "Qwen/Qwen3-8B-GGUF",
"messages": [{"role": "user", "content": "Hello"}]}'
One key works for every model and pool in your account, the same way an
OpenAI key does. GET /v1/models lists what you have running.
Choosing a model
Paste any Hugging Face repository, or the link to one GGUF file. The dashboard reads the repository and lists the quantizations with their sizes; from the API, name the file with a colon:
"model": "unsloth/Qwen3.8-27B-GGUF:Q4_K_S"
Rule of thumb for speed, and it is just division:
tokens/s ≈ GPU memory bandwidth ÷ model file size × 0.7
Measured, same 27B model at Q4 (16 GB), same prompt:
| GPU | Bandwidth | Tokens/s |
|---|---|---|
| NVIDIA A40 | 696 GB/s | 30 |
| RTX PRO 4500 SE | 800 GB/s | 37 |
| RTX 6000 Ada | 960 GB/s | 45 |
Memory size decides what fits; bandwidth decides how fast. A 48 GB card is not faster than a 24 GB one — it just holds more.
JSON output
Give it a schema and the runtime generates only valid JSON. No prose to strip, no half-parsed answers.
r = client.chat.completions.create(
model="unsloth/Qwen3.8-27B-GGUF",
messages=[{"role": "user", "content": "Extract the fields:\n\n" + document}],
max_tokens=1500,
temperature=0,
response_format={"type": "json_schema", "json_schema": {
"name": "fields",
"schema": {
"type": "object",
"properties": {
"date": {"type": "string"},
"place": {"type": "string"},
"summary": {"type": "string", "maxLength": 200},
},
"required": ["date", "place", "summary"],
},
}},
)
maxLength is worth setting. The field the model is free to ramble
in is usually most of the time spent.
Reasoning models
A reasoning model writes its thinking before the answer, and you wait for all of it. On a 27B extracting six fields: 1175 tokens and 26.1 s thinking, 176 tokens and 4.5 s with thinking off. Same answer.
Set it once on the deployment — Advanced options → Tuning → Thinking — and every call to that model gets it, with nothing to remember in your code. A single call can ask for the opposite:
# this one needs to think
client.chat.completions.create(
model="unsloth/Qwen3.8-27B-GGUF",
messages=[...],
max_tokens=3000,
extra_body={"chat_template_kwargs": {"enable_thinking": True},
"reasoning_effort": "low"}, # low | medium | xhigh
)
Measured on the same question: thinking off, 1081 tokens; low,
795 tokens and faster than off — a little thinking makes the model
organise the answer instead of rambling.
Many documents at once
Generating for eight requests costs almost what one costs: the GPU reads the model once and serves the whole batch. Measured on an RTX 6000 Ada, 25 000 characters per document (~7 000 tokens):
| At the same time | Total | Per document |
|---|---|---|
| 1 | 4.7 s | 4.7 s |
| 4 | 7.8 s | 1.9 s |
| 8 | 14.0 s | 1.7 s |
The one rule: the context is shared between the conversations running at the same time.
max_model_len ≥ parallel requests × tokens of your largest document
Four 7 000-token documents on a 16 384 context give you "Context size has been exceeded". Set Max model len to 40 960 and Conversations at once to 4 and they fit with room to spare.
Routing between models
A pool is two or more deployments behind one name. Call
model="router" and the cheap model answers what it can; the big one
is woken when it has to be.
client.chat.completions.create(
model="router",
messages=[...],
extra_body={"session_id": "user-42"}, # keeps the routing state per conversation
)
The answer tells you who served it:
"model": "mix → Qwen/Qwen3-4B-GGUF (escalated)".
Five modes: escalation (a judge model escalates), classifier (a judge picks one of N), advisor (the judge runs in the background and counts for the next turn), stage (keyword rules, no judge) and round robin. We measured escalation on 100 GSM8K problems: a 0.6B alone got 69 %, the 8B alone 93 %, and the pool 87 % at a third of the time — and a judge smaller than 1.7B was worse than a coin flip, so give the pool a decent judge.
Your own servers (Pro)
Add a machine by pasting the line your provider shows you — the whole thing, port included:
ssh [email protected] -p 22133 -i ~/.ssh/id_ed25519
SursumAI installs itself there, opens a private SSH tunnel and the models on
that machine join your pools and your API. The URL you call does not
change: it stays http://localhost:8001/v1, because your
machine is the front door and the remote model is never exposed — it listens on
127.0.0.1 on the server, and the tunnel is the only way in.
Before adding one, ssh user@host has to work from the machine
running SursumAI, without a password. That means the key of that
machine, not your laptop's, if they are different.
Running it as a service
On a server, the ports and the address are yours to choose:
SURSUMAI_BIND=0.0.0.0 \
SURSUMAI_WEB_PORT=80 \
SURSUMAI_CENTRAL_PORT=8080 \
sursumai
The base URL in the dashboard follows what you chose. A port outside 1-65535, two services on the same port, or anything inside 9000-9099 (where deployed models listen) stops the process with a sentence instead of a collision.
A common shape: a small always-on box (an EC2 t3.small is enough,
it needs no GPU) running SursumAI, and a rented GPU attached over SSH that you
turn on and off. The box holds the database, the keys and the URL; the GPU is
the muscle.
Command line
sursumai # start everything and open the dashboard
sursumai status # what is running
sursumai list # your deployments
sursumai deploy Qwen/Qwen3-4B-GGUF
sursumai chat <id> "hello"
sursumai logs <id> --follow
sursumai update # upgrade to the latest release
sursumai uninstall # remove it from this computer
When something breaks
- "Context size has been exceeded"
- Your parallel requests together need more context than the deployment has. Raise Max model len, or lower Conversations at once.
- The card says the model stopped running
- Press Redeploy. If it keeps happening, open Logs — the last lines are the runtime's own words.
- "This model was not published quantized that way"
- You asked vLLM for AWQ or GPTQ on a repository that ships full weights. Leave the quantization alone, or pick a repository that was published quantized.
- The answer comes back empty
- A reasoning model spent the whole budget thinking. Raise
max_tokensto 1500 or more, or turn thinking off on the deployment. - The model is slower than I expected
- Check the arithmetic in Choosing a model. If it matches, the card is at its ceiling — fewer output tokens or a card with more bandwidth are the only levers.
- A machine went offline
- Rented GPUs change address when they restart. Remove the machine and add it again with the new SSH line.
Install it
$ curl -fsSL https://github.com/Ga0512/SursumAI/raw/v1.0.18/install.sh | bash
Something missing here? Open an issue — the docs are part of the repository.