- Runtime
- vLLM 0.30.0
- Install recipe
- Recipe ready · unverified
- VRAM
- 48 GB minimum / 48 GB recommended
- RAM
- 80 GB
- Disk
- 62.58 GB
- Platforms
- linux-nvidia, wsl2-nvidia
- Expected speed
- About 47 ms p 50 per decision (raw, serial) on the measured GPU; the author reports 56 ms median on its own run. Provisioning class: RTX PRO 6000 96 GB; this names a memory class, not a latency measurement or confirmation of architecture compatibility. CPU/Mac speed is estimated only in separate portable variants.
Run locally / Model details
decisio v0.8.0 on gemma-4-31B-it
31B text model.
Hardware
Will it run on my machine?
Requirements are recorded per variant. A dash means the catalogue does not provide that value.
| Variant | Precision | Runtime | Benchmarked | Install recipe | Minimum / recommended VRAM | RAM | Disk | Platforms | Expected speed |
|---|---|---|---|---|---|---|---|---|---|
| bf16-fp8-on-load-vllm | fp8-on-load (base bf16) | vLLM 0.30.0 | Yes | Recipe ready · unverified | 48 GB / 48 GB | 80 GB | 62.58 GB | linux-nvidia, wsl2-nvidia | About 47 ms p 50 per decision (raw, serial) on the measured GPU; the author reports 56 ms median on its own run. Provisioning class: RTX PRO 6000 96 GB; this names a memory class, not a latency measurement or confirmation of architecture compatibility. CPU/Mac speed is estimated only in separate portable variants. |
dm-local plan decisio-gemma-4-31b-v080Checks memory, disk, platform, and runtime against the catalogue.Quick start
Choose where to run it
Installer commands use the pinned model revision. Confirm the model terms before proceeding.
An install recipe is available, but it has not been verified by an install test.
- 1
curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH" - 2
dm-local install decisio-gemma-4-31b-v080
Call your endpoint
Use a local typed endpoint
The API accepts typed questions and returns typed answers with probabilities.
Request
curl http://127.0.0.1:8484/v1/systemone -H 'Content-Type: application/json' -d '{"state":"Mia owns a red bicycle.","questions":{"color":{"type":"choice","instructions":"Which color is the bicycle?","criteria":{"red":null,"blue":null}}}}'Example output
{"id":"dec_…","model":"decisio-gemma-4-31b-v080","answers":{"color":{"type":"choice","choice":"red","confidence":0.97,"probabilities":{"red":0.97,"blue":0.03}}},"usage":{"input_tokens":64,"output_tokens":0,"decisions":1}}Python
import json
import urllib.request
payload = {"state": "Mia owns a red bicycle.",
"questions": {"color": {"type": "choice", "instructions": "Which color is the bicycle?", "criteria": {"red": None, "blue": None}}}}
request = urllib.request.Request("http://127.0.0.1:8484/v1/systemone", data=json.dumps(payload).encode(), headers={"Content-Type": "application/json"})
with urllib.request.urlopen(request, timeout=30) as response:
result = json.load(response)
print(result["answers"]["color"])Read the hosted API reference →Uninstall
Remove the local files
Use the installer to stop the service and remove this model's downloaded files.
dm-local uninstall decisio-gemma-4-31b-v080Model terms
Licence details
Commercial use is allowed by the model’s listed licence.
The local installer has a separate free and commercial licence.
Troubleshooting
Common setup issues
Driver or CUDA version is too old
Update to a driver/runtime combination listed for the selected variant, then run dm-local plan again.
Out of memory
Choose a smaller quantised variant when one is listed, or use a remote GPU with enough recommended memory.
NVIDIA Container Toolkit is missing
Install and configure the NVIDIA Container Toolkit for your host before retrying the container runtime.
The port is already in use
Stop the service using port 8484 or configure a different local port, then rerun the self-test.
A download was interrupted
Run the install command again. Completed downloads resume when the artifact server supports range requests.
Checksum verification failed
Do not use the file. Remove the incomplete download and retry from the pinned official source.
Apple Silicon memory pressure
Close other memory-heavy apps and choose a smaller supported variant if the catalogue lists one.
WSL2 cannot see the GPU
Check the Windows GPU driver and WSL2 GPU support, then verify nvidia-smi inside the WSL distribution.
SSH or firewall blocks the endpoint
Keep the service on loopback and use an SSH tunnel. Avoid exposing the API port directly to the public internet.