Run locally / Model details

NeoHorse Jev 4B

4B text and image model.

Recipe ready · unverifiedtextimage#5 on ImageJevBench (Capability, v0.3.0)

Hardware

Will it run on my machine?

Requirements are recorded per variant. A dash means the catalogue does not provide that value.

VariantPrecisionRuntimeBenchmarkedInstall recipeMinimum / recommended VRAMRAMDiskPlatformsExpected speed
bf16-author-bundlebf16PyTorchNoRecipe ready · unverified12 GB / 12 GB16 GB9.1 GBlinux-nvidiaNot measured for this installer recipe. Historical published benchmark timings describe a separate runtime.
bf16-torch-cpubf16PyTorchNoRecipe ready · unverifiedUses system/unified RAM / Uses system/unified RAM16 GB9.1 GBlinux-cpu, windows-cpuNot measured for this installer recipe. Historical published benchmark timings describe a separate runtime.
bf16-torch-mpsbf16PyTorchNoRecipe ready · unverifiedUses system/unified RAM / Uses system/unified RAM24 GB9.1 GBmacos-arm64Not measured for this installer recipe. Historical published benchmark timings describe a separate runtime.
q8_0-llamacpp-author-runtimegguf-q8_0llama.cppNoDocumentation only12 GB / 12 GB16 GB5.17 GBlinux-nvidiaUnbenchmarked estimate: short text decisions may take fractions of a second to a few seconds on NVIDIA GPUs, depending on context and question count. Author tested SM 90/Hopper; RTX 4090/Metal/CPU parity not established. NVIDIA memory class example: RTX 4090 24 GB; author only tested SM 90/Hopper, so Ada build/operator parity remains unverified.
q4_k_m-llamacpp-author-runtimegguf-q4_k_mllama.cppNoDocumentation only12 GB / 12 GB16 GB3.39 GBlinux-nvidiaUnbenchmarked estimate: short text decisions may take fractions of a second to a few seconds on NVIDIA GPUs, depending on context and question count. Author tested SM 90/Hopper; RTX 4090/Metal/CPU parity not established. NVIDIA memory class example: RTX 4090 24 GB; author only tested SM 90/Hopper, so Ada build/operator parity remains unverified.
bf16-author-bundlebf16
Runtime
PyTorch
Install recipe
Recipe ready · unverified
VRAM
12 GB minimum / 12 GB recommended
RAM
16 GB
Disk
9.1 GB
Platforms
linux-nvidia
Expected speed
Not measured for this installer recipe. Historical published benchmark timings describe a separate runtime.
bf16-torch-cpubf16
Runtime
PyTorch
Install recipe
Recipe ready · unverified
VRAM
Uses system/unified RAM minimum / Uses system/unified RAM recommended
RAM
16 GB
Disk
9.1 GB
Platforms
linux-cpu, windows-cpu
Expected speed
Not measured for this installer recipe. Historical published benchmark timings describe a separate runtime.
bf16-torch-mpsbf16
Runtime
PyTorch
Install recipe
Recipe ready · unverified
VRAM
Uses system/unified RAM minimum / Uses system/unified RAM recommended
RAM
24 GB
Disk
9.1 GB
Platforms
macos-arm64
Expected speed
Not measured for this installer recipe. Historical published benchmark timings describe a separate runtime.
q8_0-llamacpp-author-runtimegguf-q8_0
Runtime
llama.cpp
Install recipe
Documentation only
VRAM
12 GB minimum / 12 GB recommended
RAM
16 GB
Disk
5.17 GB
Platforms
linux-nvidia
Expected speed
Unbenchmarked estimate: short text decisions may take fractions of a second to a few seconds on NVIDIA GPUs, depending on context and question count. Author tested SM 90/Hopper; RTX 4090/Metal/CPU parity not established. NVIDIA memory class example: RTX 4090 24 GB; author only tested SM 90/Hopper, so Ada build/operator parity remains unverified.
q4_k_m-llamacpp-author-runtimegguf-q4_k_m
Runtime
llama.cpp
Install recipe
Documentation only
VRAM
12 GB minimum / 12 GB recommended
RAM
16 GB
Disk
3.39 GB
Platforms
linux-nvidia
Expected speed
Unbenchmarked estimate: short text decisions may take fractions of a second to a few seconds on NVIDIA GPUs, depending on context and question count. Author tested SM 90/Hopper; RTX 4090/Metal/CPU parity not established. NVIDIA memory class example: RTX 4090 24 GB; author only tested SM 90/Hopper, so Ada build/operator parity remains unverified.
dm-local plan neohorse-jev-4bChecks memory, disk, platform, and runtime against the catalogue.

Quick start

Choose where to run it

Installer commands use the pinned model revision. Confirm the model terms before proceeding.

An install recipe is available, but it has not been verified by an install test.

  1. 1curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH"
  2. 2dm-local install neohorse-jev-4b

Call your endpoint

Use a local typed endpoint

The API accepts typed questions and returns typed answers with probabilities.

Request

curl http://127.0.0.1:8484/v1/systemone -H 'Content-Type: application/json' -d '{"state":"Mia owns a red bicycle.","questions":{"color":{"type":"choice","instructions":"Which color is the bicycle?","criteria":{"red":null,"blue":null}}}}'

Example output

{"id":"dec_…","model":"neohorse-jev-4b","answers":{"color":{"type":"choice","choice":"red","confidence":0.97,"probabilities":{"red":0.97,"blue":0.03}}},"usage":{"input_tokens":64,"output_tokens":0,"decisions":1}}

Python

import json
import urllib.request
payload = {"state": "Mia owns a red bicycle.",
           "questions": {"color": {"type": "choice", "instructions": "Which color is the bicycle?", "criteria": {"red": None, "blue": None}}}}
request = urllib.request.Request("http://127.0.0.1:8484/v1/systemone", data=json.dumps(payload).encode(), headers={"Content-Type": "application/json"})
with urllib.request.urlopen(request, timeout=30) as response:
    result = json.load(response)
print(result["answers"]["color"])

Image input

curl http://127.0.0.1:8484/v1/multimodal -H 'Content-Type: application/json' -d '{"state":"Read the image and answer the question.","images":["data:image/png;base64,…"],"questions":{"object":{"type":"choice","instructions":"What is shown?","criteria":{"bicycle":null,"car":null}}}}'
Read the hosted API reference →

Uninstall

Remove the local files

Use the installer to stop the service and remove this model's downloaded files.

dm-local uninstall neohorse-jev-4b

Model terms

Licence details

apache-2.0Commercial use: yes

Commercial use is allowed by the model’s listed licence.

Model card

The author has not published where the training data came from. We found no statement or evidence that it was distilled from Jev outputs.

The local installer has a separate free and commercial licence.

Troubleshooting

Common setup issues

Driver or CUDA version is too old

Update to a driver/runtime combination listed for the selected variant, then run dm-local plan again.

Out of memory

Choose a smaller quantised variant when one is listed, or use a remote GPU with enough recommended memory.

NVIDIA Container Toolkit is missing

Install and configure the NVIDIA Container Toolkit for your host before retrying the container runtime.

The port is already in use

Stop the service using port 8484 or configure a different local port, then rerun the self-test.

A download was interrupted

Run the install command again. Completed downloads resume when the artifact server supports range requests.

Checksum verification failed

Do not use the file. Remove the incomplete download and retry from the pinned official source.

Apple Silicon memory pressure

Close other memory-heavy apps and choose a smaller supported variant if the catalogue lists one.

WSL2 cannot see the GPU

Check the Windows GPU driver and WSL2 GPU support, then verify nvidia-smi inside the WSL distribution.

SSH or firewall blocks the endpoint

Keep the service on loopback and use an SSH tunnel. Avoid exposing the API port directly to the public internet.

NeoHorse Jev 4B local setup — Decision Models