Ornith 1.5 9B — VRAM requirements

Ornith AI · 9.65B params · up to 256K context

QuantizationVRAM at 4K contextVRAM at 256K context
FP16 / BF1622.8 GB61.7 GB
FP811.7 GB50.6 GB
INT811.7 GB50.6 GB
INT46.17 GB45.1 GB
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.

Compatible GPUs at 256K context

FP16 / BF16 (61.7 GB)

FP8 (50.6 GB)

INT8 (50.6 GB)

INT4 (45.1 GB)

paramsB (9.65B) is the exact safetensors total from the Hugging Face API for ornith-ai/Ornith-1.5-9B (ungated); architecture fields from that repo's config.json. Dense model. maxContextLength 262,144 per the Ornith 1.5 card, which also documents YaRN rope-scaling at factor 4 to reach roughly 1M tokens — not modelled here. Provenance note: Ornith's own card states the 1.0 family was built on top of Qwen3.5 and Gemma 4 with continued pretraining, so this is a derivative rather than a from-scratch architecture. Pulled 2026-09-19.

← Interactive VRAM calculator · All GPUs