Ornith 1.5 9B — VRAM requirements
Ornith AI · 9.65B params · up to 256K context
| Quantization | VRAM at 4K context | VRAM at 256K context |
|---|---|---|
| FP16 / BF16 | 22.8 GB | 61.7 GB |
| FP8 | 11.7 GB | 50.6 GB |
| INT8 | 11.7 GB | 50.6 GB |
| INT4 | 6.17 GB | 45.1 GB |
Weights = params × bytes-per-parameter for the selected quantization.KV cache is always computed at FP16 and assumes batch size 1. A flat +15% is added on top for CUDA context/activation overhead — a rule of thumb, not a measured figure. Full methodology on the interactive calculator.
Compatible GPUs at 256K context
FP16 / BF16 (61.7 GB)
FP8 (50.6 GB)
INT8 (50.6 GB)
INT4 (45.1 GB)
paramsB (9.65B) is the exact safetensors total from the Hugging Face API for ornith-ai/Ornith-1.5-9B (ungated); architecture fields from that repo's config.json. Dense model. maxContextLength 262,144 per the Ornith 1.5 card, which also documents YaRN rope-scaling at factor 4 to reach roughly 1M tokens — not modelled here. Provenance note: Ornith's own card states the 1.0 family was built on top of Qwen3.5 and Gemma 4 with continued pretraining, so this is a derivative rather than a from-scratch architecture. Pulled 2026-09-19.