No per-message fees, ever. Pick a plan that fits your knowledge and traffic.
Loading plans…
Prefer not to upload anything to the cloud? Buy a Knowly appliance and run the whole assistant — chat, embeddings and reranking — inside your own building. Your data never leaves the box. Each server ships with Ubuntu and Knowly preinstalled, ready to plug in.
Small firms & single teams that want private AI without a data center.
5–10 concurrent users
Growing companies and multi-department deployments running larger models.
10–25 concurrent users
Enterprises and regulated industries needing maximum throughput and scale.
25–50+ concurrent users
Prices are starting configurations in USD and reflect the current (2026) GPU and memory market, where RTX GPUs, DDR5 and NVMe have all risen sharply. Final GPU, memory and storage are sized to your models and user count and quoted per order. Unlike a bare workstation, every Knowly appliance ships with the complete private RAG assistant — chat, embeddings and reranking — preinstalled and tuned; comparable custom on-premise RAG builds run $80,000–$300,000+ in software and services alone. Talk to us for a detailed quote.