Skip to content

Running AI locally: what hardware do you need?

· 7 minute read

Running a language model locally sounds like a data center job. For most offices, though, a single AI server is enough. The deciding factor is GPU memory: it determines how large the model can be. Two more questions follow: how many people use the AI at the same time, and how many documents should it search?

The key figure: GPU memory

Language models run fastest on graphics processors (GPUs). How large the model can be is limited mainly by GPU memory (VRAM). The model has to fit in completely, otherwise it becomes very slow.

As a rule of thumb a model needs about 2 GB of memory per billion parameters at full precision (16 bit). Quantising to 4 bit cuts this to about a quarter, with little loss in quality. On top of that comes memory for the context, meaning the text the model is currently processing.

  • A model with 8 billion parameters needs about 5 to 6 GB at 4 bit and about 16 GB at 16 bit.
  • A model with 32 billion parameters needs about 20 GB at 4 bit.
  • A model with 70 billion parameters needs about 40 GB at 4 bit, which usually means two GPUs or one card with a lot of memory.

Concurrent users and context length

A model that answers smoothly for one person can slow down noticeably with ten concurrent requests. Each parallel request takes additional memory for its context. Processing long documents in one go also needs more memory.

For planning, what counts is therefore not the number of employees but how many of them typically ask at the same time and how long the texts they work with are.

RAM, CPU and storage

  • RAM: at least as much as GPU memory, ideally twice as much. It holds model files, search index and operating system.
  • CPU: less critical than the GPU. It mainly handles text recognition, indexing and managing requests.
  • Storage: fast NVMe SSDs with enough space for models, original documents and the search index, which grows with the document collection.
  • Cooling and noise: AI servers produce heat. In the office quiet operation matters, in the server room a 19-inch rack enclosure.

Hardware alone is not enough

A model on a GPU does not yet answer questions about your own files. That takes more building blocks:

  • Inference software that runs the model, for example vLLM or llama.cpp.
  • Document search (retrieval augmented generation, RAG) that finds the relevant passages for each question and passes them to the model.
  • Text recognition for scanned documents.
  • An interface with login and permissions, so everyone only sees what they are allowed to see.
  • Updates for model and software without data leaving the building.

Build your own or buy a ready-made AI server

An organisation with its own IT team experienced in Linux, GPU drivers and machine learning can build an AI server itself. The effort lies less in the hardware than in choosing, setting up and maintaining the software, and in making sure answers are reliably backed by sources.

A preconfigured appliance such as the Sinabox takes this work away: hardware, model, document search and interface are matched and set up. The Sinabox is available as a standard device and as a 19-inch rack for the server room. We size it up front based on user numbers and document volume.

Ready for AI without the cloud?

Talk to our team about deployment, configuration and an individual quote for your organization.

Frequently asked questions on this topic

Short answers to what is asked most often about this.

As a rule of thumb about 2 GB per billion parameters at 16 bit and about a quarter of that with 4-bit quantisation. A model with 8 billion parameters therefore runs from about 6 GB, one with 70 billion parameters needs around 40 GB.

Small models also run on the CPU but answer much more slowly. For several users in an office and for work with many documents, a GPU is practically a must.

No. After setup a local language model runs fully offline. The Sinabox needs no internet connection after initial setup and produces no outbound traffic.