Run AI in your studio. Not someone else's cloud.
On-premise AI workstations that run local language models on hardware you own — your data never leaves the building.
Renting AI means renting the risk. Every prompt, script, frame and client asset you send to a cloud AI service is processed on someone else’s hardware, under someone else’s terms. For studios under NDA, that’s real exposure — and the bill never stops. On-premise AI puts the model, the data and the compute inside your own building.
What on-premise AI gives you
Your data never leaves the building
No per-token bills
Runs the open models you already use
Built for sustained creative work
Plug it in and it's already an AI machine.
Every Renderboxes AI workstation ships with RBOS — our own tuned build of Ubuntu 24.04 LTS with a local AI runtime built in. A headless appliance: you never touch a Linux shell. It auto-discovers on your network and is managed from a browser. Ships GPU-ready (CUDA/ROCm pre-configured, no driver install) and exposes a local OpenAI-compatible API so your existing tools point at the on-prem box instead of the cloud. Built-in health dashboard: live per-GPU temp, VRAM, utilisation and power — fully offline.
Machine configurations
Machine configurations
| Machine | GPU option | Config | Total VRAM |
|---|---|---|---|
| Molecule Air | AMD Radeon AI PRO R9700 32GB | 8× | 256GB |
| Molecule Air | NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 96GB | 4× / 8× | 384 / 768GB |
| Nano Pro | NVIDIA RTX PRO 5000 Blackwell 72GB | 4× | 288GB |
| Nano Pro | NVIDIA RTX PRO 5000 Blackwell 48GB | 4× | 192GB |
More GPU memory = larger models run locally; memory pools across GPUs.
Benchmarks
[Benchmark placeholder — see Benchmarks system report below. AI benchmark data to be added once confirmed with Rich.]
Own it once. Stop paying per token.
Cloud AI is rented by the token and never stops. An on-premise workstation is a fixed cost you own outright, with only electricity ongoing — typically paying back the recurring cloud spend within months.
AI by workflow
Workstation vs server vs cloud
-
Where data lives
AI workstation: On the machine | On-prem server: In your building | Cloud AI: Third-party cloud
-
Who it serves
AI workstation: A studio/team | On-prem server: Whole organisation | Cloud AI: Anyone, metered
-
Ongoing cost
AI workstation: Electricity only | On-prem server: Electricity only | Cloud AI: Per token, forever
-
You own it
AI workstation: Yes | On-prem server: Yes | Cloud AI: No
FAQ
-
Does anything leave my network?
No. RBOS runs fully offline, air-gap capable — no telemetry, no cloud calls.
-
Which models can it run?
Open-weight GGUF models — Llama, Qwen, Mistral, Gemma, DeepSeek — locally, out of the box.
-
How fast is it?
On a Molecule Air with 8× AMD Radeon AI PRO R9700, a 30B model runs ~84 tok/s single-user and serves 2,400+ tok/s across a studio.
-
Do I need Linux skills?
No — RBOS is managed entirely from a browser.
-
Do I install GPU drivers?
No — ships GPU-ready with CUDA/ROCm pre-configured.
-
Can I point my own tools at it?
Yes — a local OpenAI-compatible API.
-
How much VRAM to run a local LLM?
Larger models need more; these pool 192GB–768GB across GPUs to run models that won’t fit on a single-GPU workstation.
Own your AI. Own the hardware.
Tell us your studio, your pipeline and the models you want to run. We’ll configure an on-premise AI workstation, build it in the UK, and ship it ready to run.