Product · AI Box
A language model in the office. No cloud, no account, no internet.
The AI Box is a finished machine with a strong open language model on it. It stands in your building, sits on your network, and needs no connection to the outside. Not a sentence your staff type, and not a document they upload, leaves the building.
- 3,000 € once
- 399 € a month for operations
- Delivered in 3 to 4 weeks
First we work out whether the machine pays off at your number of users.
The problem is not the vendor. It is the documents.
The tasks where a language model saves the most time nearly all hang on files nobody likes sending out of the building: quotes, contracts, personnel records, costings, customer correspondence. That is exactly where AI adoption stalls after the first trial.
A cloud service solves it contractually, and that works. But it stays a promise somebody else has to keep. The AI Box solves the same problem physically: no vendor who could read along, no transfer to a third country, no question about training. If somebody asks where a sentence was processed, you point at the machine in the next room.
No account
Nobody signs up anywhere. Access runs through the browser on the local network, and user management stays yours.
No subscription that can be cancelled
The model sits on the machine as a file. It keeps working even if the maker withdraws it or changes its pricing.
Demonstrable, not promised
You can pull the network cable and carry on working. That is the proof, and it takes five seconds to show.
What is in the box
Deliberately off-the-shelf hardware, no server parts. That is why the machine costs 3,000 € and not 15,000 €: a graphics card from the gaming shelf runs a quantized 27B model just as fast as a card with a datacenter sticker on it.
- 24 GB VRAMGraphics cardThis is where the model sits while it works. About 17 GB of model file fits into 24 GB of video memory with room to spare, which is exactly why an off-the-shelf card is enough.
- Air cooledProcessor and coolerThe processor takes the requests, the arithmetic happens on the graphics card. A normal tower cooler is enough: no water cooling, nothing that needs regular servicing.
- 64 GBMemoryHeadroom for long documents and several requests at once. Whatever does not fit into video memory is held here, without the machine starting to crawl.
- 2 TB NVMeStorageThe model file takes about 17 GB. The rest is room for further models and for the documents you give the machine to look things up in.
- 60 to 450 WPower supplyIdle, the machine draws about as much as a light bulb. It only reaches full load while an answer is being written, so in short bursts.
- Office quietFansLarge slow fans rather than small fast ones. The machine is allowed to stand next to a desk, and needs neither a server room nor air conditioning.
Specifications
- Model
- Qwen3.6-27B, quantized (Q4_K_M, about 17 GB)
- Graphics card
- 24 GB VRAM, consumer class
- Memory
- 64 GB
- Storage
- 2 TB NVMe
- Speed
- around 30 tokens per second, faster than most people read
- Concurrent requests
- 3 to 5, sized for 10 to 20 staff
- Connection
- Gigabit Ethernet, used through the browser
- Power draw
- about 60 W idle, up to 450 W under load
- Siting
- Office-friendly, no server room and no air conditioning needed
What your staff do with it
The interface is a chat window in the browser, reachable at an address on the company network. Anyone who has used ChatGPT needs no training.
Summarize documents
A forty-page contract down to ten lines of substance, without the file leaving the building.
Draft correspondence
Quotes, rejections, reminders, tender responses. The draft comes from the model, the responsibility stays with a person.
Translate
Technical text between German and English, without a translation service ever seeing the file.
Search your own files
Optionally with access to a directory you release to it. Ask questions of your own material instead of hunting through folders.
Code and formulas
Macros, scripts, spreadsheet formulas. In administration this is reliably the first use case that spreads by word of mouth.
What the monthly fee covers
A local model rarely fails on the technology. It fails because after six months nobody updates models, nobody checks whether the answers still hold up, and nobody picks up the phone when it jams. That is why we do not sell the machine without the operations.
Model updates
We roll in new model versions after testing them against your own cases. Never untested, never automatically overnight.
Answer quality
We keep a small set of your typical tasks and measure each update against it. If an update is not actually better, the old version stays.
Monitoring and on-call
We see it when the machine goes down, usually before you do. A number where somebody answers is part of it.
Hardware
If a part fails during the term, we replace it. The graphics card is the only part that ages in any meaningful way.
The access stays yours
Remote maintenance runs only through an access you enable and can switch off again at any time. If you would rather not, maintenance happens on site.
Variant · Hosted in Germany
When it grows: the same build in a datacenter.
Beyond roughly twenty-five active users a single consumer card stops being enough. At that point we put the same build into a German datacenter as a 2U server, on hardware of its own, with no cloud underneath and no shared machine.
The difference from the box is capacity, not principle. It stays your model on your machine, the machine simply no longer stands in your building but behind redundant uplinks and backup power. For companies without a suitable technical room, this is usually the quieter path.
Specifications
- Form factor
- 2U, 19 inch, hardware of its own
- Graphics card
- 48 GB VRAM
- Location
- Datacenter in Germany
- Users
- 25 to 50 in daily use
- Connection
- Leased line or VPN into your network
- Operations
- Redundant power supplies, backup power, monitoring around the clock
- Price
- 490 € a month, everything included
The three paths side by side.
| The three paths side by side. | AI Box on site | Server in Germany | Cloud subscription per user |
|---|---|---|---|
| Where your data sits | On the machine in your building | Datacenter in Germany | With the vendor, EU region bookable |
| Outside connection | None needed | Leased line or VPN | Needed constantly |
| Entry cost | 3,000 € once | No one-off cost | No one-off cost |
| Running cost | 399 € a month | 490 € a month | Per user per month |
| Fits | 10 to 20 staff | 25 to 50 staff | Any size |
| The model | Stays until you change it | Stays until you change it | Changes when the vendor changes it |
| If the internet goes down | Nothing happens | Access is gone | Everything stops |
Common questions about the AI Box.
- Is the model as good as ChatGPT?
- For summaries, correspondence, translation and straightforward programming tasks, Qwen3.6-27B is close enough that the difference rarely shows in daily work. On very long chains of reasoning and on specialist knowledge, the large cloud models are still ahead. We tell you up front which category your tasks fall into instead of letting you find out.
- Does the machine really have to be offline?
- No. It can also run on a normal company network with internet access. The point is that it does not have to. If you want the strictest version, it runs in its own network segment with no route to the outside, and then the third-country question is answered technically rather than contractually.
- Do we need a server room?
- No. The machine is an ordinary computer and can stand in an office or a technical cabinet. It needs a socket, a network port and air. Air conditioning is not required.
- What happens if the graphics card fails?
- We replace it as part of the maintenance. The model sits on the SSD as a file and is back in service half an hour after the swap. For companies that cannot absorb a day of downtime, we keep a spare machine ready.
- What does it cost over three years?
- 3,000 € once plus thirty-six monthly payments, about 17,400 € in total. Whether that beats licences depends on your number of users: from roughly fifty to a hundred intensive users, owned infrastructure comes in under comparable subscriptions, below that usually not. Most buyers do not choose the box because it is cheaper, though, but because the documents stay in the building.
- Can we expand the machine later?
- Yes. A second graphics card or a larger model is the usual next step, as long as the user count stays in range. If it grows well beyond that, the server in the datacenter is the sensible route, and the box stays on site as a fallback.