LocallyAI

Local AI machines,
sized to your team.

Four Apple machines, from a Mac mini to a Mac Studio, each with the LocallyAI harness pre-installed — here is exactly what's inside, what it runs and what it costs installed.

16–128 GB unified · Serves 1–10+ · Fast and Smart lanes + Whisper, on the box · Unlimited seats

Model 01

The Compact

Solo practitioners · 1–2 people

For one or two people who mainly want consults and meetings turned into finished notes and letters, drafted in your templates.

Machine
Apple Mac mini
Memory
16 GB unified
Serves
1–2 people
Installed from
$2,500 ex GST
Technical line drawing of the machine supplied as The Compact

16 GB, allocated

  • 7 GB Fast lane models
  • 1.6 GB transcription
  • 1 GB document search
  • 3 GB reserved for the system
  • 3.4 GB kept spare for conversation context
Model 02

The Office

Small teams · 2–5 people

The most popular starting point. Both model lanes fit, so the router can send the hard questions to the Smart lane.

Machine
Apple Mac mini or AMD AI 395+
Memory
32–48 GB unified
Serves
2–5 people at once
Installed from
$5,000 ex GST
Technical line drawing of the machine supplied as The Office

32 GB, allocated

  • 16 GB Smart + Fast lane models
  • 3 GB transcription
  • 1 GB document search
  • 3 GB reserved for the system
  • 9 GB kept spare for conversation context
Model 03

The Practice

Whole offices · 5–10 people

For an office where several things happen at once — a consult transcribing in one room while a contract summary drafts in another. Reads full agreements and complete histories in a single pass.

Machine
Apple Mac Studio or AMD AI 395+
Memory
64 GB unified
Serves
5–10 people, several jobs at once
Installed from
$8,000 ex GST
Technical line drawing of the machine supplied as The Practice

64 GB, allocated

  • 25 GB Smart lane models
  • 6 GB Fast lane models
  • 3 GB transcription
  • 1 GB document search
  • 3 GB reserved for the system
  • 26 GB kept spare for conversation context
Model 04

The Firm

Document-heavy practices · 10+ people

Our largest standard machine, for firms buried in discovery or clinics with many practitioners. Both lanes run at full quantisation with room for background jobs.

Machine
Apple Mac Studio or AMD AI 395+ (128 GB), top spec
Memory
96–128 GB unified
Serves
10+ people, heavy simultaneous use
Installed from
$10,000 ex GST
Technical line drawing of the machine supplied as The Firm

128 GB, allocated

  • 63 GB Smart lane models
  • 16 GB Fast lane models
  • 3 GB transcription
  • 1 GB document search
  • 3 GB reserved for the system
  • 42 GB kept spare for conversation context
Option

The Portable

On the move, or serving the room — depends on what the day needs

The same harness on a laptop. On your desk it serves the office like any of the machines above; picked up, it's a private AI that travels — between rooms, out to site, down to court — with no network and no box to connect to.

Machine
Apple MacBook Pro
Memory
36–64 GB unified
Serves
The whole office plugged in · one person on the move
Installed from
Quoted to the spec you pick
Technical line drawing of the laptop supplied as The Portable

The Harness, on the laptop itself

at the desk — staff connect by browserserving
in the bag — chat, files, dictationon board
out of the officestill private
Sizing

We would rather sell you the right machine than the big one.

Every tier runs the same Fast and Smart lanes — bigger machines simply run them at higher precision and fit more models at once. We typically run models at Q6–Q8, occasionally Q4 where a system needs it.

Default model options (other models on request)

The one roster, for every tier on this page. Every machine also keeps memory in reserve so it stays fast on its busiest day — and if you prefer AMD workstation hardware, it's available on request.

Fast lane — Qwen 3.5 9B · Ornith 1.5 9B · Gemma 4 12Bloaded
Smart lane — Qwen 3.8 27B · Ornith 1.5 35B MoE · GPT-OSS 20B · Gemma 4 26B-A4Bloaded
Smart lane, The Firm — adds DeepSeek V4 Flashloaded
Transcription — Whisperloaded
The Compact — Fast lane and Whisper only, at its sizethe one exception
Headroom — kept spare for conversation contextreserved

A note on the models

The models named on this page are what we install today — better ones keep arriving. Our latest recommendations are easily accessed and downloaded straight through the browser, provided the machine has an internet connection.

Every install ends with a health check across every service on the box, and the same check runs again at every visit.

Book a 20-minute chat.

We'll tell you honestly whether this suits your business — and what it would cost.

20 min · free Book a chat