STOP RENTING INTELLIGENCE. OWN IT.

Private AI for your code, documents, and daily work—running on hardware you own.

Local AI infrastructure / Built to order

Ownership / 01

Your models. Your data. Your hardware.

We select and install leading open-weight models for the hardware you buy. You choose what runs. There is no per-token charge for the models on your box.

Alfred Remembers keeps conversations, documents, recordings, photos, and notes on the workstation. Local inference and document search can run without sending prompts or files to an external model provider.

The computer is yours, set up for the way you work. Offline configurations are available. Network access is only for the connections you enable, such as web research, email, or software updates.

Alfred Remembers / 02

Give your private AI a long-term memory.

Alfred Remembers connects the conversations, documents, recordings, photos, and notes you add into a lasting memory stored on your Intelligence Box.

Ask “What did we decide about that contractor?” or “What was the restaurant Sarah recommended?” Alfred helps find the relevant history, so you can pick up where you left off.

For families, it brings together household documents, purchases, warranties, preferences, and recommendations. For businesses, it connects meetings, projects, customers, procedures, and decisions.

Your memory lives on hardware you control and can be used by supported local or private AI models. It is included with every box.

  • Conversations 01
  • Documents 02
  • Recordings 03
  • Photos 04
  • Notes 05

The Boxes / 03

Four standard boxes.

Each one is hardware you own, models selected for that hardware, and our software—installed and ready for a kind of work.

Solo

Your personal AI workspace.

A compact Mac mini with 64 GB of unified memory, configured for one person using a coding agent, private chat, and document search. A practical starting point for local AI, with compact models selected for useful quality and responsiveness.

Apple M5 Pro · 18-core CPU / 20-core GPU · 64 GB unified memory · 1 TB SSD

64 GB
Unified memory
1
Person
$4,495
From

Pro

More room for demanding work.

A Mac Studio with 128 GB of unified memory for larger coding models, longer working context, and heavier document tasks. Built for a developer or power user who wants more model choice and room to grow.

Apple M5 Max · 18-core CPU / 40-core GPU · 128 GB unified memory · 1 TB SSD

128 GB
Unified memory
1
Power user
$7,495
From

Office

One workspace for a small team.

Two Mac Studios in our custom enclosure, connected over Thunderbolt 5 and your local network. Separate model services keep coding work and everyday business requests from competing for the same machine. Designed for a mixed team of up to eight people, subject to workload validation.

2 × M5 Max Mac Studio · 128 GB unified memory and 1 TB SSD per node · 4 TB external document storage · 10 GbE

128 GB × 2
Unified, per node
Up to 8
People, designed for
$15,995
From

Team Pro

More capacity for sustained AI work.

A Linux workstation with two 96 GB NVIDIA professional GPUs for demanding coding agents, shared inference, and CUDA workflows. Designed for an eight-person team with one or two active coding agents alongside business users, with workloads and context limits agreed during setup.

2 × RTX PRO 6000 Blackwell Max-Q · 96 GB GPU memory each · Threadripper 9960X · 256 GB ECC · 8 TB NVMe

192 GB
GPU memory
8
People, designed for
$59,995
From
Show full technical comparison Hide full technical comparison
Comparison of Solo, Pro, Office, and Team Pro
Solo Pro Office Team Pro
Best for One person A developer or power user A small mixed team Sustained team AI work
System Mac mini Mac Studio 2 × Mac Studio in a custom enclosure Linux workstation
Processor M5 Pro, 18-core CPU M5 Max, 18-core CPU 2 × M5 Max, 18-core CPU Threadripper 9960X
Graphics 20-core GPU 40-core GPU 40-core GPU per node 2 × RTX PRO 6000 Blackwell Max-Q
Memory 64 GB unified 128 GB unified 128 GB unified per node 192 GB GPU, plus 256 GB ECC
Storage 1 TB SSD 1 TB SSD 1 TB SSD per node, 4 TB external 8 TB NVMe
Network 2.5 GbE Built-in Ethernet 10 GbE and Thunderbolt 5 Confirmed with the build
Software MLX stack MLX stack MLX, separate model services CUDA stack
From $4,495 $7,495 $15,995 $59,995

Team size describes the intended mix of users, not an unlimited number of simultaneous generations. Supported models, context lengths, and concurrent workloads are confirmed for each configuration. Apple unified memory and NVIDIA dedicated GPU memory have different capacity and performance characteristics. Multi-node memory totals are distributed across machines. Prices are U.S. turnkey prices before tax and shipping.

Workflows / 04

What you can do.

  1. 01

    Code with a local agent

    Explore a repository, make changes, generate tests, and run approved tools using a locally hosted coding model. Your source code can stay on your network.

  2. 02

    Ask your company documents

    Search selected files and receive answers linked to source passages. We connect your documents through retrieval, so you can update the knowledge base without retraining the language model.

  3. 03

    Support finance and operations

    Extract information, organize reports, draft explanations, and work with spreadsheet or calculation tools. Review important outputs and use the underlying tools to verify numbers.

  4. 04

    Write and research with your team

    Draft briefs, proposals, internal communications, and marketing content. Use local files privately, and enable external research or connected services when your workflow calls for them.

Software

What comes installed.

We select and pre-install leading open-weight models, configure the tools for your workflows, and provide our own software, including Alfred Remembers. Team systems add user access, request routing, workload queues, and monitoring.

Apple systems use an Apple Silicon inference stack based on MLX and compatible local runtimes. NVIDIA systems use a CUDA inference stack selected and tested for their models.

Our custom Apple enclosures organize the computers, cabling, and airflow while keeping individual nodes accessible for service. Multi-Mac systems can run independent model services or supported distributed workloads over Thunderbolt 5.

  • Open-weight models 01
  • Alfred Remembers 02
  • Browser AI workspace 03
  • Local model API 04
  • Coding-agent client 05
  • Document search 06
  • Team access and queues Teams

Why local

The intelligence infrastructure belongs to you.

Privacy

Local inference and document retrieval can operate without sending prompts or files to an external model provider.

Control

You choose the models and integrations. Connections that need the network are explicit during setup.

Predictable cost

No per-token charges for the models running on your box. The hardware is a purchase, not a meter.

Your workflows

The system is configured for coding, documents, writing, and operations—not a generic cloud dashboard.

Offline when you need it

Once the selected models and software are installed, the box can work without a network.

Ownership

You keep the hardware. Optional hosted models stay available for work that benefits from them.

Models

Current models, practical configurations.

Our current evaluation set is below. We select the model and quantization for your hardware and work, and qualify new releases before adding them to the supported catalog.

  • Qwen3.8-27B
  • Qwen3.6-35B-A3B
  • Qwen3-Coder-Next

Larger models and longer context windows need more memory and may reduce concurrency. We demonstrate the configured system on a representative workflow so you can judge its usefulness before choosing.

FAQ

Questions worth answering first.

How does my AI remember things?

Alfred Remembers connects the information you add into a persistent local memory that supported AI models can draw on. It helps recover the context behind a question across conversations, documents, and other saved material.

Is it like ChatGPT?

It provides a familiar chat experience using locally hosted models. Quality and features depend on the selected model, tools, and configuration. We demonstrate your intended workflows rather than promise identical results to every hosted assistant.

Can it work offline?

Yes, once the selected models and local software are installed. Workflows that depend on online information or cloud applications need network access.

Can my whole team use it?

Yes. Team configurations provide shared access and managed queues. We size the system around how many people generate responses at once, their context needs, and how heavily they use coding agents.

Does connecting Macs combine their memory?

Supported distributed software can split a model across Macs. Each Mac still has its own memory, and distributed execution has communication overhead. For many office workloads, running separate model services offers a better experience.

Will it replace all our AI subscriptions?

It can replace the workloads that your local models and tools handle well. Savings depend on actual usage, hardware cost, electricity, and support. Optional hosted models can be configured for work that benefits from them.

Can we add capacity later?

Multi-Mac systems can expand with additional qualified nodes. GPU workstations have expansion paths determined by their motherboard, chassis, power, and cooling. We document the supported upgrades with your configuration.

Contact

Let's size your AI.

Tell us who will use it, what they will do, and which data needs to stay local. We'll recommend a box and demonstrate a representative workflow.