Hybrid Architecture

How hybrid works

StaffGPT runs your AI employees where your data should live. Keep sensitive work on machines you control, burst everything else to the cloud, and manage it all from one place.

Blend both

One workforce, two runtimes

Every AI employee can run in the cloud or on your own hardware. You choose per employee — no rebuild, no separate tooling.

Cloud by default

New employees start in our managed cloud so you get value in minutes with nothing to install.

Local for sensitive work

Flip an employee to local and its tasks run entirely on a machine you own. Data never leaves your network.

Hybrid routing

Keep regulated tasks on-premise and let everything else scale to the cloud — automatically, by policy.

How a task gets routed

  1. 1

    An employee receives a task from a studio, schedule, or teammate.

  2. 2

    StaffGPT checks that employee's deployment mode: cloud, local, or hybrid.

  3. 3

    Local and hybrid-sensitive tasks are dispatched to a paired, active machine on your network.

  4. 4

    Everything else runs in the managed cloud. Results return to the same place regardless of where they ran.

Getting your machine

Four ways to get a local AI machine

You don't have to source hardware or configure anything yourself. We can set it up at your location, lease you a ready-to-run machine, build one for you, or pair the hardware you already own. Setup and software are included in every option except bring-your-own.

We set up at your location

On-site installation

Starting at $1,200

one-time, per site

A StaffGPT engineer comes to you, installs the machine on your network, loads the models, and hands it over running.

  • On-site install and network config
  • Models installed and benchmarked
  • Machine paired and verified with you
  • Team walkthrough before we leave

Local AI Machine Leasing

Hardware as a subscription

Starting at $249

per month, per machine

Lease a configured local AI machine with setup, software, and support included. No capital purchase and no procurement cycle.

  • Machine, setup, and software included
  • Hardware refresh at end of term
  • Ongoing support and model updates
  • Cancel or scale machines by term

We build it for you

Custom build, shipped

Starting at $3,400

one-time, hardware included

We spec and assemble the machine, install the models, burn it in, and ship it ready to pair — you own the hardware outright.

  • We source and assemble the parts
  • Models preinstalled and tested
  • Shipped ready to pair
  • You own the machine outright

Bring your own machine

Use hardware you own

Included

no additional cost

Already have a capable workstation or GPU box? Pair it yourself in minutes with the Desktop Agent — no purchase required.

  • Works with hardware you already own
  • Self-serve pairing in minutes
  • Free Desktop Agent download
  • Meet the requirements below
Talk to us about a machine

Prices are illustrative starting points to help you budget — not a quote. Final pricing depends on your chosen specs, number of machines, site location, and term length. Contact us for an exact quote.

Model catalog

Choose the models your machines run

Every model here is open-weight and runs fully offline once installed. Pick the ones that match the work your departments actually do — you can run several on one machine if it has the memory, and add more later without a rebuild.

Recommended tier — single 24 GB GPU

Qwen3 Coder 30B

Default

Our default build-anything model

  • Code
  • Websites
  • Dashboards
  • Tables
Download
18.6 GB
VRAM
17 GB

Qwen3 32B

General reasoning and planning

  • Reasoning
  • Kanban
  • Writing
  • Tables
Download
20.3 GB
VRAM
18 GB

Qwen2.5 Coder 32B

Code completion and scripting

  • Code
  • Tables
  • Automation
Download
19.9 GB
VRAM
18 GB

Devstral 24B

Works across a whole codebase

  • Code
  • Automation
  • Reasoning
Download
14.3 GB
VRAM
14 GB

Gemma 3 27B

Marketing and long-form writing

  • Writing
  • Summaries
  • Reads images
Download
17.4 GB
VRAM
16 GB

Mistral Small 24B

Balanced all-rounder

  • Writing
  • Reasoning
  • Automation
Download
14.3 GB
VRAM
14 GB

DeepSeek R1 32B

Shows its working, step by step

  • Reasoning
  • Analysis
  • Finance
Download
19.9 GB
VRAM
18 GB

Muse Glimmer 30B

On request

StaffGPT house model

  • Code
  • Websites
  • Reasoning
  • Writing
Download
17 GB
VRAM
17 GB

Qwen 3.8 27B

On request

Lighter house alternative

  • Code
  • Reasoning
  • Writing
Download
16 GB
VRAM
16 GB

Runs on modest hardware

Gemma 3 12B

Fast everyday drafting

  • Writing
  • Reasoning
  • Reads images
Download
8.1 GB
VRAM
9 GB

Nomic Embed Text

Search across your own documents

  • Search
  • Recall
Download
0.3 GB
VRAM
1 GB

Workstation tier — 32 GB+ GPU

Qwen2.5 VL 32B

Reads screenshots, scans and invoices

  • Reads images
  • Documents
  • Tables
Download
21.7 GB
VRAM
20 GB

Llama 3.3 70B

Highest quality, biggest machine

  • Reasoning
  • Writing
  • Code
  • Tables
Download
42.5 GB
VRAM
44 GB

Sizes are the actual download for the 4-bit build of each model; the VRAM figure is what we recommend having free to run it comfortably. The one-click installer suggests a small starter set rather than everything — the full catalog is well over 100 GB. House models marked "On request" are installed for you by StaffGPT.

Which model for which job

Most business output — websites, project boards, spreadsheets, game logic, reports — is generated as text or code, so one strong model covers a lot of ground. This is where to start for each kind of work.

Apps, websites and landing pages
Qwen3 Coder 30B
Refactoring across many files
Devstral 24B
Kanban boards, plans and roadmaps
Qwen3 32B
Spreadsheets, tables and reports
Qwen3 Coder 30B or DeepSeek R1 32B
Game logic and simulations
Qwen3 Coder 30B
Marketing copy and long documents
Gemma 3 27B
Financial analysis and audits
DeepSeek R1 32B
Reading scans, invoices and screenshots
Qwen2.5 VL 32B
Searching your own documents
Nomic Embed Text
Images and video

Generating images and video needs a second runtime

The local setup above runs language and code models. Image, graphics and video generation uses a completely different kind of model that the standard runtime does not serve — so it is a separate add-on rather than part of the one-click installer. Ask us if you need it and we will scope the machine for it.

Separate runtime required

SDXL

Images

Product shots, marketing images, general graphics

~10-12 GB

FLUX.1 dev

Images

Higher-fidelity images, including text inside the image

~16-24 GB

Wan 2.2

Video

Short video clips at lower resolutions

~16-24 GB

Memory figures are approximate — image and video models vary enormously with resolution, precision and optimisations, so treat these as planning estimates rather than firm requirements. Note that a model that reads images, like Qwen2.5 VL above, cannot draw them.

Equipment requirements

What you need to run locally

Each machine runs the StaffGPT Desktop Agent plus whichever models you pick from the catalog above. GGUF quantization lets the same weights fit a range of hardware — heavier quantization needs less GPU memory, higher precision needs more.

Minimum

CPU-only, patient

OS
macOS 13+, Windows 11, or Ubuntu 22.04+
CPU
8 cores
Memory
16 GB RAM
GPU
Optional (CPU inference)
Model
Starter models such as Gemma 3 12B (Q4, ~8 GB)
Storage
50 GB free SSD
Network
Broadband, outbound HTTPS

Runs the smaller models on CPU/RAM at reduced speed — fine for drafting and review. Best paired with cloud for heavier work.

Recommended

Smooth local employees

OS
macOS 14+, Windows 11, or Ubuntu 22.04+
CPU
12+ cores
Memory
32 GB RAM
GPU
NVIDIA 24 GB VRAM (RTX 4090) or Apple M-series 32 GB
Model
Recommended-tier models (Q4_K_M, ~14-20 GB VRAM)
Storage
100 GB free SSD
Network
Broadband, outbound HTTPS

Runs any recommended-tier model at a 4-bit quant fully on the GPU with low latency — the sweet spot for most departments.

Workstation

Full-precision on-prem

OS
Windows 11 or Ubuntu 22.04+
CPU
16+ cores
Memory
64 GB RAM
GPU
NVIDIA 32 GB+ VRAM (A6000 / dual RTX 4090)
Model
Workstation models, or several at once (~20-44 GB VRAM)
Storage
200 GB+ NVMe SSD
Network
Broadband; air-gap supported

Runs the largest models, or several smaller ones at once, entirely in-house — including air-gapped deployments.

Every model in the catalog is open-weight — free to download and fully offline once installed. Only outbound HTTPS is required — no inbound ports. The agent authenticates with its own key and never exposes your machine to the internet.

Pair a machine

Connect a computer in four steps

Pairing binds a machine to your account with its own cryptographic identity. It takes about a minute.

01

Add the computer

On your Machines page, click Add Computer to generate a one-time pairing code that is valid for five minutes.

02

Install the desktop agent

Download and open the StaffGPT Desktop Agent on the machine you want to use. It generates a private key that never leaves the device.

03

Enter the code

Paste the pairing code into the agent. It registers the machine and binds its public key to your account.

04

Go active

The machine starts sending heartbeats and appears as Active — Online. You can revoke it at any time, instantly and permanently.

One-click setup

Download one file. Double-click it. Done.

Every script and piece of software your machine needs is bundled into a single installer. Download it from your Machines page, double-click, and the machine appears online in your workspace — no terminal, no commands, no configuration.

Windows

StaffGPT-Setup.cmd

Double-click the file

macOS

StaffGPT-Setup-macos.zip

Unzip, then double-click

Linux

StaffGPT-Setup-linux.zip

Unzip, then double-click or run

What the installer handles for you

  • Checks your CPU, memory, disk, and GPU against the tiers below
  • Installs the correct Node.js runtime if it isn't already present
  • Downloads the StaffGPT Desktop Agent and verifies it runs
  • Pairs the machine to your workspace automatically
  • Offers to install the model runtime and pull your chosen model
  • Registers a background service so the agent restarts with the computer
  • Confirms the first heartbeat reached your workspace before finishing

The installer is per-user: it never asks for an administrator password and installs nothing system-wide. Everything lands in your home folder and can be fully removed with a single uninstall command. Your pairing code is already embedded, so there is nothing to type or copy.

Local install

Installation & onboarding, step by step

A repeatable standard operating procedure for turning a machine you own into a working local runtime. Most machines are live in under 15 minutes.

Prefer to do it manually?

The individual scripts remain available for air-gapped sites, fleet imaging, and administrators who want to inspect or automate each step. The standard operating procedure below documents that path in full — the one-click installer performs exactly these steps for you.

  1. 1

    Pre-qualify the machine

    Run the pre-qualification script to confirm the CPU, memory, disk, and GPU meet a tier. It only reads specs and sends nothing.

  2. 2

    Install prerequisites

    Install Node.js 18+ and a local model runtime. No admin agents or inbound ports are required — only outbound HTTPS.

  3. 3

    Download & launch the Desktop Agent

    Get the agent from your Machines page and run it on the target machine. It generates a private key that never leaves the device.

  4. 4

    Pair with a one-time code

    Click Add Computer, copy the five-minute code, and paste it into the agent to bind the machine to your account.

  5. 5

    Validate

    Confirm the machine shows Active — Online, then assign a test employee in local mode and run one task end to end.

  6. 6

    Harden & keep online

    Install the agent as a service so it restarts on boot and stays online. Revoke instantly from the web if a machine is lost.

What local costs

  • Guided install & onboardingContact sales

    Optional white-glove setup, hardening, and first-employee validation per site.

  • Machine pairingFree

    Unlimited paired machines on every plan — no per-seat device fee.

  • Local runtimeContact sales

    Priced below cloud compute; work that runs locally offsets your cloud usage.

  • Cloud computeUsage-based

    Charged only for tasks that actually run in the managed cloud.

Figures are finalized with sales based on scale and support needs. Self-serve pairing and the Desktop Agent are always free.

Test your machine

Pre-qualify, then verify connectivity

Two real scripts you download from your Machines page: one checks the hardware, the other proves the machine can pair and stay online.

First run the pre-qualification check to confirm the machine meets a tier. Then pair the Desktop Agent with a one-time code and leave it running — it registers, sends signed heartbeats, and the machine turns Active �� Online.

prequalify.mjs + staffgpt-agent.mjs — download both from your Machines page
# 1. Pre-qualify: reads CPU / RAM / disk / GPU, sends nothing
node prequalify.mjs
# → prints the tier the machine qualifies for

# 2. On your Machines page: Add Computer → copy the pairing code, then:
node staffgpt-agent.mjs pair --url https://your-staffgpt-url --code XXXX-XXXX
# → "Paired. Machine id: ..."

# 3. Keep it online and watch it turn Active — Online
node staffgpt-agent.mjs run
# → repeating "heartbeat ok"

A healthy machine passes all of these

  • Pre-qualify reports at least the Minimum tier.
  • Pairing succeeds and returns a machine ID.
  • The heartbeat challenge is signed and accepted.
  • The machine appears as Active — Online on your Machines page.
  • Revoking the machine immediately blocks further heartbeats.
Pricing

Hybrid keeps your cloud bill down

Every plan can run hybrid. Work that runs on your own machines does not consume cloud compute — so bringing hardware lowers your usage.

Cloud

Included

Managed employees on our infrastructure. Start here, add machines whenever you want.

  • No hardware needed
  • Elastic scale
  • Usage-based compute

Bring your own machine

Your hardware

Pair machines you already own and shift local-eligible work off the cloud.

  • Lower cloud usage
  • Data stays on-prem
  • Unlimited paired machines

Enterprise / on-prem

Custom

Fully local or air-gapped deployments with dedicated support and controls.

  • Air-gap ready
  • SSO & audit
  • Dedicated support
Compare all plans

Pairing machines is free on every plan. See the full plan comparison for included compute and limits.

Put your own hardware to work

Pair a machine in a minute, or talk to us about a fully on-premise deployment.