How hybrid works
StaffGPT runs your AI employees where your data should live. Keep sensitive work on machines you control, burst everything else to the cloud, and manage it all from one place.
One workforce, two runtimes
Every AI employee can run in the cloud or on your own hardware. You choose per employee — no rebuild, no separate tooling.
Cloud by default
New employees start in our managed cloud so you get value in minutes with nothing to install.
Local for sensitive work
Flip an employee to local and its tasks run entirely on a machine you own. Data never leaves your network.
Hybrid routing
Keep regulated tasks on-premise and let everything else scale to the cloud — automatically, by policy.
How a task gets routed
- 1
An employee receives a task from a studio, schedule, or teammate.
- 2
StaffGPT checks that employee's deployment mode: cloud, local, or hybrid.
- 3
Local and hybrid-sensitive tasks are dispatched to a paired, active machine on your network.
- 4
Everything else runs in the managed cloud. Results return to the same place regardless of where they ran.
Four ways to get a local AI machine
You don't have to source hardware or configure anything yourself. We can set it up at your location, lease you a ready-to-run machine, build one for you, or pair the hardware you already own. Setup and software are included in every option except bring-your-own.
We set up at your location
On-site installation
Starting at $1,200
one-time, per site
A StaffGPT engineer comes to you, installs the machine on your network, loads the models, and hands it over running.
- On-site install and network config
- Models installed and benchmarked
- Machine paired and verified with you
- Team walkthrough before we leave
Local AI Machine Leasing
Hardware as a subscription
Starting at $249
per month, per machine
Lease a configured local AI machine with setup, software, and support included. No capital purchase and no procurement cycle.
- Machine, setup, and software included
- Hardware refresh at end of term
- Ongoing support and model updates
- Cancel or scale machines by term
We build it for you
Custom build, shipped
Starting at $3,400
one-time, hardware included
We spec and assemble the machine, install the models, burn it in, and ship it ready to pair — you own the hardware outright.
- We source and assemble the parts
- Models preinstalled and tested
- Shipped ready to pair
- You own the machine outright
Bring your own machine
Use hardware you own
Included
no additional cost
Already have a capable workstation or GPU box? Pair it yourself in minutes with the Desktop Agent — no purchase required.
- Works with hardware you already own
- Self-serve pairing in minutes
- Free Desktop Agent download
- Meet the requirements below
Prices are illustrative starting points to help you budget — not a quote. Final pricing depends on your chosen specs, number of machines, site location, and term length. Contact us for an exact quote.
Choose the models your machines run
Every model here is open-weight and runs fully offline once installed. Pick the ones that match the work your departments actually do — you can run several on one machine if it has the memory, and add more later without a rebuild.
Recommended tier — single 24 GB GPU
Qwen3 Coder 30B
Our default build-anything model
- Code
- Websites
- Dashboards
- Tables
- Download
- 18.6 GB
- VRAM
- 17 GB
Qwen3 32B
General reasoning and planning
- Reasoning
- Kanban
- Writing
- Tables
- Download
- 20.3 GB
- VRAM
- 18 GB
Qwen2.5 Coder 32B
Code completion and scripting
- Code
- Tables
- Automation
- Download
- 19.9 GB
- VRAM
- 18 GB
Devstral 24B
Works across a whole codebase
- Code
- Automation
- Reasoning
- Download
- 14.3 GB
- VRAM
- 14 GB
Gemma 3 27B
Marketing and long-form writing
- Writing
- Summaries
- Reads images
- Download
- 17.4 GB
- VRAM
- 16 GB
Mistral Small 24B
Balanced all-rounder
- Writing
- Reasoning
- Automation
- Download
- 14.3 GB
- VRAM
- 14 GB
DeepSeek R1 32B
Shows its working, step by step
- Reasoning
- Analysis
- Finance
- Download
- 19.9 GB
- VRAM
- 18 GB
Muse Glimmer 30B
StaffGPT house model
- Code
- Websites
- Reasoning
- Writing
- Download
- 17 GB
- VRAM
- 17 GB
Qwen 3.8 27B
Lighter house alternative
- Code
- Reasoning
- Writing
- Download
- 16 GB
- VRAM
- 16 GB
Runs on modest hardware
Gemma 3 12B
Fast everyday drafting
- Writing
- Reasoning
- Reads images
- Download
- 8.1 GB
- VRAM
- 9 GB
Nomic Embed Text
Search across your own documents
- Search
- Recall
- Download
- 0.3 GB
- VRAM
- 1 GB
Workstation tier — 32 GB+ GPU
Qwen2.5 VL 32B
Reads screenshots, scans and invoices
- Reads images
- Documents
- Tables
- Download
- 21.7 GB
- VRAM
- 20 GB
Llama 3.3 70B
Highest quality, biggest machine
- Reasoning
- Writing
- Code
- Tables
- Download
- 42.5 GB
- VRAM
- 44 GB
Sizes are the actual download for the 4-bit build of each model; the VRAM figure is what we recommend having free to run it comfortably. The one-click installer suggests a small starter set rather than everything — the full catalog is well over 100 GB. House models marked "On request" are installed for you by StaffGPT.
Which model for which job
Most business output — websites, project boards, spreadsheets, game logic, reports — is generated as text or code, so one strong model covers a lot of ground. This is where to start for each kind of work.
- Apps, websites and landing pages
- Qwen3 Coder 30B
- Refactoring across many files
- Devstral 24B
- Kanban boards, plans and roadmaps
- Qwen3 32B
- Spreadsheets, tables and reports
- Qwen3 Coder 30B or DeepSeek R1 32B
- Game logic and simulations
- Qwen3 Coder 30B
- Marketing copy and long documents
- Gemma 3 27B
- Financial analysis and audits
- DeepSeek R1 32B
- Reading scans, invoices and screenshots
- Qwen2.5 VL 32B
- Searching your own documents
- Nomic Embed Text
Generating images and video needs a second runtime
The local setup above runs language and code models. Image, graphics and video generation uses a completely different kind of model that the standard runtime does not serve — so it is a separate add-on rather than part of the one-click installer. Ask us if you need it and we will scope the machine for it.
SDXL
ImagesProduct shots, marketing images, general graphics
~10-12 GB
FLUX.1 dev
ImagesHigher-fidelity images, including text inside the image
~16-24 GB
Wan 2.2
VideoShort video clips at lower resolutions
~16-24 GB
Memory figures are approximate — image and video models vary enormously with resolution, precision and optimisations, so treat these as planning estimates rather than firm requirements. Note that a model that reads images, like Qwen2.5 VL above, cannot draw them.
What you need to run locally
Each machine runs the StaffGPT Desktop Agent plus whichever models you pick from the catalog above. GGUF quantization lets the same weights fit a range of hardware — heavier quantization needs less GPU memory, higher precision needs more.
Minimum
CPU-only, patient
- OS
- macOS 13+, Windows 11, or Ubuntu 22.04+
- CPU
- 8 cores
- Memory
- 16 GB RAM
- GPU
- Optional (CPU inference)
- Model
- Starter models such as Gemma 3 12B (Q4, ~8 GB)
- Storage
- 50 GB free SSD
- Network
- Broadband, outbound HTTPS
Runs the smaller models on CPU/RAM at reduced speed — fine for drafting and review. Best paired with cloud for heavier work.
Recommended
Smooth local employees
- OS
- macOS 14+, Windows 11, or Ubuntu 22.04+
- CPU
- 12+ cores
- Memory
- 32 GB RAM
- GPU
- NVIDIA 24 GB VRAM (RTX 4090) or Apple M-series 32 GB
- Model
- Recommended-tier models (Q4_K_M, ~14-20 GB VRAM)
- Storage
- 100 GB free SSD
- Network
- Broadband, outbound HTTPS
Runs any recommended-tier model at a 4-bit quant fully on the GPU with low latency — the sweet spot for most departments.
Workstation
Full-precision on-prem
- OS
- Windows 11 or Ubuntu 22.04+
- CPU
- 16+ cores
- Memory
- 64 GB RAM
- GPU
- NVIDIA 32 GB+ VRAM (A6000 / dual RTX 4090)
- Model
- Workstation models, or several at once (~20-44 GB VRAM)
- Storage
- 200 GB+ NVMe SSD
- Network
- Broadband; air-gap supported
Runs the largest models, or several smaller ones at once, entirely in-house — including air-gapped deployments.
Every model in the catalog is open-weight — free to download and fully offline once installed. Only outbound HTTPS is required — no inbound ports. The agent authenticates with its own key and never exposes your machine to the internet.
Connect a computer in four steps
Pairing binds a machine to your account with its own cryptographic identity. It takes about a minute.
Add the computer
On your Machines page, click Add Computer to generate a one-time pairing code that is valid for five minutes.
Install the desktop agent
Download and open the StaffGPT Desktop Agent on the machine you want to use. It generates a private key that never leaves the device.
Enter the code
Paste the pairing code into the agent. It registers the machine and binds its public key to your account.
Go active
The machine starts sending heartbeats and appears as Active — Online. You can revoke it at any time, instantly and permanently.
Download one file. Double-click it. Done.
Every script and piece of software your machine needs is bundled into a single installer. Download it from your Machines page, double-click, and the machine appears online in your workspace — no terminal, no commands, no configuration.
Windows
StaffGPT-Setup.cmdDouble-click the file
macOS
StaffGPT-Setup-macos.zipUnzip, then double-click
Linux
StaffGPT-Setup-linux.zipUnzip, then double-click or run
What the installer handles for you
- Checks your CPU, memory, disk, and GPU against the tiers below
- Installs the correct Node.js runtime if it isn't already present
- Downloads the StaffGPT Desktop Agent and verifies it runs
- Pairs the machine to your workspace automatically
- Offers to install the model runtime and pull your chosen model
- Registers a background service so the agent restarts with the computer
- Confirms the first heartbeat reached your workspace before finishing
The installer is per-user: it never asks for an administrator password and installs nothing system-wide. Everything lands in your home folder and can be fully removed with a single uninstall command. Your pairing code is already embedded, so there is nothing to type or copy.
Installation & onboarding, step by step
A repeatable standard operating procedure for turning a machine you own into a working local runtime. Most machines are live in under 15 minutes.
Prefer to do it manually?
The individual scripts remain available for air-gapped sites, fleet imaging, and administrators who want to inspect or automate each step. The standard operating procedure below documents that path in full — the one-click installer performs exactly these steps for you.
- 1
Pre-qualify the machine
Run the pre-qualification script to confirm the CPU, memory, disk, and GPU meet a tier. It only reads specs and sends nothing.
- 2
Install prerequisites
Install Node.js 18+ and a local model runtime. No admin agents or inbound ports are required — only outbound HTTPS.
- 3
Download & launch the Desktop Agent
Get the agent from your Machines page and run it on the target machine. It generates a private key that never leaves the device.
- 4
Pair with a one-time code
Click Add Computer, copy the five-minute code, and paste it into the agent to bind the machine to your account.
- 5
Validate
Confirm the machine shows Active — Online, then assign a test employee in local mode and run one task end to end.
- 6
Harden & keep online
Install the agent as a service so it restarts on boot and stays online. Revoke instantly from the web if a machine is lost.
What local costs
- Guided install & onboardingContact sales
Optional white-glove setup, hardening, and first-employee validation per site.
- Machine pairingFree
Unlimited paired machines on every plan — no per-seat device fee.
- Local runtimeContact sales
Priced below cloud compute; work that runs locally offsets your cloud usage.
- Cloud computeUsage-based
Charged only for tasks that actually run in the managed cloud.
Figures are finalized with sales based on scale and support needs. Self-serve pairing and the Desktop Agent are always free.
Pre-qualify, then verify connectivity
Two real scripts you download from your Machines page: one checks the hardware, the other proves the machine can pair and stay online.
First run the pre-qualification check to confirm the machine meets a tier. Then pair the Desktop Agent with a one-time code and leave it running — it registers, sends signed heartbeats, and the machine turns Active �� Online.
# 1. Pre-qualify: reads CPU / RAM / disk / GPU, sends nothing
node prequalify.mjs
# → prints the tier the machine qualifies for
# 2. On your Machines page: Add Computer → copy the pairing code, then:
node staffgpt-agent.mjs pair --url https://your-staffgpt-url --code XXXX-XXXX
# → "Paired. Machine id: ..."
# 3. Keep it online and watch it turn Active — Online
node staffgpt-agent.mjs run
# → repeating "heartbeat ok"A healthy machine passes all of these
- Pre-qualify reports at least the Minimum tier.
- Pairing succeeds and returns a machine ID.
- The heartbeat challenge is signed and accepted.
- The machine appears as Active — Online on your Machines page.
- Revoking the machine immediately blocks further heartbeats.
Hybrid keeps your cloud bill down
Every plan can run hybrid. Work that runs on your own machines does not consume cloud compute — so bringing hardware lowers your usage.
Cloud
Managed employees on our infrastructure. Start here, add machines whenever you want.
- No hardware needed
- Elastic scale
- Usage-based compute
Bring your own machine
Pair machines you already own and shift local-eligible work off the cloud.
- Lower cloud usage
- Data stays on-prem
- Unlimited paired machines
Enterprise / on-prem
Fully local or air-gapped deployments with dedicated support and controls.
- Air-gap ready
- SSO & audit
- Dedicated support
Pairing machines is free on every plan. See the full plan comparison for included compute and limits.
Put your own hardware to work
Pair a machine in a minute, or talk to us about a fully on-premise deployment.