Not every job is a bigger model.
Twelve real AI workload types, worked backwards from the human and the job to be done first - then the machine. A small, specialized model that deeply knows one world can be the right tool precisely because it is small. Find your workload below, then use the calculator to see the real buy-vs-rent verdict.
Nine of these are honestly local. Three are honestly cloud.
Most AI-workload guidance starts with "which model is biggest." We start with the job - and the unit that actually measures it, because tokens/sec is not the right yardstick for a world model or a voice agent.
Chat & reasoning
LocalExamples: General assistant conversation, Q&A, summarization, everyday agentic tasks
Measured in: tokens/sec
Best fit: A laptop-class Puppy (MacBook Pro 16 M5 Max 128GB)
The default job the calculator above is built around - interactive, single-user, exactly what a 128GB unified-memory machine is sized for.
Coding agents
LocalExamples: Repo-aware generation, code review, test generation, agentic development loops
Measured in: tokens/sec, quality-weighted
Best fit: The same laptop-class Puppy, or a DGX Spark-class desk machine for larger repos
Coder-tuned models in the 30B-class MoE range fit the 128GB envelope comfortably and keep private repositories private by never leaving the device.
Small Specialist Models (SLM)
LocalExamples: Routing, classification, extraction, guardrails, structured output, high-frequency narrow tasks
Measured in: tokens/sec (very high) & time-to-first-token
Best fit: Anything - even the phone already in your pocket
Small is not a compromise here, it is the correct tool: a 1-4B model fine-tuned for one narrow job runs at hundreds of tokens/sec on almost any hardware, with near-zero latency. Most of the actual jobs a business runs thousands of times a day are this shape, not a giant chat model.
World models & simulation
LocalExamples: A model trained to deeply understand one world - a game environment, a robot's physical sensorimotor space, a business's own operational patterns - and run it forward to predict, plan, and simulate
Measured in: simulation steps/sec or frames/sec, not tokens
Best fit: Edge/robotics-class hardware (NVIDIA Jetson Orin Nano Super)
World models earn their keep by being small and specialized, not huge and general - they recreate the one world they were trained on well enough to plan inside it. Real-time response matters more than raw parameter count, which is exactly what edge silicon is built for.
Vision & multimodal understanding
LocalExamples: Reading documents, screenshots, photos, and video frames - understanding, not generating
Measured in: tokens/sec on the vision-language path
Best fit: A laptop-class Puppy with a vision-capable model
Vision-language models in the 30B class fit the same unified-memory envelope as text models and keep sensitive images (financial documents, medical scans) off the network entirely.
Voice & real-time streaming
LocalExamples: Transcription, live conversation, always-on voice agents
Measured in: real-time factor & time-to-first-token, not raw throughput
Best fit: Any laptop-class or above Puppy
What matters here is never falling behind the live audio stream, not maximum tokens/sec - a different number to optimize for, and one local hardware handles well since there is no network round-trip in the loop.
Embeddings & retrieval (RAG indexing)
LocalExamples: Turning documents, emails, and records into a searchable private index
Measured in: documents indexed/sec
Best fit: Almost anything, even a Mac mini or a Jetson - run it overnight
This is a batch, background-friendly job with no latency pressure - the cheapest, smallest machine that can hold the index comfortably is the right one.
Fine-tuning & LoRA (teaching a model your world)
LocalExamples: Adapting a small or mid-size model to a business's own documents, tone, or domain
Measured in: training steps/sec - a different memory shape than inference (gradients and optimizer state, not just weights)
Best fit: A prosumer 24/7 Puppy (NVIDIA DGX Spark) or a Mac Studio M3 Ultra
LoRA/QLoRA adaptation of 8B-30B models fits comfortably in 128GB-class unified memory - full fine-tunes of frontier-scale models are the one part of this job that still needs the cloud.
Image & short-form generation
HybridExamples: Distilled, fast diffusion models for stills and short clips - not studio VFX
Measured in: images or seconds of video per minute
Best fit: A laptop-class Puppy with a real discrete GPU (Lenovo Legion Pro 7i, RTX 5090)
Lightweight, distilled generation models genuinely run well locally on real GPU silicon - it is only the heaviest, highest-fidelity, or studio-scale generation that needs to burst to the cloud, not generation as a category.
Studio-scale video, VFX & codec at scale
CloudExamples: Feature-length VFX render farms, studio transcoding pipelines, live broadcast encoding, real-time ray-traced previsualization
Measured in: frames/sec, at scale
Best fit: Cloud GPU/TPU burst or purpose-built encoder hardware
Throughput-bound batch workloads with enormous parallel frame counts - exactly what data-center clusters exist for. A single machine's unified-memory advantage does not apply; this is compute-bound, not memory-bandwidth-bound.
Music & audio production at professional scale
HybridExamples: Multi-track mastering farms, large-catalog audio restoration, AI-assisted scoring across a full session library
Measured in: sessions/hour at catalog scale
Best fit: Local for the interactive session, cloud burst for the batch render
The interactive DAW session genuinely benefits from Apple silicon's media engines and low local latency. Batch rendering across a catalog is a parallel workload that scales better in the cloud.
Frontier scientific research & simulation
CloudExamples: Climate and physics simulation, protein folding and molecular dynamics, large-scale numerical methods that push toward new discovery
Measured in: varies by domain - HPC-class metrics, not tokens
Best fit: Cloud/HPC-class compute for the frontier run; local for the modeling, prep, and analysis around it
This is precisely the work that helps humanity understand the universe and our own consciousness a little better - and it needs real HPC clusters, not a laptop. Puppy's honest job here is the surrounding work: private data prep, local analysis, drafting, and orchestration - burst-routing the actual frontier compute to the right cloud partner, at the right price, with a receipt.
Every number moves live - nothing is fixed but the formula.
This tool models the token-shaped workloads above - chat, coding, small specialist models, vision, voice - the ones a bandwidth-and-power calculation actually applies to.
Puppy · Buy vs Rent
Own your compute, or rent it? Drag the dials.
Cloud is 38.1× cheaper per token - you're not running the box enough to amortize it.
Estimates from the memory-bound decode model (614 GB/s M5 Max, community-verified). Confidence: high on method, moderate on absolute tok/s - to be replaced with our own measured Compute Passport numbers.
Nine machines, three jobs to be done - the first decision reduced to what it should be. Every one of the 🤫 Puppy 100 is still one tap away behind "See all 100." Each shows whether its bandwidth and power figures are a published vendor spec or an honest rung-level estimate; rungs 7-10 in the full list are rentals, not machines to own.
Where this actually shows up in your life.
Your CPA workflow, at tax time
Any American household or small business ownerEvery January you dig through a dozen accounts and email threads for every 1099, W-2, 1098, and K-1 tied to your SSN or your business's ITIN or EIN. By your consent, Agent One connects to the financial accounts you choose to link, watches for the tax documents those institutions issue, and assembles them into one packet for your CPA - with a receipt showing exactly what was accessed and when.
Your SSN or ITIN is used only to match documents a payer already issued to you - never transmitted to a new party without your explicit, scoped consent. Every access is written to your own receipt ledger, readable and revocable by you at any time.
Quarterly books, prepared for the bookkeeper
Small business ownersBy consent, Agent One categorizes expenses from linked business accounts and hands a clean quarterly packet to your bookkeeper or accountant - the same consent-and-receipt architecture as the tax workflow above, running continuously instead of once a year.
One consolidated view, for the advisor
Family office staff and their principalsAccount statements scattered across multiple institutions, consolidated by consent into one packet the principal's RIA or wealth advisor can actually use - never held by us, never seen by us, read only by the people the principal names.
Token generation, for engineers and creators
Engineers and creators who need real computeThe workload this whole page and the calculator above are built around: own the baseline, burst to the ceiling, pay cloud prices only for the minutes you actually need them.
One is a product of Hushh Technologies Corporation (brand: 🤫 “hussh”), an independent company. One runs on third-party silicon, systems, and cloud; all company names are used solely to describe the platforms on which One software runs. Hushh Technologies is not affiliated with, endorsed by, sponsored by, or partnered with any company named.