🤫husshhussh
🤫husshhusshOnePuppy
🤫 Puppy One · work backwards from your workload

Not every job is a bigger model.

Twelve real AI workload types, worked backwards from the human and the job to be done first - then the machine. A small, specialized model that deeply knows one world can be the right tool precisely because it is small. Find your workload below, then use the calculator to see the real buy-vs-rent verdict.

Jump to the calculatorThe Puppy One Max reference machine
The ontology

Nine of these are honestly local. Three are honestly cloud.

Most AI-workload guidance starts with "which model is biggest." We start with the job - and the unit that actually measures it, because tokens/sec is not the right yardstick for a world model or a voice agent.

Chat & reasoning

Local

Examples: General assistant conversation, Q&A, summarization, everyday agentic tasks

Measured in: tokens/sec

Best fit: A laptop-class Puppy (MacBook Pro 16 M5 Max 128GB)

The default job the calculator above is built around - interactive, single-user, exactly what a 128GB unified-memory machine is sized for.

Coding agents

Local

Examples: Repo-aware generation, code review, test generation, agentic development loops

Measured in: tokens/sec, quality-weighted

Best fit: The same laptop-class Puppy, or a DGX Spark-class desk machine for larger repos

Coder-tuned models in the 30B-class MoE range fit the 128GB envelope comfortably and keep private repositories private by never leaving the device.

Small Specialist Models (SLM)

Local

Examples: Routing, classification, extraction, guardrails, structured output, high-frequency narrow tasks

Measured in: tokens/sec (very high) & time-to-first-token

Best fit: Anything - even the phone already in your pocket

Small is not a compromise here, it is the correct tool: a 1-4B model fine-tuned for one narrow job runs at hundreds of tokens/sec on almost any hardware, with near-zero latency. Most of the actual jobs a business runs thousands of times a day are this shape, not a giant chat model.

World models & simulation

Local

Examples: A model trained to deeply understand one world - a game environment, a robot's physical sensorimotor space, a business's own operational patterns - and run it forward to predict, plan, and simulate

Measured in: simulation steps/sec or frames/sec, not tokens

Best fit: Edge/robotics-class hardware (NVIDIA Jetson Orin Nano Super)

World models earn their keep by being small and specialized, not huge and general - they recreate the one world they were trained on well enough to plan inside it. Real-time response matters more than raw parameter count, which is exactly what edge silicon is built for.

Vision & multimodal understanding

Local

Examples: Reading documents, screenshots, photos, and video frames - understanding, not generating

Measured in: tokens/sec on the vision-language path

Best fit: A laptop-class Puppy with a vision-capable model

Vision-language models in the 30B class fit the same unified-memory envelope as text models and keep sensitive images (financial documents, medical scans) off the network entirely.

Voice & real-time streaming

Local

Examples: Transcription, live conversation, always-on voice agents

Measured in: real-time factor & time-to-first-token, not raw throughput

Best fit: Any laptop-class or above Puppy

What matters here is never falling behind the live audio stream, not maximum tokens/sec - a different number to optimize for, and one local hardware handles well since there is no network round-trip in the loop.

Embeddings & retrieval (RAG indexing)

Local

Examples: Turning documents, emails, and records into a searchable private index

Measured in: documents indexed/sec

Best fit: Almost anything, even a Mac mini or a Jetson - run it overnight

This is a batch, background-friendly job with no latency pressure - the cheapest, smallest machine that can hold the index comfortably is the right one.

Fine-tuning & LoRA (teaching a model your world)

Local

Examples: Adapting a small or mid-size model to a business's own documents, tone, or domain

Measured in: training steps/sec - a different memory shape than inference (gradients and optimizer state, not just weights)

Best fit: A prosumer 24/7 Puppy (NVIDIA DGX Spark) or a Mac Studio M3 Ultra

LoRA/QLoRA adaptation of 8B-30B models fits comfortably in 128GB-class unified memory - full fine-tunes of frontier-scale models are the one part of this job that still needs the cloud.

Image & short-form generation

Hybrid

Examples: Distilled, fast diffusion models for stills and short clips - not studio VFX

Measured in: images or seconds of video per minute

Best fit: A laptop-class Puppy with a real discrete GPU (Lenovo Legion Pro 7i, RTX 5090)

Lightweight, distilled generation models genuinely run well locally on real GPU silicon - it is only the heaviest, highest-fidelity, or studio-scale generation that needs to burst to the cloud, not generation as a category.

Studio-scale video, VFX & codec at scale

Cloud

Examples: Feature-length VFX render farms, studio transcoding pipelines, live broadcast encoding, real-time ray-traced previsualization

Measured in: frames/sec, at scale

Best fit: Cloud GPU/TPU burst or purpose-built encoder hardware

Throughput-bound batch workloads with enormous parallel frame counts - exactly what data-center clusters exist for. A single machine's unified-memory advantage does not apply; this is compute-bound, not memory-bandwidth-bound.

Music & audio production at professional scale

Hybrid

Examples: Multi-track mastering farms, large-catalog audio restoration, AI-assisted scoring across a full session library

Measured in: sessions/hour at catalog scale

Best fit: Local for the interactive session, cloud burst for the batch render

The interactive DAW session genuinely benefits from Apple silicon's media engines and low local latency. Batch rendering across a catalog is a parallel workload that scales better in the cloud.

Frontier scientific research & simulation

Cloud

Examples: Climate and physics simulation, protein folding and molecular dynamics, large-scale numerical methods that push toward new discovery

Measured in: varies by domain - HPC-class metrics, not tokens

Best fit: Cloud/HPC-class compute for the frontier run; local for the modeling, prep, and analysis around it

This is precisely the work that helps humanity understand the universe and our own consciousness a little better - and it needs real HPC clusters, not a laptop. Puppy's honest job here is the surrounding work: private data prep, local analysis, drafting, and orchestration - burst-routing the actual frontier compute to the right cloud partner, at the right price, with a receipt.

The calculator

Every number moves live - nothing is fixed but the formula.

This tool models the token-shaped workloads above - chat, coding, small specialist models, vision, voice - the ones a bandwidth-and-power calculation actually applies to.

🤫

Puppy · Buy vs Rent

Own your compute, or rent it? Drag the dials.

↗RENT — use cloud

Cloud is 38.1× cheaper per token - you're not running the box enough to amortize it.

Local (this machine)$22.87 / 1M tok
Cloud alternative$0.600 / 1M tok
Does not pay back within the 3-yr hardware life on this workload (~19.4B tokens).
WorkloadQuality · reference workload
How hard you run it
Active generation8 hr / day
0.5-3 hr = personal · 8 hr = workday · 24 hr = server duty
Electricity$0.11 / kWh
Cloud alternative
Cloud price$0.600 / 1M tok
Data must stay private
Can't send to cloud - the sovereignty case.
Decode speed
11.1
tok / sec
Energy efficiency
568
tokens / Wh
Marginal cost
$0.194
/ 1M · energy only
Amortized all-in
$22.87
/ 1M · at this usage
Pick your Puppy9 machines, 3 jobs to be done
Start here
Zero to near-zero new spend - the lowest-friction way to feel a Puppy work today.
Do real work daily
Hours-a-day interactive use - coding, drafting, research, agentic loops that run while you work.
Never sleeps
24/7 duty, heavy models, the machines built to never power down.
✓ Vendor spec · Apple M5 Max published spec
The machinetap to edit
Apple M5 Max · 614 GB/s · 70 W · $7,903 landed · 3-yr life

Estimates from the memory-bound decode model (614 GB/s M5 Max, community-verified). Confidence: high on method, moderate on absolute tok/s - to be replaced with our own measured Compute Passport numbers.

Nine machines, three jobs to be done - the first decision reduced to what it should be. Every one of the 🤫 Puppy 100 is still one tap away behind "See all 100." Each shows whether its bandwidth and power figures are a published vendor spec or an honest rung-level estimate; rungs 7-10 in the full list are rentals, not machines to own.

Real workflows, working backwards from you

Where this actually shows up in your life.

Your CPA workflow, at tax time

Any American household or small business owner

Every January you dig through a dozen accounts and email threads for every 1099, W-2, 1098, and K-1 tied to your SSN or your business's ITIN or EIN. By your consent, Agent One connects to the financial accounts you choose to link, watches for the tax documents those institutions issue, and assembles them into one packet for your CPA - with a receipt showing exactly what was accessed and when.

Your SSN or ITIN is used only to match documents a payer already issued to you - never transmitted to a new party without your explicit, scoped consent. Every access is written to your own receipt ledger, readable and revocable by you at any time.

Quarterly books, prepared for the bookkeeper

Small business owners

By consent, Agent One categorizes expenses from linked business accounts and hands a clean quarterly packet to your bookkeeper or accountant - the same consent-and-receipt architecture as the tax workflow above, running continuously instead of once a year.

One consolidated view, for the advisor

Family office staff and their principals

Account statements scattered across multiple institutions, consolidated by consent into one packet the principal's RIA or wealth advisor can actually use - never held by us, never seen by us, read only by the people the principal names.

Token generation, for engineers and creators

Engineers and creators who need real compute

The workload this whole page and the calculator above are built around: own the baseline, burst to the ceiling, pay cloud prices only for the minutes you actually need them.

One is a product of Hushh Technologies Corporation (brand: 🤫 “hussh”), an independent company. One runs on third-party silicon, systems, and cloud; all company names are used solely to describe the platforms on which One software runs. Hushh Technologies is not affiliated with, endorsed by, sponsored by, or partnered with any company named.

Products

  • Agent One
  • The 🤫 One app
  • Puppy One
  • Which Puppy is right for you?
  • The Puppy 100
  • Tag One
  • The 🤫 Store
  • The 🤫 One Card
  • Pricing
  • Claim your One
  • The product roadmap

🤫 Yellow Pages

  • The 🤫 Yellow Pages
  • Discover in the feed
  • Find a local expert
  • Coverage & markets
  • Connect - in Agent One
  • Ping an expert

Business & Enterprise

  • 🤫 for Business
  • Small & medium business
  • 🤫 Concierge (VVIP)
  • 🤫 for the Enterprise
  • Industry solutions
  • Federal government & agencies
  • 🇺🇸 Defense & national security
  • For advisors (RIAs)
  • Partner Portal
  • One for Sellers
  • Developers

Watch, read & learn

  • The media library
  • The 🤫 Feed
  • See it in a minute
  • Listen - the podcasts
  • Blogs
  • Research & papers
  • Guides - by topic
  • Academy
  • The Heartbeat
  • Wiki

Company & open

  • About
  • Team
  • Investors
  • Fund A
  • Building in the open
  • Newsroom & press
  • Release notes
  • Careers
  • Contact
  • Explore - the whole site, mapped
  • Sitemap

Trust, rights & gratitude

  • The Hussh Protocol (PCHP)
  • Day 0 Trusted Circle
  • The case - a right, made enforceable
  • Data-rights landscape
  • Accessibility
  • 🤫 Champions of the Community
  • 🤫 Faculty - the professors
  • Gratitude - people we admire
  • The 1024 - humans of the world
  • Search every page
  • Browse (developer view)
🤫husshhusshKirkland, WAPrivacyTerms

© 2026 Hushh Technologies Corporation - an independent company.