Puppy One keeps your most sensitive context close. You still hold the keys.
It thinks locally whenever it can, and reaches the cloud only when you permit it. You own the device, the keys, the data, and the routing policy. This is not a laptop we resell - it is the reference engineering machine behind a private-AI system: provisioning, a curated local model pack, encrypted local RAG, hybrid-compute routing, and a transparent cost-and-energy accounting for every token it generates.
Yes - one unit, bought now, as the reference and Founder's Edition machine.
- An exceptional portable private-AI workstation - not the right default machine for every Puppy One customer.
- This first unit is a Hushh-owned development, demonstration, and benchmarking machine - not the start of a resale-inventory financing strategy. Apple's retail terms are for end-user purchase, and 0% installment credit is real consumer credit risk, not a wholesale facility.
- Commercial units route through an authorized Apple business procurement relationship, or the customer buys and owns the Mac directly - never bulk retail-card purchases intended for resale.
Where it is genuinely excellent, acceptable, or the wrong tool.
Excellent fit
- • Private RAG across financial documents, email, CRM exports, meeting notes, research, and contracts
- • Advisor and family-office meeting preparation
- • Financial-statement and tax-document extraction
- • Drafting client communications where source data should stay local
- • Compliance pre-review and policy comparison, with human approval
- • Private-repository coding agents
- • Local multimodal document and screenshot understanding
- • Speech transcription, embeddings, semantic search, and local knowledge graphs
- • Offline and travel use
- • Single-user and small-team interactive AI
- • LoRA / QLoRA adaptation of small and medium models
Acceptable, with limits
- • One heavy 70B-120B model session
- • Two simultaneous medium-model sessions
- • Moderate long-context analysis
- • Local image generation
- • Experimental fine-tuning of 30B-class models
- • Temporary internal API serving
The wrong machine
- • Training foundation models from scratch
- • High-concurrency multi-tenant SaaS
- • A 24/7 production service with a contractual uptime requirement
- • CUDA-only research and production stacks
- • Large-scale reinforcement learning
- • Large generative-video batches
- • Hundreds of parallel agents
- • Dense models above the practical 128GB memory envelope
- • Any workload requiring redundant nodes and automatic failover
A future Mac Studio or redundant desktop configuration is the right shape for an always-on Puppy Station. This machine is the portable, owner-grade Puppy.
BYO everything - compute, model, memory, infrastructure, intelligence, and above all, subscription.
Puppy One and Agent One are built to bring your own: your own compute (this machine, a cloud instance you already pay for, or both), your own model choice and the memory/context budget it needs, your own cloud infrastructure and identity stack, your own intelligence provider under your own keys - and above all, your own subscription. If you already pay for Apple One, Google One, ChatGPT Plus, Claude Pro, or a committed cloud credit, the router burns that quota down first. We are never the vendor standing between you and a service you already pay for - we are the orchestration and consent layer on top of it.
The four BYO pillars →Quota attainment - the honest upsell signal
When a customer's existing subscription quota keeps running out - the ChatGPT Plus message cap hit every week, the Claude Pro allowance gone by Wednesday, committed cloud credits exhausted early - that consistent pattern is the signal to offer owned compute, never a sales push based on guesswork. This is how the business model meets the customer where they actually are: start on what they already pay for, and let their own measured usage tell us when a Puppy purchase is the honest next step.
Private by default. Your own subscription next. Metered cloud last.
Financial information, family data, confidential documents, private repositories, regulated workflows.
No prompt or source data leaves the machine. Local models only.
Work that starts local and occasionally needs more headroom.
The local model handles retrieval, preparation, redaction, and ordinary reasoning. When local capability, context, concurrency, or latency is exceeded, the router burns down a subscription quota you already pay for before it ever reaches for a metered route - and only the minimum necessary context is transmitted, every route logged and visible to you.
Non-sensitive work where the frontier model is materially better.
An approved frontier cloud model - your own subscription or API keys first, a metered route only if nothing you already pay for has headroom - used for non-sensitive work or explicit customer-approved use, with estimated token cost and data boundary shown before execution.
Seven questions, in order, every single time.
Data sensitivity
Is cloud transmission permitted at all?
Model capability
Can a certified local model meet the task-quality threshold?
Memory & context
Will the model and its KV cache fit safely in the certified envelope?
Latency & concurrency
Is the local queue within the service target?
BYO subscription quota
Does an existing subscription you already pay for - a message allowance, a storage tier, a committed cloud credit - have headroom to burn down before anything metered?
Economics
If nothing you already own or pay for has headroom, which metered route has the lower expected cost for an accepted result?
Consent
Has the customer permitted this specific route?
Audit
Record the model, route, data boundary, estimated energy, and dollar cost - every time.
Cloud becomes the default for
- • Foundation-model training
- • CUDA-dependent software
- • High-concurrency workloads
- • Large public or synthetic batches
- • Very long contexts beyond the local certified envelope
- • Frontier-model quality requirements
- • Multi-tenant production services
- • Workloads requiring high availability or geographically distributed serving
The Puppy Compute Passport
Every certified model publishes a passport, not a marketing number: exact model and hash, license and permitted commercial use, quantization, runtime and version, macOS version, maximum certified context, memory consumption, prefill speed at 1K/8K/32K input tokens, decode speed at 128/512/2,048 output tokens, time-to-first-token (P50/P95), 1/2/4-session concurrency, idle/prefill/decode wall power, watt-hours per million output tokens, hardware-plus-energy cost per million tokens, four-hour thermal stability, a 72-hour service soak, and whether a given task ran local or used cloud burst. Every figure states plainly whether it is measured or estimated.
How to measure it honestly
The adapter's wattage is not inference power. A defensible number requires: battery already charged, fixed display brightness, the same macOS and runtime version, the same prompt/context/output length, separate idle/prefill/decode measurements, and at least three runs after thermal stabilization. The right units are tokens per joule, tokens per watt-hour, watt-hours per million output tokens, and million accepted tokens per kilowatt-hour - accepted tokens, not merely generated ones, because a token the user rejects is wasted compute.
The figures below are early community benchmarks on comparable M5 Max 128GB hardware - directional, not yet a Hushh-measured Compute Passport. They will be replaced by our own wall-power and runtime measurements before launch.
| Model, quantization | Community decode rate | Illustrative $ / million tokens |
|---|---|---|
| Llama 3.1 8B, 4-bit | 117 t/s | ~$2.10 hardware + electricity |
| Qwen3 MoE 30B, 4-bit | 176 t/s | ~$1.40 hardware + electricity |
| Gemma 4 26B-A4B, 4-bit | 151 t/s | ~$1.63 hardware + electricity |
| GPT-OSS 120B, 4-bit | ~114 t/s | ~$2.16 hardware + electricity |
| Qwen3-VL 32B dense, 4-bit | 27 t/s | ~$9.11 hardware + electricity |
Local ownership versus cloud rental - two different comparisons
Same architecture (local Apple silicon versus a hosted Apple-silicon instance) and same business outcome (local versus a data-center GPU running the same open-weight model) are not the same comparison. Modeled at an illustrative local cost near $0.89/productive hour (capital plus electricity, three-year life, eight productive hours a day), a comparable hosted Apple-silicon cloud instance runs roughly 7x more per allocated hour, and a data-center GPU rental runs roughly 4-5x more per rented hour at public list price - cited for comparison only, never as a claim that we operate or partner with any provider named. In practice the first cloud route we reach for is never a fresh, metered account - it's the quota you already burn down under BYO Subscription.
Local wins for repeated, private, interactive, low-concurrency work. Your own existing subscription wins next, for as long as it has headroom. Metered cloud is the last resort, reserved for real parallelism a single interactive user rarely uses enough of to capture its economics.
Ten gates, honestly tracked.
One is a product of Hushh Technologies Corporation (brand: 🤫 “hussh”), an independent company. One runs on third-party silicon, systems, and cloud; all company names are used solely to describe the platforms on which One software runs. Hushh Technologies is not affiliated with, endorsed by, sponsored by, or partnered with any company named.
Own the machine. Own the routing policy.
This page will be replaced, line by line, with our own measured Compute Passport as the reference unit is benchmarked - not the other way around.
One is a product of Hushh Technologies Corporation (brand: 🤫 “hussh”), an independent company. One runs on third-party silicon, systems, and cloud; all company names are used solely to describe the platforms on which One software runs. Hushh Technologies is not affiliated with, endorsed by, sponsored by, or partnered with any company named.