A useful assistant starts with the right fit. Compare the computer, choose a model, and leave room for everything else.
14computer options
30curated models
01setup that suits you
One everyday workhorse is enough to get started.
YOUR WORKBENCH
A setup, shaped around you.
▣
Pick the platform, then the capacity.
A researched shortlist of established and specialist brands, not a market ranking. Local price, stock and support vary.
Your roles. Only as many as you need.
Start with one model or build the full nine-role architecture. Matching selections share one serialized instance.
Format changes apply to eligible generative models. GPT-OSS stays MXFP4; CPU encoders keep their own format. Cache precision is not changed.
BEFORE YOUR FIRST DOWNLOAD
Know what you’re bringing home.
Sources checked 12 September 2026
01 / THE FILE
A model name isn’t a file format.
GGUF, MLX and Safetensors
GGUF is the route used by this planner for generative models, with llama.cpp through Vulkan, Metal, CUDA or CPU. Start with Q4_K_M for eligible models. Q5/Q8 cost more disk and RAM; quality improvements depend on the task. GPT-OSS uses its own MXFP4 representation.
MLX is an Apple Silicon alternative. Use an MLX-compatible converted model and a runtime that supports that exact architecture. GGUF and MLX artifacts are not interchangeable. Safetensors is a storage format used by Transformers and MLX; the model configuration and runtime still have to match.
Use the official model card, then inspect its linked quantizations. Check uploader, exact family/size, Instruct or IT variant, quantization and any vision projector. A source link here is not an automatic download.
02 / THE WORKFLOW
Get one thing working first.
Windows, Mac and Linux routes
LM Studio offers a convenient desktop path on supported Windows and Apple Silicon systems. Use Vulkan for the AMD baseline; Metal on Mac; CUDA for a supported NVIDIA GPU. ROCm support depends on the exact hardware, OS, driver and framework combination.
GB10 devices use Linux/Arm64. Follow the NVIDIA software guidance and an appropriate CUDA build. Do not download x86 Windows tools for an Arm Linux machine.
For a local RAG application, use a local endpoint, Python, a single-process Qdrant local store, and SQLite metadata/BM25. CPU encoders use their own Transformers/Sentence Transformers path. Bind to localhost and disable cloud fallback if the intent is fully local.
03 / THE DECISION
Fit is the beginning.
Try your own ten representative prompts. Compare answer quality, time to first response, memory pressure and mistakes. A larger model or a higher TOPS figure does not automatically mean a better everyday assistant.
Origins, licenses & limitations
Home includes all origins in this curated catalog. Company applies Windows-compatible hardware and excludes Chinese-developed model families and derivatives; it does not certify company approval. The Apache/MIT profile filters model licenses, not hardware origins. An uploader’s location does not change the original model’s developer or base lineage.
License labels summarize model cards. Check the exact weights, original terms and derivative requirements. The catalog includes established practical options; it is not exhaustive or a claim to rank the newest models.
Product names belong to their respective makers. Device cards link manufacturer specifications; announced products are marked. Prices, local availability, warranties, shipping SKUs and delivery dates must be checked before purchase. Bandwidth is a published specification, not a measured generation speed.
This private site is a planner. It does not install software, download model weights, run inference, or read your files. Settings stay in this browser.