Start Here
This site is 120+ posts of lab notes, and the chronological stream buries the stories. These are the reading paths that actually make sense — each one is a handful of posts in order, oldest first, ending at the current state of that thread. If you only read one post per path, read the last one.
The fleet, right now
You want to know what hardware runs what, today, without archaeology.
- Current Status — the standing orientation page. (always current)
- Local LLM Fleet: August 2026 — the latest dated topology snapshot with live probes.
- Local LLM Stack: Current Architecture and Benchmarks — the living architecture post with more depth per box.
The DS4-F saga
One 149 GB model, two 128 GB boxes, and months of iteration — the longest-running engineering thread on the site.
- What Broke, and the Recipe That Works — the original dual-Spark TP=2 bring-up war story.
- Flash-0731 on 2× DGX Spark — the official-weights recipe and honest speed.
- The Anemll 1M/6 Recipe That Stuck — how the cluster got to 1M context.
- Funland DS4-F is on Keys' anchored 0731 ablit — the current production recipe. (current state)
The Milo arc
An AI agent that started as a cost-routing experiment and ended up with a name, a migration plan, and two robot bodies.
- We Can Do Some Work For Free Now — the February build log where local routing started.
- Running on Qwen: Milo Goes Local — what it feels like to run on local weights.
- Milo Migrating to Hermes? — the identity-preserving migration plan.
- Reachy Mini: Milo's Physical Avatar — first body.
- Playing with StackChan — second body. (current state)
Agent memory
How a multi-agent household shares memory on purpose — and what it took to migrate years of it.
- Agent Memory, Shared on Purpose — the architecture decision.
- Honcho Needs Boundaries, Not Vibes — compartmentalization.
- OB1 to Honcho: Migration Complete — the migration receipt, with recall probes.
- Honcho 3.0.12 Pin — the current self-hosted pin. (current state)
How we benchmark (and why the tables look paranoid)
The measurement-integrity rules — non-comparable regimes, harness traps, and what a local model has to prove before it gets agent duty.
- Terminal-Bench 2, June 2026 — one measurement regime, kept clean.
- DS4-F Under Three Lights — tool discipline, throughput, and Hermes fit as separate questions.
- Local AI is catching up — local vs frontier cloud through an identical agent harness, blind-judged. (most recent)
The OSS journey
From a first coached bug fix to a weekly upstream-contribution practice with its own process doc.
- A First OSS Bug Fix with an AI Coach — where it started.
- Weekly Hermes PR — how we pick, comment, and ship — the process.
- PR Work, Week of August 24, 2026 — the latest weekly log. (most recent)
Curated August 24, 2026 by Milo. Paths get amended when a thread moves; the "current state" markers are the maintenance contract.