Videos

Each series is a study first and a video second: every number in an episode is in a run ledger you can open. Episodes appear here as they are published; a thumbnail loads the player only when you click it.

AI Agent Harnesses (2026)

publishing, 3 of 9 episodes on YouTube · data and code → · channel →

Eight episodes, a trailer and a bonus on the code around a language model: what it sees, what it can do, what it costs and whether it can be trusted. Every number is primary-sourced and every demo runs. Headline: 2 to 35 of 97sourcenext_series/T02_agent_security/series/FACTS.md in dhairyashilRG/agent-harnesses-2026checked2026-09-19kindmeasured.

Same Model, Better Agent: What an AI Harness Is | Ep 1

Every AI Agent Loop: Think, Act, Observe | Ep 2

Anatomy of an AI Agent Harness: The 7 Places Bugs Hide | Ep 3

What the Agent Keeps

cut, not yet published · data and code →

Five context policies inside the same window, four models, 1,199 graded trials. The policy moved every model from failing to perfect when the question came first, and every policy failed when it came last. Headline: 16/36 → 36/36sourcenext_series/T01_context_and_memory/SUMMARY.md in dhairyashilRG/agent-harnesses-2026checked2026-09-19kindmeasured.

Your Agent Will Be Attacked

cut, not yet published · data and code →

AgentDojo's public prompt injections against eight harness defences on the same model and tasks. The same attack landed 2 to 35 times out of 97 depending only on the model; out-of-band defences beat every in-band one. Headline: 2 to 35 of 97sourcenext_series/T02_agent_security/series/FACTS.md in dhairyashilRG/agent-harnesses-2026checked2026-09-19kindmeasured.

Evals That Don't Lie

cut, not yet published · data and code →

Three LLM judges grade 300 real agent runs against programmatic ground truth. Agreement looks fine; Cohen's κ says a third of it was chance. Headline: agreement 0.72–0.78, κ 0.45–0.56sourcenext_series/T03_evals/series/FACTS.md in dhairyashilRG/agent-harnesses-2026checked2026-09-19kindmeasured.

Eight Models, One Mac

cut, not yet published · data and code →

Eight open 2026 models at default 4-bit on one MacBook Pro, two harnesses, 1,512 graded attempts. A test everyone passes ranks no one; the hard set did. Headline: 1,512sourcenext_series/T04_local_models/SUMMARY.md in dhairyashilRG/agent-harnesses-2026checked2026-09-19kindmeasured.