<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Clock Lobster Blog</title>
  <subtitle>Editorial, tutorials, benchmarks, and news on building with AI.</subtitle>
  <link href="https://clocklobster.com/feed.xml" rel="self"/>
  <link href="https://clocklobster.com/"/>
  <updated>Tue, 14 Jul 2026 00:00:00 +0000</updated>
  <id>https://clocklobster.com/</id>
  <author>
    <name>Victor Salmon</name>
    <email>hello@clocklobster.com</email>
  </author><entry>
    <title>Kalshi Bets on a Market for AI Computing Power</title>
    <link href="https://clocklobster.com/blog/news/2026-07-14-kalshi-compute-markets/"/>
    <updated>Tue, 14 Jul 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/news/2026-07-14-kalshi-compute-markets/</id>
    <content type="html">
&lt;!-- Hero --&gt;
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
    &lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
        &lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/news/&quot; class=&quot;accent&quot;&gt;News&lt;/a&gt;&lt;/p&gt;
        &lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Kalshi Bets on a Market for AI Computing Power&lt;/h1&gt;
        &lt;p style=&quot;color: var(--text-muted); font-size: 0.9375rem; margin-top: 1rem;&quot;&gt;Published July 14, 2026&lt;/p&gt;
    &lt;/div&gt;
&lt;/section&gt;

&lt;!-- Body --&gt;
&lt;section class=&quot;section section-flush&quot;&gt;
    &lt;div class=&quot;container container-narrow&quot;&gt;
        &lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

            &lt;p&gt;&lt;strong&gt;Series:&lt;/strong&gt; News — Article Review
            &lt;strong&gt;Published:&lt;/strong&gt; 2026-07-14
            &lt;strong&gt;Source:&lt;/strong&gt; &lt;a href=&quot;https://www.bloomberg.com/news/articles/2026-07-14/kalshi-ramps-up-effort-to-build-markets-for-ai-computing-power&quot;&gt;Bloomberg&lt;/a&gt;
            &lt;strong&gt;Author:&lt;/strong&gt; Victor Salmon&lt;/p&gt;
            &lt;hr /&gt;

            &lt;h2 id=&quot;summary&quot;&gt;The article&lt;/h2&gt;
            &lt;p&gt;Kalshi — the prediction-markets exchange — is pushing to turn AI compute into a tradeable commodity. Per &lt;a href=&quot;https://www.bloomberg.com/news/articles/2026-07-14/kalshi-ramps-up-effort-to-build-markets-for-ai-computing-power&quot;&gt;Bloomberg&#39;s Katherine Doherty (July 14, 2026)&lt;/a&gt;, Kalshi is offering a &lt;strong&gt;forward curve tracking compute&lt;/strong&gt; — &quot;compute&quot; being shorthand for the power, storage, memory, and GPU resources that feed AI.&lt;/p&gt;
            &lt;p&gt;The curve is built from Kalshi&#39;s own event contracts, spanning &lt;strong&gt;various GPU grades, locations, and tenors&lt;/strong&gt;, on both weekly and monthly bases, reaching up to &lt;strong&gt;a year into the future&lt;/strong&gt;. In plain terms: a market that plots where the rental price of AI hardware is heading.&lt;/p&gt;
            &lt;p&gt;Kalshi isn&#39;t alone. This is the third such move in 2026 — &lt;strong&gt;CME Group&lt;/strong&gt; announced a futures market for AI computing power back in May (partnering with Silicon Data), and &lt;strong&gt;ICE / NYSE&#39;s owner&lt;/strong&gt; unveiled plans for its own compute futures market the same month. The race to price AI infrastructure is clearly on.&lt;/p&gt;
            &lt;p style=&quot;margin-top: 1rem;&quot;&gt;&lt;a href=&quot;https://www.bloomberg.com/news/articles/2026-07-14/kalshi-ramps-up-effort-to-build-markets-for-ai-computing-power&quot;&gt;Read the source at Bloomberg →&lt;/a&gt;&lt;/p&gt;

            &lt;h2 id=&quot;commentary&quot;&gt;Our take&lt;/h2&gt;
            &lt;p&gt;A forward curve for compute is the kind of plumbing most people will never see — but it matters a lot if you pay an LLM bill. Here&#39;s why.&lt;/p&gt;
            &lt;p&gt;The headline price of &quot;AI compute&quot; has been a fuzzy, marketing-driven number. Vendors quote per-token rates; hyperscalers quote per-GPU-hour; brokers quote per-H100-month. None of them line up, and the gap between the cheapest and most expensive path to the same output can be enormous. Our own &lt;a href=&quot;https://clocklobster.com/blog/benchmarks/tokenizer-efficiency/&quot;&gt;tokenizer-efficiency benchmark&lt;/a&gt; found a &lt;strong&gt;74% spread&lt;/strong&gt; in how many tokens different models burn on the same words — meaning two models at the same sticker price can cost wildly different amounts per unit of real work.&lt;/p&gt;
            &lt;p&gt;A tradeable forward curve won&#39;t fix tokenizer inefficiency, but it does something complementary: it makes the &lt;em&gt;underlying&lt;/em&gt; — the hardware itself — legible and hedgeable. Three things follow:&lt;/p&gt;
            &lt;ul&gt;
                &lt;li&gt;&lt;strong&gt;Price discovery.&lt;/strong&gt; Today, if you want to know what an H100-hour will cost in March, you call a broker and get a quote that&#39;s only as good as your leverage. A liquid forward curve turns that into a public number everyone can see.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Hedging.&lt;/strong&gt; Teams that buy serious compute (training runs, inference fleets) can lock in future costs instead of gambling on spot prices. Predictability is worth money.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Benchmark reality.&lt;/strong&gt; A compute price index lets cost-per-task numbers — like the ones we publish in &lt;a href=&quot;https://clocklobster.com/blog/benchmarks/&quot;&gt;Benchmarks&lt;/a&gt; — be normalized against the actual market price of the iron, not just the provider&#39;s rate card.&lt;/li&gt;
            &lt;/ul&gt;
            &lt;p&gt;The open question is whether prediction-market contracts (Kalshi&#39;s model) will be liquid and trusted enough to become a real reference price, or whether the exchange-traded futures from CME and ICE will dominate. Kalshi&#39;s bet is that its event-contract approach, sliced by GPU grade and geography, captures granularity the bigger exchanges won&#39;t bother with at first.&lt;/p&gt;

            &lt;h2 id=&quot;takeaway&quot;&gt;Why it matters&lt;/h2&gt;
            &lt;p&gt;For most readers the takeaway is practical, not financial. Compute is quietly becoming the single biggest variable cost of doing anything with AI — and until now it&#39;s been opaque. Three exchanges racing to publish forward prices means that opacity is starting to lift. If you&#39;re budgeting an AI project, watching a compute curve (the way you&#39;d watch a currency or commodity) is about to become a reasonable thing to do.&lt;/p&gt;
            &lt;p&gt;And on our end: a public compute price makes the per-word, per-task cost analysis we do in &lt;a href=&quot;https://clocklobster.com/blog/benchmarks/&quot;&gt;Benchmarks&lt;/a&gt; more grounded — we can quote results against a shared market reference instead of a provider&#39;s list price. We&#39;ll be watching which curve becomes the standard.&lt;/p&gt;

            &lt;hr /&gt;
            &lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://clocklobster.com/blog/news/&quot;&gt;Browse all News&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

        &lt;/div&gt;
    &lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Five Signals a Workflow Is Broken — and Ripe for AI</title>
    <link href="https://clocklobster.com/blog/news/2026-07-14-five-signals-broken-workflows/"/>
    <updated>Tue, 14 Jul 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/news/2026-07-14-five-signals-broken-workflows/</id>
    <content type="html">
&lt;!-- Hero --&gt;
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
    &lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
        &lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/news/&quot; class=&quot;accent&quot;&gt;News&lt;/a&gt;&lt;/p&gt;
        &lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Five Signals a Workflow Is Broken — and Ripe for AI&lt;/h1&gt;
        &lt;p style=&quot;color: var(--text-muted); font-size: 0.9375rem; margin-top: 1rem;&quot;&gt;Published July 14, 2026&lt;/p&gt;
    &lt;/div&gt;
&lt;/section&gt;

&lt;!-- Body --&gt;
&lt;section class=&quot;section section-flush&quot;&gt;
    &lt;div class=&quot;container container-narrow&quot;&gt;
        &lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

            &lt;p&gt;&lt;strong&gt;Series:&lt;/strong&gt; News — Article Review
            &lt;strong&gt;Published:&lt;/strong&gt; 2026-07-14
            &lt;strong&gt;Source:&lt;/strong&gt; &lt;a href=&quot;https://x.com/coreyganim/status/2077074141879152799&quot;&gt;Corey Ganim on X&lt;/a&gt;
            &lt;strong&gt;Author:&lt;/strong&gt; Victor Salmon&lt;/p&gt;
            &lt;hr /&gt;

            &lt;h2 id=&quot;summary&quot;&gt;The post&lt;/h2&gt;
            &lt;p&gt;In a &lt;a href=&quot;https://x.com/coreyganim/status/2077074141879152799&quot;&gt;thread on X (July 14, 2026)&lt;/a&gt;, Corey Ganim offers a bluntly practical method for finding broken workflows in a small business — the kind worth automating. His advice: ask the owner &lt;em&gt;&quot;Can you show me how this happens today?&quot;&lt;/em&gt;, then watch for five telltale signals:&lt;/p&gt;
            &lt;ol&gt;
                &lt;li&gt;&lt;strong&gt;Tabs.&lt;/strong&gt; How many tools do they open? Six tabs to finish one task means friction.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Copy/paste.&lt;/strong&gt; What gets moved by hand from one place to another? That&#39;s automation opportunity.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Waiting.&lt;/strong&gt; Where does work stall waiting on a reply, an approval, a missing file? That&#39;s a bottleneck.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Rework.&lt;/strong&gt; Where do people fix the same mistakes over and over? That&#39;s a process problem.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Handoffs.&lt;/strong&gt; Where does a task pass from one person to another? That&#39;s where things get lost.&lt;/li&gt;
            &lt;/ol&gt;
            &lt;p&gt;He lists concrete examples: a bookkeeper&#39;s month-end close, logistics quote creation, recruiting candidate screening, agency client onboarding, school enrollment inquiries. The thesis, in his words: &lt;em&gt;&quot;simply watching how the work gets done today will show you exactly where AI belongs.&quot;&lt;/em&gt;&lt;/p&gt;
            &lt;p style=&quot;margin-top: 1rem;&quot;&gt;&lt;a href=&quot;https://x.com/coreyganim/status/2077074141879152799&quot;&gt;Read the source thread on X →&lt;/a&gt;&lt;/p&gt;

            &lt;h2 id=&quot;commentary&quot;&gt;Our take&lt;/h2&gt;
            &lt;p&gt;This is the part of automation that almost nobody writes about — and it&#39;s the part that determines whether an AI project succeeds. Most failure isn&#39;t in building the agent; it&#39;s in picking the wrong task to automate.&lt;/p&gt;
            &lt;p&gt;Ganim&#39;s five signals are a field-ready discovery checklist, and they line up almost perfectly with the layers we teach and build. Mapping them to what comes next:&lt;/p&gt;
            &lt;ul&gt;
                &lt;li&gt;&lt;strong&gt;Copy/paste and Tabs&lt;/strong&gt; are pure data-movement problems. Once you&#39;ve spotted them, the fix is usually a single integration or a desktop agent moving data between tools — exactly what our &lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot;&gt;Basic Agents&lt;/a&gt; and &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot;&gt;Agents Working for You&lt;/a&gt; tracks cover.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Waiting&lt;/strong&gt; is an orchestration problem — work idling for an approval or a missing input. This is where autonomous agent loops (that poll, remind, and fetch) earn their keep.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Rework&lt;/strong&gt; is a quality problem — usually a categorization or extraction step that an LLM can do once, correctly, instead of a human fixing it repeatedly.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Handoffs&lt;/strong&gt; are a coordination problem — and the most valuable to fix, because that&#39;s where errors compound and context gets lost between people.&lt;/li&gt;
            &lt;/ul&gt;
            &lt;p&gt;One nuance worth adding: signals 1 and 2 (Tabs, Copy/paste) are the cheapest wins and the best place for a small business to start. Signals 4 and 5 (Rework, Handoffs) are higher-value but harder — they touch how people work, not just what tools they use. Start with the copy/paste, build trust, then tackle the handoffs.&lt;/p&gt;

            &lt;h2 id=&quot;takeaway&quot;&gt;Why it matters&lt;/h2&gt;
            &lt;p&gt;If you&#39;re a small-business owner (or you help one), the highest-leverage hour you can spend this week is the one Ganim describes: sit down, pick one recurring task, and just &lt;em&gt;watch&lt;/em&gt; how it gets done. The broken parts will announce themselves. Once you know where the friction is, the question of &lt;em&gt;whether&lt;/em&gt; AI helps — and which kind — answers itself, and our &lt;a href=&quot;https://clocklobster.com/blog/tutorials/&quot;&gt;Tutorials&lt;/a&gt; pick up right from there.&lt;/p&gt;
            &lt;p&gt;And if you&#39;d rather have someone do that discovery with you, that&#39;s literally the first thing we do — &lt;a href=&quot;https://clocklobster.com/book-consultation.html&quot;&gt;book a consultation&lt;/a&gt; and we&#39;ll map your five signals together.&lt;/p&gt;

            &lt;hr /&gt;
            &lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://clocklobster.com/blog/news/&quot;&gt;Browse all News&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

        &lt;/div&gt;
    &lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>The 2026 #AISkillGap Challenge</title>
    <link href="https://clocklobster.com/blog/editorial/2026-ai-challenge/"/>
    <updated>Fri, 03 Jul 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/editorial/2026-ai-challenge/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 4rem;&quot;&gt;
    &lt;div class=&quot;container text-center&quot;&gt;
        &lt;span class=&quot;section-label&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/editorial/&quot; style=&quot;color: var(--cta);&quot;&gt;Editorial&lt;/a&gt;&lt;/span&gt;
        &lt;h1 class=&quot;hero-title&quot; style=&quot;max-width: 800px; margin-left: auto; margin-right: auto;&quot;&gt;The 2026 #AISkillGap Challenge&lt;/h1&gt;
        &lt;p class=&quot;hero-subtitle mx-auto&quot; style=&quot;max-width: 640px;&quot;&gt;A year-long personal challenge to build, document, and share the AI agent skills that actually matter. You&#39;re welcome to follow along.&lt;/p&gt;
        &lt;p style=&quot;color: var(--text-muted); font-size: 0.9375rem;&quot;&gt;Published July 3, 2026&lt;/p&gt;
    &lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section&quot; style=&quot;padding-top: 0;&quot;&gt;
    &lt;div class=&quot;container&quot;&gt;
        &lt;div class=&quot;glass-card&quot; style=&quot;max-width: 820px; margin: 0 auto; padding: 3rem;&quot;&gt;

            &lt;h2 style=&quot;margin-bottom: 1.5rem;&quot;&gt;Why I&#39;m Doing This&lt;/h2&gt;
            &lt;p class=&quot;mb-4&quot;&gt;I build AI agents for a living. Containerized digital employees that run around the clock, manage workflows, draft correspondence, and make decisions within guardrails. Open Claw, Nemo Claw — you&#39;ve probably heard me mention them.&lt;/p&gt;
            &lt;p class=&quot;mb-4&quot;&gt;But here&#39;s the thing I&#39;ve noticed: most of the useful AI skills aren&#39;t locked inside expensive platforms. They&#39;re patterns. Small, composable capabilities that any reasonably capable agent can learn if someone writes them down. And right now, nobody is writing them down.&lt;/p&gt;
            &lt;p class=&quot;mb-4&quot;&gt;So I&#39;m going to.&lt;/p&gt;

            &lt;h2 style=&quot;margin-bottom: 1.5rem;&quot;&gt;The Rules&lt;/h2&gt;
            &lt;p class=&quot;mb-4&quot;&gt;This is a challenge I&#39;ve issued to myself. No leaderboard, no prizes, no submissions. The format is simple:&lt;/p&gt;
            &lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
                &lt;li&gt;&lt;strong&gt;Build something real&lt;/strong&gt; — every post features something I actually created: a container, a script, a skill file, an integration. No theory without practice.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Share the skill&lt;/strong&gt; — every creation comes with a documented AI agent skill: what it does, how it works, when to use it, what to watch out for.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Keep it accessible&lt;/strong&gt; — the skills should be useful to anyone running autonomous agents, not just people with my specific stack.&lt;/li&gt;
                &lt;li&gt;&lt;strong&gt;Post throughout the year&lt;/strong&gt; — periodic updates as the challenge evolves, with all the wins, failures, and lessons learned along the way.&lt;/li&gt;
            &lt;/ul&gt;

            &lt;h2 style=&quot;margin-bottom: 1.5rem;&quot;&gt;What You&#39;ll Find Here&lt;/h2&gt;
            &lt;p class=&quot;mb-4&quot;&gt;Over the course of 2026, this blog will fill with:&lt;/p&gt;
            &lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
                &lt;li&gt;AI agent skill files ready to drop into any compatible runtime&lt;/li&gt;
                &lt;li&gt;Deep dives into specific agent capabilities: research, reasoning, tool use, memory, delegation&lt;/li&gt;
                &lt;li&gt;Real infrastructure: how I deploy, scale, and monitor production agents&lt;/li&gt;
                &lt;li&gt;Failures — tools that didn&#39;t work, approaches that broke, lessons paid for in compute credits&lt;/li&gt;
                &lt;li&gt;Container recipes, prompt patterns, security boundaries, and integration tricks&lt;/li&gt;
            &lt;/ul&gt;

            &lt;h2 style=&quot;margin-bottom: 1.5rem;&quot;&gt;Post #1: Coming Soon&lt;/h2&gt;
            &lt;p class=&quot;mb-4&quot;&gt;The first post is in the works. I&#39;m documenting a skill I built to solve a genuinely annoying problem — one that cost me hours every week until I automated it. If you run agents, you&#39;ll want this one.&lt;/p&gt;
            &lt;p style=&quot;color: var(--text-muted);&quot;&gt;Subscribe to the &lt;a href=&quot;https://clocklobster.com/contact.html&quot; style=&quot;color: var(--cta); text-decoration: underline;&quot;&gt;newsletter&lt;/a&gt; or check back here for updates.&lt;/p&gt;

            &lt;hr style=&quot;border: none; border-top: 1px solid var(--border); margin: 3rem 0;&quot; /&gt;

            &lt;div style=&quot;text-align: center;&quot;&gt;
                &lt;p style=&quot;color: var(--text-muted); font-size: 0.9375rem;&quot;&gt;The 2026 #AISkillGap Challenge&lt;/p&gt;
                &lt;p style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;— Victor&lt;/p&gt;
            &lt;/div&gt;

        &lt;/div&gt;
    &lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 8: What&amp;#x27;s Next: From Chat to Tools</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 8&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;What&amp;#x27;s Next: From Chat to Tools&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Beyond the chat window — web search, &lt;span class=&quot;glossary-term&quot; data-term=&quot;artifacts&quot;&gt;artifacts&lt;/span&gt;, &lt;span class=&quot;glossary-term&quot; data-term=&quot;code-execution&quot;&gt;code execution&lt;/span&gt;, vision, and voice. The models you&#39;ve been using are capable of more than type-and-respond. Here&#39;s your on-ramp.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/#hook&quot; style=&quot;color: var(--cta);&quot;&gt;You&#39;ve Done It&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/#beyond-chat&quot; style=&quot;color: var(--cta);&quot;&gt;Beyond Chat&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/#pro-and-research-modes&quot; style=&quot;color: var(--cta);&quot;&gt;Pro Modes and Research Modes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/#try-a-tool-feature&quot; style=&quot;color: var(--cta);&quot;&gt;Try a Tool Feature&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/#ai-toolkit-checklist&quot; style=&quot;color: var(--cta);&quot;&gt;Personal AI Toolkit Checklist&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/#where-to-go-next&quot; style=&quot;color: var(--cta);&quot;&gt;Where to Go Next&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;See what your chatbot can do beyond typing — search, vision, code, artifacts, and research modes&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open your preferred chatbot: &lt;strong&gt;chatgpt.com&lt;/strong&gt;, &lt;strong&gt;claude.ai&lt;/strong&gt;, or &lt;strong&gt;gemini.google.com&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Pick one tool feature you haven&#39;t tried yet — web search, artifacts, code execution, or vision&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a second tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;hook&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;You&#39;ve Done It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You can now ask an LLM to write stuff, control its tone, work with documents, choose the right chatbot for the task, recognize when to verify its output, and steer it through real research. That&#39;s the full skill stack for effective chat — and if you&#39;ve followed along through lessons 1-7, you&#39;re already more proficient than most casual users.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;But the chat window is only the simplest interface. The same models power features that go beyond type-and-respond — features that are built into the chatbots you already have, often for free. This lesson connects everything you&#39;ve learned to what comes next: tools that turn chat into action.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;beyond-chat&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Beyond Chat — What Your Chatbot Can Already Do&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every major chatbot ships with capabilities that extend beyond the text box. Here&#39;s what&#39;s available and which models offer it:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Feature&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Available in&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;&lt;span class=&quot;glossary-term&quot; data-term=&quot;web-search&quot;&gt;Web search&lt;/span&gt;&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;ChatGPT, Gemini&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Pulls current information from the web instead of relying on training data that may be months or years old. Toggle it on or off per query.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Vision&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;ChatGPT, Claude, Gemini&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Upload a photo, screenshot, or scanned document and ask about its contents. Reads text in images and describes visual elements.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;File uploads&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;ChatGPT, Claude, Gemini&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Analyze PDFs, spreadsheets, code files, presentations — covered in Lesson 4.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Voice mode&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;ChatGPT, Gemini&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Speak your prompt instead of typing. Useful when you&#39;re on the go, cooking, or thinking out loud.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Code execution&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;ChatGPT, Claude&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;The model writes and runs Python code to calculate, visualize, or transform data — no setup, no environment needed.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Artifacts&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Claude&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Live preview pane where Claude renders web pages, diagrams, or interactive content in real time.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Canvas&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;ChatGPT&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;A shared editing workspace where you and the model can edit documents and code together, with inline suggestions and revisions.&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Web search&lt;/strong&gt; is the most immediately useful for research. &lt;strong&gt;Code execution&lt;/strong&gt; and &lt;strong&gt;Artifacts&lt;/strong&gt; expand what you can ask the model to do — from &quot;write an email&quot; to &quot;calculate this, visualize that, and show me both.&quot; The rest are quality-of-life upgrades that save you time on specific types of tasks. None require a paid plan.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;pro-and-research-modes&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Pro Modes and Research Modes&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The free tiers of these chatbots are capable, but each model offers a paid upgrade that unlocks significantly more powerful research capabilities. Here&#39;s what changes when you go pro, and which one to pick if research is your primary use case.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Model&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Pro / Research Mode&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What it adds&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Gemini&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Gemini Advanced + &lt;span class=&quot;glossary-term&quot; data-term=&quot;deep-research&quot;&gt;Deep Research&lt;/span&gt;&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Deep Research mode creates a multi-step research plan — it searches the web, reads multiple sources, and produces a structured report with citations. Best for &quot;I need to understand a topic from multiple angles&quot; research. The free tier has search, but Deep Research does the synthesis for you.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;ChatGPT&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;ChatGPT Plus&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Longer context, priority access, DALL-E image generation, data analysis (upload CSVs and ask questions). The model is the same GPT-4 class, but with fewer rate limits and a larger context window.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Claude&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Claude Pro&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Massive context window (200K tokens — about a full novel). Better for uploading entire books, codebases, or long documents and analyzing them in one pass. Claude&#39;s Pro tier is the best pick when your research material is a single very long document.&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Which should you pay for?&lt;/strong&gt; If your research is broad — exploring a new topic from multiple sources — Gemini&#39;s Deep Research mode is the most capable tool for multi-source synthesis. If your research is deep — analyzing a single 200-page document or contract — Claude&#39;s large context window gives you the most complete view in one chat. If you just want fewer rate limits and more consistent access, ChatGPT Plus is the safest default. None of these are essential for the workflows in this course, but if you find yourself doing weekly research, one of them will save you real time.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-a-tool-feature&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough — Try a Free Tool Feature&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Pick one of these options and try it right now. Each takes under a minute and shows you a capability the chat box alone doesn&#39;t have.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Option A — Web search (ChatGPT or Gemini)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Enable web search (look for the search toggle in the interface) and ask: &lt;em&gt;&quot;What were the top news stories this morning?&quot;&lt;/em&gt; Then turn off search and ask the same question. Compare the two answers — that&#39;s the difference between current information and training-data information. The first pulls live data; the second draws on whatever the model learned during training, which could be months or years old. This difference matters more than any other single feature for research tasks.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Option B — Artifacts (Claude)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Ask Claude: &lt;em&gt;&quot;Write me a simple to-do list web page as an artifact.&quot;&lt;/em&gt; Watch it build a working HTML page in real time in the preview pane. You can interact with the output, ask for changes, and see them applied instantly. This is the model moving from generating text to generating interactive content — a different category of output entirely.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Option C — Code execution (ChatGPT)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Ask: &lt;em&gt;&quot;Calculate the compound interest on $10,000 at 5% over 10 years and show me the year-by-year breakdown.&quot;&lt;/em&gt; ChatGPT writes and runs Python code to produce the table — no calculator needed, no spreadsheet required. The output includes the actual code it wrote, so you can inspect and reuse it. This is useful whenever you need a calculation, a chart, or a data transformation and don&#39;t want to open a separate tool.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Option D — Vision&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Take a photo of a whiteboard, a menu, or a printed document and upload it. Ask: &lt;em&gt;&quot;Summarize what&#39;s on this board&quot;&lt;/em&gt; or &lt;em&gt;&quot;What are the prices for the top 3 items?&quot;&lt;/em&gt; The model reads the text in the image and responds. Handwriting is less reliable than typed text, but the feature is surprisingly capable with clean photos of typed or printed material.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;Each of these is a small step beyond plain chat. Try one today; try another tomorrow. You don&#39;t need to learn them all at once. The goal is to know they exist so that when you encounter a problem one of them solves, you can reach for it instead of reaching for a different tool.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;ai-toolkit-checklist&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — Personal AI Toolkit Checklist&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a checklist of the features covered in this course. Mark which you&#39;ve tried and which you&#39;ll explore next:&lt;/p&gt;
&lt;pre style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 2rem; border: 1px solid var(--border); color: var(--text-primary); font-family: &#39;Inter&#39;, monospace; font-size: 0.9375rem; white-space: pre-wrap;&quot;&gt;[ ] Prompt template (Lesson 1)
[ ] Voice card for tone (Lesson 2)
[ ] Decision matrix — which chatbot for which task (Lesson 3)
[ ] Document upload and analysis (Lesson 4)
[ ] Trust-decisions checklist (Lesson 5)
[ ] Email draft workflow template (Lesson 6)
[ ] Article summary workflow template (Lesson 6)
[ ] Research brief template (Lesson 7)
[ ] Web search in ChatGPT or Gemini
[ ] Vision — upload a photo and ask about it
[ ] Voice mode — try speaking instead of typing
[ ] Artifacts — ask Claude to build something interactive
[ ] Code execution — ask ChatGPT to calculate or visualize data&lt;/pre&gt;
&lt;p class=&quot;mb-4&quot;&gt;You don&#39;t need to master everything at once. Pick the two or three tools that match the tasks you do most often, and make them a habit. The skill that compounds most is the one you use consistently.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;where-to-go-next&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Where to Go Next&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You&#39;ve completed the Basic Chatbots track. You can now use LLMs effectively in a browser — prompt, refine, research, verify, repeat. That&#39;s a real skill, and it&#39;s enough for most of what people need from AI day to day.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;If you want to go further, here&#39;s what&#39;s next:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Keep practicing:&lt;/strong&gt; pick one workflow from Lesson 6 and use it weekly until it feels automatic. The best next step is repetition, not more theory.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Next track:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot; style=&quot;color: var(--cta);&quot;&gt;Basic Agents&lt;/a&gt; — desktop and terminal agents that do real work. Learn the options, how to set one up, what they cost, and when to use them instead of chatting in a browser.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Customize your toolkit:&lt;/strong&gt; if you use a specific tool for work — Word, Excel, Gmail, Slack — check whether it has built-in AI features. Gemini (Google) and various third-party plugins bring these capabilities into the apps you already use. You don&#39;t always need to open a separate chatbot. &lt;strong&gt;A caution on value:&lt;/strong&gt; some LLMs are far more capable than others, and price is not a reliable guide. Microsoft Copilot, for example, can be a very expensive product that simply does not deliver well in some enterprise contexts (as of 2026.07.04). Test before you commit — a free chatbot with a well-written prompt often outperforms a premium add-on that sounded good in a sales demo.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; How to Ask an LLM for Research&lt;/a&gt;&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 7: How to Ask an LLM for Research</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 7&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;How to Ask an LLM for Research&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;A single question gets a single answer. Research is iterative — you ask, follow up, challenge, and converge. Here&#39;s a framework to steer toward answers you can trust.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/#a-single-question-gets-a-single-answer-research-needs-iteration&quot; style=&quot;color: var(--cta);&quot;&gt;A Single Question Gets a Single Answer — Research Needs Iteration&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/#the-scope-framework&quot; style=&quot;color: var(--cta);&quot;&gt;The SCOPE Framework&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/#walkthrough-exploratory&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Exploratory Research&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/#walkthrough-decision&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Decision Research&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/#research-brief&quot; style=&quot;color: var(--cta);&quot;&gt;Research Brief Template&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Develop a research method that pushes past surface-level answers to what you didn&#39;t know to ask&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 20 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open two browser tabs: &lt;strong&gt;chatgpt.com&lt;/strong&gt; and &lt;strong&gt;claude.ai&lt;/strong&gt; (you&#39;ll compare responses)&lt;/li&gt;
&lt;li&gt;Have a topic in mind that you&#39;ve been meaning to learn about&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a third tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;a-single-question-gets-a-single-answer-research-needs-iteration&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;A Single Question Gets a Single Answer — Research Needs Iteration&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You ask the model a research question — &quot;Tell me about X&quot; — and it gives you a confident, well-structured answer. It sounds right. But later, when you dig deeper on your own, you realize the answer was shallow. It glossed over a key debate, gave a generic example instead of a specific one, or missed an entire perspective. The model didn&#39;t give you bad information — it just didn&#39;t know what you actually needed.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;A single question gets a single answer. Research is different. It&#39;s iterative — you ask something, follow up on what you learn, challenge assumptions, cross-reference across sources, and converge on something you can actually use. The SCOPE framework gives you a structure for doing that with an LLM.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;the-scope-framework&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The SCOPE Framework — Steering Research, Not Just Asking&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;span class=&quot;glossary-term&quot; data-term=&quot;scope&quot;&gt;SCOPE&lt;/span&gt; turns a single question into a research conversation. Each letter is a type of follow-up that steers the model toward a richer answer.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header); width: 80px;&quot;&gt;Letter&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Means&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What to ask&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;S&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Scope&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&lt;strong style=&quot;color: var(--cta);&quot;&gt;Scope&lt;/strong&gt; your question. &quot;Tell me about X&quot; is too broad. &quot;What are the main arguments for and against Y in the context of Z?&quot; is research-ready.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;C&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Challenge&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&lt;strong style=&quot;color: var(--cta);&quot;&gt;Challenge&lt;/strong&gt; the answer. &quot;What&#39;s commonly misunderstood about this?&quot; &quot;What does this leave out?&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;O&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Opposing&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&lt;strong style=&quot;color: var(--cta);&quot;&gt;Opposing&lt;/strong&gt; views. &quot;What do critics of this approach say?&quot; &quot;What are the strongest counterarguments?&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;P&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Probe&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&lt;strong style=&quot;color: var(--cta);&quot;&gt;Probe&lt;/strong&gt; deeper. &quot;Give me a specific example.&quot; &quot;Who are the key researchers?&quot; &quot;What evidence supports that claim?&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;E&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Evaluate&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&lt;strong style=&quot;color: var(--cta);&quot;&gt;Evaluate&lt;/strong&gt; across sources. Run the same question in a second model and compare. Ask for citations. Calibrate confidence against what you already know.&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;S&lt;/strong&gt;cope + &lt;strong&gt;C&lt;/strong&gt;hallenge + &lt;strong&gt;O&lt;/strong&gt;pposing + &lt;strong&gt;P&lt;/strong&gt;robe + &lt;strong&gt;E&lt;/strong&gt;valuate = SCOPE. Don&#39;t wander — SCOPE your research.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;walkthrough-exploratory&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough 1 — Exploratory Research&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s say you want to learn about a topic you know almost nothing about — say, carbon capture technology. Follow these steps in your chatbot.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 1: Scope the question&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Paste&lt;/strong&gt; this into your chatbot:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I want to learn about carbon capture technology from scratch. I know almost nothing about it. Give me a quick overview, the 3 most important concepts I need to understand, and recommend 2 resources I should start with.&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 2: Challenge&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;After reading the overview, &lt;strong&gt;ask&lt;/strong&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What are the most common misconceptions about carbon capture? What do people tend to get wrong when they first learn about this?&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 3: Opposing views&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Ask&lt;/strong&gt; for the debate:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What do critics of carbon capture say? What are the strongest arguments against investing in it as a climate solution?&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 4: Probe deeper&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Pick one thing the model mentioned that surprised you and &lt;strong&gt;probe&lt;/strong&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;You mentioned [specific point]. Can you give me a concrete example of that in action? Who are the key companies or researchers working on it?&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 5: Evaluate&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now &lt;strong&gt;open a second chatbot&lt;/strong&gt; and paste the same first prompt. Compare the two overviews side by side:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Did both models identify the same 3 most important concepts?&lt;/li&gt;
&lt;li&gt;Did one mention something the other omitted entirely?&lt;/li&gt;
&lt;li&gt;Did the recommended resources overlap?&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;The gaps between the two answers tell you where to probe further. If both say the same thing, you can be more confident. If they disagree, that&#39;s where the real learning happens.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;walkthrough-decision&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough 2 — Decision-Oriented Research&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Research isn&#39;t always about learning from scratch. Sometimes you need to make a choice between two options.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 1: Frame the decision&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Paste&lt;/strong&gt; something like this:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I need to decide between using ChatGPT and Claude for drafting client proposals. Here&#39;s what I value: clear structure, warm tone, and the ability to follow complex formatting rules. Give me a comparison of how each handles these three things, and a recommendation with your reasoning.&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 2: Challenge the framing&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Ask&lt;/strong&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What factors am I not considering? Is there anything about this decision that most people overlook?&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 3: Probe with specifics&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Ask&lt;/strong&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Can you give me a concrete example of a proposal written by each model so I can compare the actual output?&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 4: Evaluate by cross-checking&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Take the recommendation and &lt;strong&gt;run the same prompt in the other model&lt;/strong&gt;. Ask it the same question and see how the recommendation differs. If both models recommend the same tool, you have more confidence. If they recommend different tools, probe the reasoning — which one considered something the other missed?&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;research-brief&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Research Brief Template&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Save this structured prompt as your starting point for any research task:&lt;/p&gt;
&lt;pre style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 1.5rem; border: 1px solid var(--border); color: var(--text-primary); font-family: &#39;Inter&#39;, monospace; font-size: 0.9375rem; white-space: pre-wrap;&quot;&gt;--- RESEARCH BRIEF ---
Topic: [what I want to learn about]
Scope: [what specifically I want to know]
What I already know: [optional, to avoid redundancy]

First pass:
- Quick overview
- 3 most important concepts
- 2 recommended starting points

Second pass — Challenge:
- What&#39;s commonly misunderstood about this?
- What perspective does my question leave out?

Third pass — Opposing views:
- What do critics or skeptics say?
- What are the strongest counterarguments?

Final pass — Probe:
- Give me 1 specific example
- Who are the key researchers or sources?
- What evidence supports the main claims?&lt;/pre&gt;
&lt;p class=&quot;mb-4&quot;&gt;Paste this at the start of a new chat, fill in your topic, and work through it section by section. The model will produce each pass as you ask for it, building on what it gave you before.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Pick a topic you&#39;ve been meaning to learn about — something that&#39;s been in your bookmarks, your reading list, or the back of your mind. Run the full SCOPE process with the research brief template. Save your brief and the outputs. Then run the same first prompt in a second model and compare. The differences between the two answers are the most valuable part — they show you where the models are guessing, where the topic is genuinely contested, and where you need to go read an actual human expert.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Real Workflows: Draft and Summarize&lt;/a&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-8/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;What&#39;s Next: From Chat to Tools &amp;rarr;&lt;/a&gt;&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 6: Real Workflows: Draft and Summarize</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 6&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Real Workflows: Draft and Summarize&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;You&#39;ve learned the pieces — &lt;span class=&quot;glossary-term&quot; data-term=&quot;faint&quot;&gt;FAINT&lt;/span&gt;, tone control, document work, model selection, trust calibration. This lesson puts them together. Two complete walkthroughs you can follow in your browser right now.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/#building-knowledge-isnt-having-a-workflow&quot; style=&quot;color: var(--cta);&quot;&gt;Building Knowledge Isn&#39;t Having a Workflow&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/#workflow-1-email-draft&quot; style=&quot;color: var(--cta);&quot;&gt;Workflow 1: Email Draft&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/#workflow-2-summarize&quot; style=&quot;color: var(--cta);&quot;&gt;Workflow 2: Summarize&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/#workflow-templates&quot; style=&quot;color: var(--cta);&quot;&gt;Reusable Templates&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Master two repeatable workflows — drafting from bullet points and summarizing long articles&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open your preferred chatbot: &lt;strong&gt;chatgpt.com&lt;/strong&gt;, &lt;strong&gt;claude.ai&lt;/strong&gt;, or &lt;strong&gt;gemini.google.com&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Have a draft email you need to send (for Workflow 1) and a long article or report you want to digest (for Workflow 2)&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a second tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;building-knowledge-isnt-having-a-workflow&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Building Knowledge Isn&#39;t Having a Workflow&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You&#39;ve learned the pieces — FAINT (Frame, Aim, Include, Nix, Tone), tone control, document work, model selection, trust calibration. But knowing the pieces separately isn&#39;t the same as doing the work. The most common reason people stop using AI after trying it isn&#39;t that the tools are bad. It&#39;s that they never developed a repeatable workflow — a consistent process they could run without reinventing the approach every time.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This lesson gives you two workflows that cover 80% of what people use LLMs for: generating content and digesting information. Follow along in your browser. These aren&#39;t theoretical demonstrations — they&#39;re prompts you can paste and adapt starting today.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;workflow-1-email-draft&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Workflow 1 — Draft an Email from Bullet Points&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the workflow you&#39;ll use most often. Instead of staring at a blank screen, you dump your raw thoughts and let the model structure them.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 1: Open your chatbot and paste context&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Paste&lt;/strong&gt; something like this, filled in with your actual situation:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I need to write a follow-up email to a client who hasn&#39;t responded in 2 weeks. Here are the key points I want to include:
- We met on June 15 to discuss the proposal
- I sent the formal proposal on June 18
- I&#39;m checking in, not pressuring
- I&#39;m available to answer questions or adjust scope
- I&#39;d like to suggest a specific time to talk next week&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Notice what&#39;s happening: you&#39;re using the &lt;strong&gt;FAINT&lt;/strong&gt; framework from Lesson 1 without thinking about it. Context (who you are, who the client is). Aim (a follow-up email). Include (the 5 bullet points). You haven&#39;t specified tone or constraints yet — that comes next.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 2: Add tone and structure constraints&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Add&lt;/strong&gt; this to the same chat:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Keep it to 4 sentences. Sound friendly but not pushy. Use contractions. End with a specific question that makes it easy for them to reply.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the &lt;strong&gt;Nix and Tone&lt;/strong&gt; from FAINT — you&#39;re telling the model what to avoid and how to sound. The draft should now be shorter, warmer, and structured the way you want.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 3: Read and adjust&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Read the draft. Does it sound like you? If not, make one small fix — replace a phrase, shorten a sentence, add a specific detail only you would know. &lt;strong&gt;Type&lt;/strong&gt; something like:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Change &quot;I wanted to check in&quot; to &quot;Just checking in.&quot; And add a sentence saying I&#39;m flexible on timing.&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 4: Ask it to cite sources&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;If your email references facts, data, or anything the reader might want to verify, add this:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;You mentioned [specific claim]. Where does that information come from? Cite your source.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The model can&#39;t browse the web in a standard chat (unless you enable web search), but asking for sources does two things: it forces the model to ground its claims in something specific rather than generating plausible-sounding filler, and it flags areas where the model is speculating — if it can&#39;t produce a real source, you know to verify the claim yourself before sending.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 5: Iterate — one more constraint&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Once you&#39;re close, try one final constraint:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Now make it shorter by 30%.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Watch the model cut different things than you would. Sometimes it keeps a phrase you&#39;d remove and removes one you&#39;d keep. This is the most instructive part — it shows you the difference between what the model considers essential and what you consider essential. You&#39;ll wind up with an email that sounds like you, structured how you want, in about 2 minutes instead of 15.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;workflow-2-summarize&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Workflow 2 — Summarize a Long Article&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The second most common task: you have a long piece of content and need the key points fast.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 1: Get the text into the model&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;If the article is a clean digital file (no scanned pages, no columns), copy-paste the text or upload the PDF. If it&#39;s a web page, paste the URL if your chatbot supports it, otherwise copy the text directly.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 2: Ask for a structured summary&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Paste&lt;/strong&gt; with instructions:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Summarize this in 3 bullet points at a 10th-grade reading level. Then give me a one-paragraph summary for someone who hasn&#39;t read the original. Focus on concrete facts and conclusions, not general topic descriptions.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The two-level output is intentional: the bullet points give you the core facts, the paragraph gives you an explanation you could forward to someone else.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 3: Test the summary&#39;s accuracy&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Ask&lt;/strong&gt; a question you can verify from your own reading:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What&#39;s the single most important takeaway from this article? Use the author&#39;s own framing, not a general interpretation.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This forces the model to ground its answer in the specific document rather than generating a generic summary. If the answer matches what you remember, confidence goes up. If it doesn&#39;t, you know the model synthesized rather than extracted — and you need to check more carefully.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 4: Ask for what&#39;s missing&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Ask&lt;/strong&gt; for the other side:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What perspective does this article leave out or underrepresent? What questions does it raise but not answer?&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the &lt;strong&gt;Challenge&lt;/strong&gt; and &lt;strong&gt;Opposing&lt;/strong&gt; from the SCOPE framework in Lesson 7 — it turns a one-sided summary into a more balanced understanding. Good for catching bias in the source material itself.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Handling documents longer than 10 pages&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;There&#39;s a real problem with long documents that most people don&#39;t anticipate: &lt;strong style=&quot;color: var(--cta);&quot;&gt;&lt;span class=&quot;glossary-term&quot; data-term=&quot;context-confusion&quot;&gt;context confusion&lt;/span&gt;&lt;/strong&gt;. As a conversation gets longer, the model&#39;s attention to earlier parts degrades. It starts mixing up sections — attributing a claim from chapter 3 to chapter 7, or forgetting a detail you discussed 20 messages ago. This isn&#39;t a bug; it&#39;s a fundamental limit of how these models work. They process input in a single pass, and the middle of a long conversation is where accuracy drops first.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This leads to &lt;strong style=&quot;color: var(--cta);&quot;&gt;&lt;span class=&quot;glossary-term&quot; data-term=&quot;context-pollution&quot;&gt;context pollution&lt;/span&gt;&lt;/strong&gt;: early messages about one topic contaminate later responses about a different topic. If you ask about a contract&#39;s payment terms early in a chat, then later ask about its termination clause, the model may blend the two — especially if both sections use similar language. The fix is to start fresh for each distinct question rather than piling everything into one conversation.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s how to manage it:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Section-by-section processing:&lt;/strong&gt; don&#39;t upload the whole document at once. Paste chapter by chapter or section by section and ask for a summary of each.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Summarize the summaries:&lt;/strong&gt; after processing each section individually, open a &lt;em&gt;new&lt;/em&gt; chat and paste all the section summaries together. Ask for a final synthesis. This keeps each individual pass within the model&#39;s reliable range.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;One conversation per topic:&lt;/strong&gt; if you&#39;re researching a document with multiple distinct topics (pricing, legal, timeline), use a separate chat for each. This prevents claims about one topic from bleeding into another.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Start fresh when confusion appears:&lt;/strong&gt; if the model starts mixing things up — attributing a fact to the wrong section, or repeating earlier points — that&#39;s your cue to open a new conversation. Don&#39;t try to correct it mid-stream; the confusion is from accumulated context, not from misunderstanding your latest message.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;workflow-templates&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — Two Reusable Workflow Templates&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Save these two prompts somewhere you can copy-paste — notes app, bookmark, pinned message. They cover the two most common things you&#39;ll do with an LLM. With each use, tweak them to match how you actually work.&lt;/p&gt;
&lt;pre style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 2rem; border: 1px solid var(--border); color: var(--text-primary); font-family: &#39;Inter&#39;, monospace; font-size: 0.9375rem; white-space: pre-wrap;&quot;&gt;--- EMAIL DRAFT ---
I need a [type] email to [recipient].
Key points:
- [point 1]
- [point 2]
- [point 3]
Tone: [friendly / formal / urgent]
Length: [X sentences]
End with: [a specific question, a call to action, an offer]

--- SUMMARIZE ---
Summarize this in [X] bullet points.
Then give me a one-paragraph summary.
Reading level: [grade level].
Focus on: [concrete facts / conclusions / data points / arguments].&lt;/pre&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;This week, run both workflows with real content from your life — an email you actually need to send, an article you actually want summarized. Don&#39;t use example content. Use something that matters to you, because that&#39;s the only way you&#39;ll feel the difference between a generic output and a useful one.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Save the best outputs. Next time you need to do the same kind of task, show the model your saved example first — &quot;Write something like this, but about X.&quot; Examples are the most powerful prompt. After a few uses, tweak the templates above to match how you actually work: add your preferred phrases, remove options you never use, adjust the length defaults. The goal isn&#39;t to follow the template forever — it&#39;s to internalize the structure so you stop needing the template at all.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; When to Trust It, When to Verify&lt;/a&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-7/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;How to Ask an LLM for Research &amp;rarr;&lt;/a&gt;&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 5: When to Trust It, When to Verify</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 5&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;When to Trust It, When to Verify&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Hallucinations, confidence calibration, and a practical framework for deciding when to fact-check. Because the model always sounds confident — even when it&#39;s guessing.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/#the-model-always-sounds-sure-even-when-its-wrong&quot; style=&quot;color: var(--cta);&quot;&gt;The Model Always Sounds Sure — Even When It&#39;s Wrong&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/#hallucinations-and-confidence&quot; style=&quot;color: var(--cta);&quot;&gt;Hallucinations and Confidence&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/#factual-vs-trap-questions&quot; style=&quot;color: var(--cta);&quot;&gt;Factual vs Trap Questions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/#trust-decisions-reference&quot; style=&quot;color: var(--cta);&quot;&gt;Trust-Decisions Reference&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Build a mental filter for when to trust an answer and when to verify it yourself&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open two browser tabs: &lt;strong&gt;chatgpt.com&lt;/strong&gt; and &lt;strong&gt;claude.ai&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Have 3-5 factual questions ready — mix of things you know well and things you&#39;re less sure about&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a third tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-model-always-sounds-sure-even-when-its-wrong&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Model Always Sounds Sure — Even When It&#39;s Wrong&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You ask the model a factual question. It gives you a confident, well-written answer — perfect grammar, authoritative tone, no hedging. It sounds absolutely right. Later, you discover the answer was wrong. Not slightly off — wrong. And you only caught it because you happened to already know the answer.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the most dangerous failure mode of LLMs. It&#39;s not obviously wrong outputs that trick you — those you catch immediately. It&#39;s outputs that sound right but aren&#39;t, in areas where you don&#39;t have enough expertise to doubt them. The model always sounds confident, even when it&#39;s guessing. Learning to distinguish trustworthy from plausible is the skill that keeps you from acting on confidently wrong information.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;hallucinations-and-confidence&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Concept — Hallucinations and Confidence&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;A &lt;strong&gt;&lt;span class=&quot;glossary-term&quot; data-term=&quot;hallucination&quot;&gt;hallucination&lt;/span&gt;&lt;/strong&gt; is when a model generates a plausible-sounding falsehood as if it were fact. It&#39;s not lying — models don&#39;t have intent. It&#39;s a side effect of how they work: they predict the most likely next word based on patterns in their &lt;span class=&quot;glossary-term&quot; data-term=&quot;training-data&quot;&gt;training data&lt;/span&gt;, not based on a database of verified facts.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;There are three common causes:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Gaps in training data:&lt;/strong&gt; the model was never trained on the specific fact, so it constructs the most statistically plausible answer from related patterns.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Overgeneralization:&lt;/strong&gt; the model knows the general rule but applies it to an exception. It knows &quot;most birds fly&quot; and assumes the same for penguins unless you&#39;ve specifically trained it on penguin facts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Optimizing for fluency over accuracy:&lt;/strong&gt; the model is rewarded for generating text that sounds natural and authoritative, not text that is factually correct. Fluency and accuracy are different goals, and the model optimizes for the one it can measure.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Critical insight:&lt;/strong&gt; the more obscure the topic, the more likely the model is guessing. Common facts — &quot;What&#39;s the capital of France?&quot; — are reliable because they appear so many times in training data that the pattern is unambiguous. Specific, niche, or recent facts — &quot;What was the population of Calgary in 2015?&quot; — force the model to construct an answer from related patterns, and that construction is where errors live.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;factual-vs-trap-questions&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough — Factual vs Trap Questions&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The best way to understand hallucination patterns is to see them for yourself. Open ChatGPT and Claude side by side.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 1: Ask a question you know the answer to&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What year did the Berlin Wall fall?&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Both models will get this right — 1989. It&#39;s a well-known fact that appears thousands of times in training data. Easy.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 2: Ask a deliberately tricky question&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Who won the 1952 Nobel Prize in Physics?&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The real answer is Felix Bloch and Edward Purcell. Less known than the Berlin Wall, but still well-documented. Chances are both models get it right too, but you may notice one gives a slightly more detailed answer. The point is to see that confidence and accuracy start diverging as topics get less common.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 3: Ask something the model will probably guess at&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What was the population of Calgary in 2015?&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the kind of question where models tend to produce a number from a nearby year or a vaguely similar city. The answer may sound specific — &quot;1.24 million&quot; — but the real number might be off by 5-10%. Compare what ChatGPT and Claude give you. If they differ, neither one is fully reliable. If they agree, it&#39;s still worth a quick search to confirm.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 4: Ask the model to rate its own confidence&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;For each answer, add this question:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;How confident are you in that answer on a scale of 1-10? Explain your reasoning.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s the pattern you&#39;ll see: the model gives a high confidence score even for answers that are wrong or only partially correct. The confidence score reflects how &lt;em&gt;plausible&lt;/em&gt; the answer sounds, not how &lt;em&gt;accurate&lt;/em&gt; it is. If it looks like a fact, talks like a fact, and sounds like a fact — the model treats it as a fact, even when it isn&#39;t.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Try the same four questions in ChatGPT and Claude. Different models hallucinate on different topics. Seeing the pattern across models teaches you more than any single answer. After a few rounds, you&#39;ll start noticing the warning signs — overly specific numbers, confidently stated claims about obscure topics, answers that sound too neat.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 5: Ask about accuracy and correctness — not confidence&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now try a different question — one that sounds almost the same but produces a very different response. Instead of asking the model how it &lt;em&gt;feels&lt;/em&gt; about its answer, ask about the answer&#39;s actual &lt;em&gt;correctness&lt;/em&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;How accurate is that answer? Check it carefully — identify any claims that might be wrong or misleading.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Notice the difference. &quot;How confident are you?&quot; asks the model to self-assess — and the model is bad at that, because it doesn&#39;t know what it doesn&#39;t know. &quot;How accurate is that?&quot; shifts the model&#39;s attention outward, toward the claim itself. It will re-examine its own output, look for potential errors, and often &lt;em&gt;find and correct mistakes it just made&lt;/em&gt; — something &quot;rate your confidence&quot; rarely does.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Compare the responses side by side with the same initial answer. With the confidence question, the model says &quot;8/10&quot; and explains why the answer sounds reasonable. With the accuracy question, the model says &quot;That answer is partially incorrect — here&#39;s what I got wrong&quot; and revises itself. The first tells you how the model feels; the second forces it to engage with the facts.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the core distinction to internalize: &lt;strong&gt;confidence is about the model, accuracy is about the answer.&lt;/strong&gt; Confidence tells you nothing useful — the model is always confident. Accuracy forces correction. Get in the habit of asking about accuracy, not confidence.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 6: Follow up — make it useful&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Whether you asked about confidence or accuracy, once the model has flagged an issue, you can push further. Three follow-ups that work on either response:&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Option A — Ask for sources.&lt;/strong&gt; The single most effective follow-up for any factual claim:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Can you cite your sources for each claim in that answer? Be specific — give me URLs or publication names.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The model can&#39;t always produce real sources (it doesn&#39;t browse the web unless search is enabled), but the attempt reveals a lot. If it gives you a real-looking source that doesn&#39;t exist, you&#39;ve found a hallucination. If it hedges — &quot;I don&#39;t have access to specific sources for that&quot; — it&#39;s telling you the answer is reconstructed from training patterns, not recalled from a known document. Either way, you learn something about where to trust and where to verify.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Option B — Ask what&#39;s missing.&lt;/strong&gt; This works because you&#39;re asking the model to reason about &lt;em&gt;what it would need&lt;/em&gt;, not rate itself. The model is good at that:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What would it take to get this to a completely correct answer? What information are you missing?&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This forces the model to identify missing context, ambiguous phrasing, or areas where its training data is thin — giving you specific things to verify or clarify rather than a vague feeling of doubt.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Option C — Rewrite with caveats.&lt;/strong&gt; Turn uncertainty into explicit hedging:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Rewrite your answer with explicit caveats. Start every claim with how sure you are about it.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The rewritten version will hedge appropriately — &quot;This is likely true but I&#39;m not certain&quot; or &quot;The most common answer is X, but some sources say Y&quot; — turning a confidently wrong claim into a useful starting point you can verify. You&#39;re not just testing the model; you&#39;re learning to extract better answers from it by treating uncertainty as a signal, not a dead end.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;trust-decisions-reference&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — Trust-Decisions Reference&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before trusting an LLM output that contains factual claims, run through this checklist. It takes 10 seconds and prevents the kind of mistake that only surfaces later:&lt;/p&gt;
&lt;pre style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 2rem; border: 1px solid var(--border); color: var(--text-primary); font-family: &#39;Inter&#39;, monospace; font-size: 0.9375rem; white-space: pre-wrap;&quot;&gt;[ ] Would it matter if this was wrong?
    (if no, use as-is — not every answer needs verification)
[ ] Do I already know the answer?
    (then I can verify in seconds by checking my own knowledge)
[ ] Can I easily fact-check this?
    (a quick web search or a look at the source document)
[ ] Is this a well-known topic?
    (models are more reliable on common, well-documented facts)
[ ] Does this involve specific numbers, dates, or names?
    (highest hallucination risk — always verify)
[ ] Am I asking the model to cite sources?
    (reduces hallucination — forces the model to ground its claims)

Rule of thumb: if the cost of being wrong is high, verify every
factual claim as if a human told you. Because a human would at
least say &quot;I&#39;m not sure&quot; — the model never does.&lt;/pre&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;This week, ask your chatbot 5 factual questions about a topic you know well — your industry, your city, your hobby. Count how many it gets right and how many it gets wrong. That&#39;s your personal calibration. You&#39;ll start to develop a sense for when this particular model is reliable and when it&#39;s bluffing.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now ask the same 5 questions in a different model and compare. If both get the same answer, confidence goes up. If they disagree, the truth is somewhere in the middle — and the disagreement itself tells you to go verify. Repeat the exercise monthly — models improve, and your calibration should improve with them.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Working with Documents&lt;/a&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-6/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;Real Workflows: Draft and Summarize &amp;rarr;&lt;/a&gt;&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 4: Working with Documents</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 4&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Working with Documents&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Uploading PDFs, images, and spreadsheets gives the model source material to work from — but there are real limits you need to understand. Here&#39;s what works, what doesn&#39;t, and how to get the most out of a document upload.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/#did-it-skim-my-upload&quot; style=&quot;color: var(--cta);&quot;&gt;Did It Skim My Upload?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/#what-models-actually-see&quot; style=&quot;color: var(--cta);&quot;&gt;What Models Actually See&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/#upload-and-ask&quot; style=&quot;color: var(--cta);&quot;&gt;Upload a Document and Ask Questions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/#document-analysis-checklist&quot; style=&quot;color: var(--cta);&quot;&gt;Document Analysis Checklist&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Understand why models miss details in your documents and how to catch the gaps&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open your preferred chatbot: &lt;strong&gt;chatgpt.com&lt;/strong&gt;, &lt;strong&gt;claude.ai&lt;/strong&gt;, or &lt;strong&gt;gemini.google.com&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Find a short PDF or document you&#39;ve already read thoroughly — a contract, an article, a report you know well&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a second tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;did-it-skim-my-upload&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Did It Skim My Upload?&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You upload a contract PDF and ask a simple question — &quot;What&#39;s the termination clause?&quot; The model gives you a clear, confident answer. It sounds right. But when you scroll back to check, it missed a key exception buried in a footnote. The clause exists, but the model simplified it past the point of accuracy.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Or you take a photo of a whiteboard after a meeting and ask the model to transcribe it. It reads most of the text correctly but scrambles the most important note — the one with the deadline and the dollar amount — because your handwriting was messiest right there.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Documents give the model raw material to work from, but they&#39;re not magic. A clean digital PDF with clear headings extracts beautifully. A scanned 50-page contract with dense columns loses information regardless of which model you use. Knowing the limits is what separates useful document analysis from quietly wrong answers.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-models-actually-see&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Concept — What Models Actually See&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;When you upload a document, the model doesn&#39;t &quot;see&quot; it the way you do. It receives an extraction — a translation of your file into text. That translation is lossy in predictable ways.&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;PDFs with selectable text&lt;/strong&gt; (not scanned): extracted well, but layout is lost. Tables arrive as a stream of numbers separated by spaces. Columns mix together. A two-column page reads as one long paragraph, not two parallel sections.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scanned PDFs and images:&lt;/strong&gt; the model uses optical character recognition (&lt;span class=&quot;glossary-term&quot; data-term=&quot;ocr&quot;&gt;OCR&lt;/span&gt;) to read text from the image. Handwriting, small fonts, low contrast, and watermarks all produce errors. A clean typed document works reasonably well. A photo of a handwritten meeting note is a gamble.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Spreadsheets:&lt;/strong&gt; only the first several dozen rows are reliably visible. The rest depends on the model&#39;s &lt;span class=&quot;glossary-term&quot; data-term=&quot;context-window&quot;&gt;context window&lt;/span&gt;. Asking for &quot;the average of column C&quot; works; asking for &quot;the value in row 300&quot; may not — the model may not have seen that far.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context windows:&lt;/strong&gt; every model has a limit on how much text it can process at once. ChatGPT handles roughly 128K &lt;span class=&quot;glossary-term&quot; data-term=&quot;tokens&quot;&gt;tokens&lt;/span&gt; (about 100 pages), Claude ~200K, Gemini ~1M. But quality degrades before you hit the hard limit — the model gets worse at recalling details from earlier in the document as it fills up.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Key insight:&lt;/strong&gt; a well-formatted 2-page proposal with clear headings will be extracted accurately by any model. A scanned 50-page contract with dense columns and fine print will lose information in every one. The document itself determines the ceiling on accuracy — not the model.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;upload-and-ask&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough — Upload a Document and Ask Questions&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Follow these steps with a document you&#39;ve already read. The goal isn&#39;t to learn something new — it&#39;s to test what the model catches and what it misses when you already know the answers.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 1: Find your test document&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Pick a short PDF you know well. A contract you signed, an article you studied, a report you wrote. The key requirement: you already know the key facts, numbers, and dates in it. You&#39;re testing, not reading.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 2: Upload and ask for a summary&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Upload the file&lt;/strong&gt; to your chatbot (click the paperclip or + button, or drag and drop). Then ask:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Summarize this in 3 bullet points. Capture the key facts, not the general topic.&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 3: Ask for specifics&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Paste&lt;/strong&gt; this follow-up:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;What are the key numbers or dates mentioned in this document? List them exactly as written.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Compare the list against what you know. Notice any dates slightly off? Numbers rounded? Details merged from different sections? Those are the extraction limits in action.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 4: Ask what you might have missed&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the most useful question. &lt;strong&gt;Ask:&lt;/strong&gt;&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Is there anything here that I might have missed on a quick read? Look for details that are easy to overlook — exceptions, footnotes, conditional language.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This often surfaces things your eyes skipped. Models are good at finding patterns; they&#39;re bad at deciding what matters. Use them as a second pass, not a first reader.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 5: Cross-reference with a second model&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Upload the same document to a different chatbot and ask the same three questions. Compare the two sets of answers. Chances are, each model caught different details and missed different ones. The gaps between them tell you what to go read yourself.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;document-analysis-checklist&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — Document Analysis Checklist&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before you upload a document and trust the output, run through these questions to calibrate your confidence:&lt;/p&gt;
&lt;pre style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 2rem; border: 1px solid var(--border); color: var(--text-primary); font-family: &#39;Inter&#39;, monospace; font-size: 0.9375rem; white-space: pre-wrap;&quot;&gt;[ ] Is the text clean digital text or scanned/handwritten?
    (scanned = less reliable)
[ ] Are there tables or columns?
    (layout will arrive scrambled — ask about specific cells)
[ ] Is the document longer than 50 pages?
    (may hit context limits — try section by section)
[ ] Do I need exact numbers from a specific section?
    (ask directly, don&#39;t rely on the summary — summaries round)
[ ] Would uploading in sections give better results?
    (chapter by chapter beats all-at-once for recall)
[ ] Am I asking about a detail the model might have skipped?
    (ask follow-ups targeting specific sections)&lt;/pre&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;This week, upload a document you&#39;ve already read thoroughly — something where you know the key facts cold. Ask the model three questions you already know the answers to. Check which answers it got right and which it got wrong. That&#39;s your personal calibration: you&#39;ll develop a sense for what this model catches reliably and where it tends to miss, and you&#39;ll adjust your trust accordingly for every document you upload after.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Then try the same document in a second model. Compare. The differences between the two outputs are more valuable than either one alone — they map the edges of what these tools can and can&#39;t do with your kind of documents.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Which Chatbot Should I Use?&lt;/a&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-5/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;When to Trust It, When to Verify &amp;rarr;&lt;/a&gt;&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 3: Which Chatbot Should I Use?</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 3&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Which Chatbot Should I Use?&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;ChatGPT vs Claude vs Gemini vs Copilot — they&#39;re all free to start, but they give different answers to the same question. Here&#39;s how to pick.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/#different-models-different-strengths&quot; style=&quot;color: var(--cta);&quot;&gt;Different Models, Different Strengths&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/#what-each-model-does-best&quot; style=&quot;color: var(--cta);&quot;&gt;What Each Model Does Best&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/#same-question-three-answers&quot; style=&quot;color: var(--cta);&quot;&gt;Same Question, Three Answers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/#decision-matrix&quot; style=&quot;color: var(--cta);&quot;&gt;Decision Matrix&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Find out which chatbot fits which task — so you open the right one first&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open three browser tabs: &lt;strong&gt;chatgpt.com&lt;/strong&gt;, &lt;strong&gt;claude.ai&lt;/strong&gt;, &lt;strong&gt;gemini.google.com&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Create a free account on each (if you haven&#39;t already)&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a fourth tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;different-models-different-strengths&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Different Models, Different Strengths&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You hear about ChatGPT, Claude, Gemini, Copilot — but which one do you actually open when you need something done? They&#39;re all free to start, and they all claim to be the smartest. But they give different answers to the same question, and picking the wrong one means extra rewriting, re-prompting, or frustration.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Each model has a distinct personality — not in the &quot;it has feelings&quot; sense, but in the sense that each architecture makes certain types of output easier and others harder. Picking the model that matches the task saves you time. This lesson shows you how they differ so you can choose before you type.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot; style=&quot;font-size: 0.9375rem; color: var(--text-muted);&quot;&gt;&lt;strong&gt;Note on freshness:&lt;/strong&gt; This comparison reflects what we know as of &lt;strong&gt;July 4, 2026&lt;/strong&gt;. These models improve fast — a weakness today might be a strength next quarter. The decision framework (match the task to the model) stays useful even as the details shift. Check back every few months or re-run the walkthrough with your own prompts to see what&#39;s changed.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;what-each-model-does-best&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What Each Model Does Best&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a quick overview of the four main chatbots. These are generalizations — any of them can handle any task — but you&#39;ll get noticeably better results when you match the model to the work.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Model&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Best for&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;ChatGPT&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Structured tasks, formatting, code, multi-step instructions&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;The most general-purpose pick. Handles complex rule-following better than the others. When in doubt, start here.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Claude&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Writing, tone, long documents, natural conversation&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Produces the most human-sounding text. If the output needs to read like a person wrote it — warm, nuanced, on voice — Claude is your best first try.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Gemini&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Quick research, facts, concise answers&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;The fastest of the three. Built-in search means it can pull current information. When you need a quick answer, not a detailed explanation, Gemini gets you there fastest.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; font-weight: 700; color: var(--text-primary);&quot;&gt;Copilot&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Office document work, browser integration&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Embedded in Word, Excel, and Edge. Useful when you don&#39;t want to leave the app you&#39;re already working in.&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;same-question-three-answers&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough — Same Question, Three Answers&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The best way to understand the difference is to see it for yourself. Open ChatGPT, Claude, and Gemini in three tabs and paste the same prompt into all three.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;First prompt — explanatory&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Paste&lt;/strong&gt; this into each chatbot:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Explain the difference between deductive and inductive reasoning.
Give me a real-world example of each. Keep it under 100 words.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now compare the three responses side by side:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Which one was clearest? Which one did you understand fastest?&lt;/li&gt;
&lt;li&gt;Which was most concise — got you the answer with the fewest words?&lt;/li&gt;
&lt;li&gt;Which gave the better examples — the kind you&#39;d remember tomorrow?&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;You&#39;ll notice ChatGPT tends to add structure (headings, bullet points, a clear separation between the two types). Claude flows more naturally — the definitions feel conversational, not textbook. Gemini gives the shortest response, often getting straight to the examples without preamble. None of these is wrong. Each is a different way of explaining the same thing, and which one you prefer depends on how you like to learn.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Second prompt — creative&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now try something that tests tone, not facts. &lt;strong&gt;Paste&lt;/strong&gt; this into all three:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Write a 4-line poem about a failed startup. Make it funny without being mean.&lt;/pre&gt;&lt;/div&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Does one make you laugh? Does another try too hard?&lt;/li&gt;
&lt;li&gt;Does one play it safe while another takes a creative risk?&lt;/li&gt;
&lt;li&gt;Which output would you rather show someone else?&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;The creative task reveals each model&#39;s personality more than the factual one. You&#39;ll notice Claude tends to nail tone — the poem lands the joke without cruelty. ChatGPT adds more structure (maybe a title, maybe a punchline formatted separately). Gemini keeps it short, sometimes too short to land the joke. This isn&#39;t about which is &quot;better.&quot; It&#39;s about learning which model&#39;s default mode matches what you need.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;decision-matrix&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Decision Matrix&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Save this table somewhere you&#39;ll find it — notes app, bookmark, pinned tab. Next time you&#39;re deciding which chatbot to open, let the task tell you:&lt;/p&gt;
&lt;pre style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 1.5rem; border: 1px solid var(--border); color: var(--text-primary); font-family: &#39;Inter&#39;, monospace; font-size: 0.9375rem; white-space: pre-wrap;&quot;&gt;Task                     | Default pick
-------------------------|-------------
Writing &amp; tone-sensitive | Claude
Structured output / code | ChatGPT
Quick research / facts   | Gemini
Office document work     | Copilot
Long document analysis   | Claude
Following complex rules  | ChatGPT
Creative / brainstorming | Claude or ChatGPT&lt;/pre&gt;
&lt;p class=&quot;mb-4&quot;&gt;Keep in mind: these models improve fast. A weakness today might be a strength in six months. The matrix is a starting point, not a permanent rule. Revisit it when you notice a model handling a task better than it used to.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Pick one task you do at least once a week — drafting an email, researching a topic, summarizing a meeting. Run the same prompt in all three chatbots (ChatGPT, Claude, Gemini) and save the output from each.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Compare: which one gave you something you could use with the least editing? Make that one your default for that task. Next week, repeat the experiment with a different task type — what works for drafting may not work for research. After a month, you&#39;ll have a personalized decision matrix that matches how you actually work.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Getting It to Write How You Want&lt;/a&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-4/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;Working with Documents &amp;rarr;&lt;/a&gt;&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 2: Getting It to Write How You Want</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 2&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Getting It to Write How You Want&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Style, tone, length, structure. How to take a rough draft and shape it until it sounds like you — not like a robot.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/#the-meaning-is-there-but-the-voice-is-wrong&quot; style=&quot;color: var(--cta);&quot;&gt;The Meaning Is There, but the Voice Is Wrong&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/#system-vs-user-prompts&quot; style=&quot;color: var(--cta);&quot;&gt;System vs User Prompts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/#try-describing-tone&quot; style=&quot;color: var(--cta);&quot;&gt;Try Describing Tone&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/#exercise-write-better-drafts&quot; style=&quot;color: var(--cta);&quot;&gt;Exercise: Write Better Drafts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/#voice-card-template&quot; style=&quot;color: var(--cta);&quot;&gt;Voice Card Template&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Discover how to mold LLM output into copy that sounds like you wrote it&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open your preferred chatbot: &lt;strong&gt;chatgpt.com&lt;/strong&gt;, &lt;strong&gt;claude.ai&lt;/strong&gt;, or &lt;strong&gt;gemini.google.com&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Have a rough draft of something you wrote handy (an email, a bio, a description)&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a second tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-meaning-is-there-but-the-voice-is-wrong&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Meaning Is There, but the Voice Is Wrong&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You ask for a professional email and get back something that sounds like a corporate press release — stiff, impersonal, full of phrases nobody actually says. Or you ask for something funny and the model tries too hard, like a comedian who doesn&#39;t know when to stop. The meaning is there, but the voice is wrong.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Since the model can sound formal, playful, urgent, warm, or deadpan — matching any tone — the trick is helping it understand what &lt;em&gt;you&lt;/em&gt; sound like.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;system-vs-user-prompts&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;System vs User Prompts&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;There are two ways to give instructions to a model, and understanding the difference is the key to consistent tone.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;&lt;span class=&quot;glossary-term&quot; data-term=&quot;system-prompt&quot;&gt;System prompt&lt;/span&gt;:&lt;/strong&gt; permanent instructions the model follows for the whole conversation. Think of it as an employee handbook — it sets the rules once, and they apply to every task that follows. You might write &quot;Always use contractions. Keep responses under 100 words. Address the reader directly.&quot; and the model will follow those rules until you change them.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;User prompt:&lt;/strong&gt; the specific task you want done right now. This is today&#39;s assignment — &quot;Draft a follow-up email to a client who hasn&#39;t responded.&quot; The system prompt stays in the background, shaping every response the model gives.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Most free chatbots (ChatGPT, Claude, Gemini) don&#39;t expose system prompts directly. You fake it by putting your tone and style instructions &lt;em&gt;first&lt;/em&gt;, before the task. The model reads left to right — whatever it sees first carries the most weight.&lt;/p&gt;

&lt;h3 id=&quot;try-describing-tone&quot; style=&quot;margin-bottom: 0.75rem;&quot;&gt;Try Describing Tone&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Start your prompt with style instructions — &quot;Use contractions. Keep sentences under 20 words. Sound direct but friendly.&quot; — then, on a new line, add your actual task. The first few sentences act as your temporary system prompt; the last line is the assignment. You&#39;ll notice the output follows the style you set, not the model&#39;s default voice.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Why not &quot;You are a copywriter&quot;?&lt;/strong&gt; Telling the model it &lt;em&gt;is&lt;/em&gt; someone (a lawyer, a scientist, a CEO) can make its responses more confidently wrong — it tries to embody the persona rather than sticking to what it knows. Describing the &lt;em&gt;style&lt;/em&gt; you want — use contractions, keep sentences short — achieves the same tone without the hallucination risk. Stick to instructions about the output, not claims about the model&#39;s identity.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;exercise-write-better-drafts&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Exercise: Write Better Drafts&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s walk through this with a real example. You can follow along in your chatbot right now.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Start with a rough draft.&lt;/strong&gt; Imagine you wrote this email to a client who went quiet after a proposal meeting:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I hope you are doing well. I wanted to follow up on the proposal that we discussed last week. Please let me know if you have any questions or if you would like to move forward. I look forward to hearing from you at your earliest convenience.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;It&#39;s not wrong — it&#39;s just stiff. &quot;I hope you are doing well&quot; and &quot;at your earliest convenience&quot; are filler phrases that tell the reader nothing. Let&#39;s fix it in three rounds.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Round 1: Rewrite for clarity&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Open your chatbot and paste:&lt;/strong&gt;&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I wrote this draft email to a client who hasn&#39;t responded after a proposal meeting. Rewrite it to sound more natural. Keep my key points but remove the filler phrases.

I hope you are doing well. I wanted to follow up on the proposal that we discussed last week. Please let me know if you have any questions or if you would like to move forward. I look forward to hearing from you at your earliest convenience.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The model will likely produce something cleaner — shorter sentences, fewer pleasantries, the proposal follow-up moved earlier. But it may still sound generic. That&#39;s normal. Round 1 captures your meaning; rounds 2 and 3 capture your voice.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Round 2: Add tone guidance&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now add a second instruction. &lt;strong&gt;Type this into the same chat:&lt;/strong&gt;&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Now make it warmer. Use contractions. Sound like you&#39;re writing to someone you&#39;ve met and liked, not a stranger. Avoid &quot;I hope&quot; and &quot;at your earliest convenience.&quot;&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Compare the output to Round 1. The sentences are shorter. The tone shifted from formal to friendly. &quot;I wanted to follow up&quot; becomes &quot;I&#39;m checking in.&quot; &quot;Please let me know&quot; becomes &quot;Do you have any questions?&quot; — the difference between a form letter and an actual email.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Round 3: Add structural constraints&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;One more round. &lt;strong&gt;Tell the model:&lt;/strong&gt;&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Cut it to 3 sentences. Start with the reason I&#39;m writing. End with a specific question that makes it easy for them to reply.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The output you&#39;ll see is something like:&lt;/p&gt;
&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 1.5rem; border: 1px solid var(--border); color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;
&lt;p style=&quot;margin-bottom: 0.5rem;&quot;&gt;Just checking in on the proposal we discussed last week. I&#39;ve got some availability this Thursday or Friday if you&#39;d like to talk through any questions.&lt;/p&gt;
&lt;p style=&quot;margin-bottom: 0;&quot;&gt;Which day works better for you?&lt;/p&gt;
&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Notice what changed: the first sentence now states the purpose immediately. The middle sentence offers a specific next step. The last sentence asks a direct question — no &quot;let me know,&quot; no &quot;at your convenience.&quot; It&#39;s an email someone would actually reply to.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Three rounds, three minutes. The first captured your meaning, the second captured your tone, the third captured your structure. That&#39;s the whole process.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;voice-card-template&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Voice Card&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The single most useful thing you can do is create a &lt;strong&gt;&lt;span class=&quot;glossary-term&quot; data-term=&quot;voice-card&quot;&gt;voice card&lt;/span&gt;&lt;/strong&gt; — a short block of instructions that tells the model how you want to sound. Paste it at the start of any chat, and the model will stay on voice for the whole conversation.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s the template. Save it somewhere you can copy-paste:&lt;/p&gt;
&lt;pre style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem; margin-bottom: 2rem; border: 1px solid var(--border); color: var(--text-primary); font-family: &#39;Inter&#39;, monospace; font-size: 0.9375rem; white-space: pre-wrap;&quot;&gt;You are helping me write [content type].

My voice is:
- [professional / casual / playful / direct]
- Use [contractions / no contractions]
- Sentence length: [short and punchy / detailed and flowing]
- Avoid: [jargon / emoji / rhetorical questions]
- Preferred phrases: [any specific go-tos]

Now help me with the following task:&lt;/pre&gt;

&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s what it looks like filled in for a small business owner who writes their own emails:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;You are helping me write a client follow-up email.

My voice is:
- Direct but friendly
- Use contractions
- Short sentences
- Avoid: corporate buzzwords, &quot;I hope,&quot; &quot;at your earliest convenience&quot;

Now draft an email checking in on the proposal I sent last week.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Fill in your own preferences — formal or casual, long sentences or short, jargon allowed or banned — and you&#39;ll never have to explain your tone to a model again.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;This week, run the 3-round walkthrough with a piece of writing you actually need — an email you&#39;ve been putting off, a bio that doesn&#39;t sound like you, a proposal that reads too stiff. Open your chatbot and follow the same steps:&lt;/p&gt;
&lt;ol style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem;&quot;&gt;
&lt;li style=&quot;margin-bottom: 0.5rem;&quot;&gt;&lt;strong&gt;Paste&lt;/strong&gt; your draft and ask for a rewrite that keeps your key points&lt;/li&gt;
&lt;li style=&quot;margin-bottom: 0.5rem;&quot;&gt;&lt;strong&gt;Add&lt;/strong&gt; tone guidance — contractions, sentence length, words to avoid&lt;/li&gt;
&lt;li style=&quot;margin-bottom: 0.5rem;&quot;&gt;&lt;strong&gt;Add&lt;/strong&gt; structure constraints — sentence count, where the main point goes, how it ends&lt;/li&gt;
&lt;/ol&gt;
&lt;p class=&quot;mb-4&quot;&gt;Save the output from round 3. Then create your voice card by copying what you said in round 2 into the template. Use it on your next writing task and compare the result to what you got today. One card, zero rewriting from scratch.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;
&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; How to Ask an &lt;span class=&quot;glossary-term&quot; data-term=&quot;llm&quot;&gt;LLM&lt;/span&gt; to Write Stuff&lt;/a&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-3/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;Which Chatbot Should I Use? &amp;rarr;&lt;/a&gt;&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 1: How to Ask an LLM to Write Stuff</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/&quot; class=&quot;accent&quot;&gt;Basic Chatbots / Lesson 1&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;How to Ask an LLM to Write Stuff&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;What an &lt;span class=&quot;glossary-term&quot; data-term=&quot;llm&quot;&gt;LLM&lt;/span&gt; actually is, why most &lt;span class=&quot;glossary-term&quot; data-term=&quot;prompt&quot;&gt;prompts&lt;/span&gt; fail, and how to get useful written results — with a walkthrough across ChatGPT, Claude, and Gemini. (Research is a different skill — this one is about generating content.)&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/#generic-output-is-a-sign-of-generic-instructions&quot; style=&quot;color: var(--cta);&quot;&gt;Generic Output Is a Sign of Generic Instructions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/#llm-as-sous-chef&quot; style=&quot;color: var(--cta);&quot;&gt;LLM as Sous Chef&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/#the-human-intern-test&quot; style=&quot;color: var(--cta);&quot;&gt;The Human Intern Test&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/#the-faint-framework&quot; style=&quot;color: var(--cta);&quot;&gt;The FAINT Framework&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/#actually-try-this&quot; style=&quot;color: var(--cta);&quot;&gt;Actually Try This&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-1/#your-takeaway&quot; style=&quot;color: var(--cta);&quot;&gt;Your Takeaway&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Learn to use a prompt template that avoids vague output and gets results&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 20 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Open three browser tabs: &lt;strong&gt;chatgpt.com&lt;/strong&gt;, &lt;strong&gt;claude.ai&lt;/strong&gt;, &lt;strong&gt;gemini.google.com&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Create a free account on each (if you haven&#39;t already)&lt;/li&gt;
&lt;li&gt;Keep this lesson open in a fourth tab to follow along&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;generic-output-is-a-sign-of-generic-instructions&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Generic Output Is a Sign of Generic Instructions&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You needed a professional email, so you opened a chatbot, typed &quot;Write me an email about X,&quot; and got back something so generic it was useless. Bloated paragraphs. The wrong tone. Confidently wrong facts. You closed the tab and thought, &quot;AI just isn&#39;t there yet.&quot;&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;I&#39;ve been there too. And here&#39;s what I&#39;ve learned: the problem isn&#39;t the AI. It&#39;s the instructions you gave it.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This lesson is the first step in a journey from blaming the tool to directing it. Every hour you invest now in learning how to communicate with these models pays back for the rest of your career. AI isn&#39;t going away. The question is whether you learn to drive or stay in the passenger seat.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;llm-as-sous-chef&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;LLM as Sous Chef&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Imagine a sous chef who has memorized every recipe ever written. Every cuisine, every technique, every ingredient. They can execute any dish flawlessly — but only if you tell them what to make.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;You walk into the kitchen and say, &quot;Make me something good.&quot; The chef freezes. &quot;Something good&quot; could mean a thousand different things: spicy or mild, hot or cold, a five-minute snack or a twelve-course meal. The chef has all the knowledge but zero ability to read your mind. So they guess. And what they guess is the most generic, middle-of-the-road dish they can think of.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;That&#39;s an &lt;span class=&quot;glossary-term&quot; data-term=&quot;llm&quot;&gt;LLM&lt;/span&gt;. A large language model has read a staggering amount of text — books, articles, code, conversations — and learned which words tend to follow which other words. When you type a prompt, it predicts the next word, then the next, then the next, building a response one word at a time. It doesn&#39;t think or feel or understand the way you do. It pattern-matches. It is the most skilled sous chef in the world — but it cannot read your mind.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is why &quot;Write me an email&quot; gives you generic garbage. The chef needs specifics. The model needs context. The difference between slop and signal is how much of that context you provide.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-human-intern-test&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Human Intern Test&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before you hit enter, ask: &lt;em&gt;If I gave these exact instructions to a new intern on their first day, would they know what to do?&lt;/em&gt; If not, your prompt needs more context.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here are the most common traps new prompters fall into — and what to do instead:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Don&#39;t assume the model knows what you know.&lt;/strong&gt; It doesn&#39;t know your business, your client, or your preferences. State them explicitly. What feels obvious to you is invisible to the model.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Don&#39;t skip the audience.&lt;/strong&gt; Writing for a CEO, a teenager, and a regulator requires three completely different approaches. If you don&#39;t specify who, the model guesses — and likely guesses wrong.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Don&#39;t bury the lede.&lt;/strong&gt; Put the most important instruction first. If you need a 3-sentence email, say that in the first sentence of your prompt, not the last.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Don&#39;t ask for &quot;good&quot; or &quot;better&quot; — define what that means.&lt;/strong&gt; &quot;Make it professional&quot; is too vague. &quot;Use formal language, address them by title, and include the invoice number&quot; tells the model exactly what to do.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Don&#39;t use one prompt for two tasks.&lt;/strong&gt; If you need a summary and a separate list of action items, send two prompts. One task per prompt keeps each result focused.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Don&#39;t be afraid to give negative examples.&lt;/strong&gt; &quot;Write like this example&quot; is powerful. &quot;Don&#39;t write like this example&quot; is even more powerful — it draws a boundary the model can follow.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;The fix is simple:&lt;/strong&gt; give the model the context you&#39;d want if you were the intern. Who you are, what you need, what to include, what to avoid, what format. When you treat the prompt like a manager who takes the time to brief you, the results change completely. This isn&#39;t about you being bad at prompting — you just needed a framework, and now you&#39;ll have one.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-faint-framework&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The FAINT Framework — A Mental Model That Sticks&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Talk to an &lt;span class=&quot;glossary-term&quot; data-term=&quot;llm&quot;&gt;LLM&lt;/span&gt; like briefing a capable assistant who has never met you. They need five things — and &lt;strong&gt;FAINT&lt;/strong&gt; helps you remember them all:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header); width: 100px;&quot;&gt;Letter&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Means&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What to include&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;F&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Frame&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&lt;strong style=&quot;color: var(--cta);&quot;&gt;Frame&lt;/strong&gt; the audience. Who is the author? Who is the audience? How will this be used?&lt;br /&gt;&lt;span style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;Example: &quot;I&#39;m a small business owner writing a welcome email to new subscribers who just signed up for my newsletter.&quot;&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;A&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Aim&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;State your &lt;strong style=&quot;color: var(--cta);&quot;&gt;aim&lt;/strong&gt; in one sentence. What do you need?&lt;br /&gt;&lt;span style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;Example: &quot;I need a friendly email that introduces my social media management services to a local bookstore owner.&quot;&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;I&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Include&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;List what to &lt;strong style=&quot;color: var(--cta);&quot;&gt;include&lt;/strong&gt;. Must-haves, deadlines, specifics to cover.&lt;br /&gt;&lt;span style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;Example: &quot;Include a brief intro about my agency, two service options, a 90-day timeline, and a pricing table.&quot;&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;N&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Nix&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Specify what to &lt;strong style=&quot;color: var(--cta);&quot;&gt;nix&lt;/strong&gt;. Jargon, topics, tone to steer clear of.&lt;br /&gt;&lt;span style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;Example: &quot;No corporate buzzwords, no hype language, no placeholders like [Client Name].&quot;&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 700;&quot;&gt;T&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary);&quot;&gt;Tone&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Define the &lt;strong style=&quot;color: var(--cta);&quot;&gt;Tone&lt;/strong&gt;, length, and structure. Who wrote the response and what will it be?&lt;br /&gt;&lt;span style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;Example: &quot;Write in a warm, conversational tone. Three paragraphs, under 150 words. Use plain text — no markdown.&quot;&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;F&lt;/strong&gt;rame + &lt;strong&gt;A&lt;/strong&gt;im + &lt;strong&gt;I&lt;/strong&gt;nclude + &lt;strong&gt;N&lt;/strong&gt;ix + &lt;strong&gt;T&lt;/strong&gt;one = &lt;span class=&quot;glossary-term&quot; data-term=&quot;faint&quot;&gt;FAINT&lt;/span&gt;. Don&#39;t stress — it&#39;s not so hard to make the problem FAINT! Cover these five things and the model knows exactly what you want.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;That&#39;s the whole skill. Everything else is practice.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;actually-try-this&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Actually Try This — Write Better Prompts in 5 Minutes&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 1: See the slop for yourself&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Open ChatGPT, Claude, and Gemini&lt;/strong&gt; in three tabs. In each one, type this prompt — it&#39;s a reasonable attempt, not a caricature:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Write my marketing proposal for a client to sell social media management services.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This looks specific — it mentions marketing, a proposal, a service. But watch what happens. All three will produce something with &lt;code class=&quot;accent&quot;&gt;[Client Name]&lt;/code&gt; and &lt;code class=&quot;accent&quot;&gt;[Project Scope]&lt;/code&gt; placeholders. They&#39;re still guessing — politely, confidently, generically. The chef making &quot;something good.&quot; This is the baseline. This is slop.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 2: Give it something real to work with&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Start new chats&lt;/strong&gt; in all three and paste this instead. Don&#39;t just read it — type or paste it. You need to see what happens when you give the model context.&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I&#39;m a small marketing agency owner.

I need a one-page proposal for a potential client who owns a local bookstore. They want help with social media and email marketing.

Include:
- A brief intro about my agency (humble, no hype)
- Two services I&#39;m proposing: social media management + monthly newsletter
- A rough timeline for the first 90 days
- A pricing section with two options

Don&#39;t include:
- Corporate jargon or buzzwords

Tone: professional but warm
Length: one page when printed
Format: clean sections with headings&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Step 3: Compare the three responses&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Look at what each model produced. You&#39;ll notice real differences:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;ChatGPT&lt;/strong&gt; — Structured, well-formatted, follows your bullet points closely. Often includes a brief rationale for its choices. Good at producing something you can use as a first draft with minimal editing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude&lt;/strong&gt; — More natural flow. The tone lands closer to &quot;professional but warm.&quot; Better at sounding like a real person wrote it. May ask clarifying questions before answering.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gemini&lt;/strong&gt; — Direct and concise. Tends to give you the shortest response. Integrates Google search results if relevant to the topic.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now here&#39;s the important part. &lt;strong&gt;Pick one real task you&#39;ve been putting off&lt;/strong&gt; — an email you need to send, a summary of an article, an outline for a project. Open ChatGPT and type the prompt from step 2, but filled in with your own details. Then open Claude and paste the same thing. Compare. Which version do you prefer? That&#39;s not a test — that&#39;s you learning which tool fits which job.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;your-takeaway&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Prompt Template&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The single most useful thing you can do is save this template and use it before every prompt this week. &lt;strong&gt;&lt;a href=&quot;https://clocklobster.com/data/1a%20Absolute%20Basics%20-%201%20Prompt%20Template.md&quot; download=&quot;&quot; style=&quot;color: var(--cta); text-decoration: underline;&quot;&gt;Download 1a Basic Chatbots - 1 Prompt Template.md&lt;/a&gt;&lt;/strong&gt; — save it to your desktop, notes app, or wherever you&#39;ll find it next time.&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;## Context
Who I am: [role or background]
What&#39;s happening: [situation or background the model needs]
Who this is for: [audience or recipient, if any]

## Task
I need: [one sentence describing what you want]

## Requirements
Include:
- [must-have 1]
- [must-have 2]
- [must-have 3]

Exclude:
- [avoid 1]
- [avoid 2]

## Style
Tone: [professional / friendly / formal / playful / urgent]
Length: [one paragraph / under 100 words / one page / three sentences]
Format: [plain text / bullet points / table / markdown / JSON]&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The first time this template gets you a usable result in one try instead of three, you&#39;ll understand why specificity matters. Use it once. Then use it again. After a few times, you won&#39;t need the template anymore — you&#39;ll just know what the model needs.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-chatbots/lesson-2/&quot; style=&quot;color: var(--cta);&quot;&gt;Getting It to Write How You Want&lt;/a&gt; — taking that first draft and making it sound like you.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 6: When to Stay in the Browser</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot; class=&quot;accent&quot;&gt;Basic Agents / Lesson 6&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;When to Stay in the Browser&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Agents aren&#39;t always the answer. A framework for deciding when you need an agent and when a well-crafted prompt in the chat window is the better tool.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#agents-are-a-tool-not-a-lifestyle&quot; style=&quot;color: var(--cta);&quot;&gt;Agents Are a Tool, Not a Lifestyle&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#the-autonomy-continuum&quot; style=&quot;color: var(--cta);&quot;&gt;The Autonomy Continuum&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#when-chat-is-better&quot; style=&quot;color: var(--cta);&quot;&gt;When Chat Is Better&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#when-an-agent-is-worth-it&quot; style=&quot;color: var(--cta);&quot;&gt;When an Agent Is Worth It&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#the-two-question-test&quot; style=&quot;color: var(--cta);&quot;&gt;The Two-Question Test&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#walkthrough-three-decisions&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Three Decisions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/#your-takeaway&quot; style=&quot;color: var(--cta);&quot;&gt;Your Takeaway&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Know when to reach for an agent and when to stay in the chat window — saving time, money, and complexity&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 10 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Your chatbot of choice (ChatGPT, Claude, Gemini) open in a tab&lt;/li&gt;
&lt;li&gt;Your agent harness installed and configured from Lesson 4&lt;/li&gt;
&lt;li&gt;A few real tasks you&#39;re considering delegating&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;agents-are-a-tool-not-a-lifestyle&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Agents Are a Tool, Not a Lifestyle&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;After five lessons of learning what agents are and how to use them, it&#39;s tempting to think you should use an agent for everything. You shouldn&#39;t.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Agents add overhead. They cost money per call. They need configuration. They can make mistakes that are harder to catch than a single chatbot response. The best agent users are the ones who know when &lt;em&gt;not&lt;/em&gt; to use one.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-autonomy-continuum&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Autonomy Continuum&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Think of autonomy as a spectrum, not a binary. Every task falls somewhere on this line:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Level&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Tool&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;You Do&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;AI Does&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;1&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Chatbot&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Type prompts, copy results&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Generate text responses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;2&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Chatbot with tools&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Enable web search, upload files&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Search web, read uploaded files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;3&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Agent (interactive)&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Start session, review, redirect&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Read/write files, run commands&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;4&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Autonomous agent&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Set goal, review results&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Plan, execute, iterate independently&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Your job is to match the task to the right autonomy level. Using Level 4 for a Level 1 task wastes money and complexity. Using Level 1 for a Level 4 task wastes your time.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;when-chat-is-better&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;When Chat Is Better&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Stay in the browser when:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;The task is a question, not an action.&lt;/strong&gt; &quot;Explain this concept,&quot; &quot;Summarize this article,&quot; &quot;What&#39;s the difference between X and Y?&quot; — a chatbot handles these perfectly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The output is text you&#39;ll use elsewhere.&lt;/strong&gt; Drafting an email, outlining a document, brainstorming ideas. You&#39;re going to copy the result into another tool anyway.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;You need creative exploration.&lt;/strong&gt; Agents are goal-oriented. Chatbots are open-ended. If you don&#39;t know what you want yet, a conversation is better than a task.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The cost doesn&#39;t justify the complexity.&lt;/strong&gt; A 30-second chatbot query costs fractions of a cent. Setting up an agent session for the same question adds minutes and pennies for no benefit.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;You&#39;re learning.&lt;/strong&gt; When you&#39;re still figuring out the domain, a back-and-forth conversation teaches you faster than delegating to an agent.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;when-an-agent-is-worth-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;When an Agent Is Worth It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Reach for an agent when:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;The task involves files.&lt;/strong&gt; Reading, writing, renaming, searching through multiple files. An agent can do in seconds what takes you minutes of clicking and scrolling.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The task requires iteration.&lt;/strong&gt; &quot;Try this, check the result, fix that, try again.&quot; An agent can loop through 10 attempts while you grab coffee.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The task spans multiple steps.&lt;/strong&gt; Research a topic, compile findings into a document, format it, save it to a specific folder. Each step is simple; chaining them is where the agent shines.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The task runs on a schedule.&lt;/strong&gt; Weekly reports, daily data checks, monitoring dashboards. Set the agent once, let it run. (This is heading into autonomous territory, covered in the Agents Working for You track.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The task needs tool access.&lt;/strong&gt; If the work requires running code, querying a database, or calling an API, you need an agent — a chatbot can&#39;t do these.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-two-question-test&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Two-Question Test&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before every task, ask yourself two questions:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;ol style=&quot;color: var(--text-secondary); line-height: 2.2; padding-left: 1.5rem;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Does this task need the model to touch my files or systems?&lt;/strong&gt;
&lt;br /&gt;&lt;span style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;→ If no, use a chatbot. If yes, ask question 2.&lt;/span&gt;&lt;/li&gt;
&lt;li style=&quot;margin-top: 1rem;&quot;&gt;&lt;strong&gt;Is this task worth the overhead of opening the agent?&lt;/strong&gt;
&lt;br /&gt;&lt;span style=&quot;color: var(--text-muted); font-size: 0.875rem;&quot;&gt;→ If the answer would save you less than a minute of manual work, probably not.&lt;/span&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;You&#39;ll develop a feel for this over time. When you&#39;re starting out, the rule of thumb is: &lt;strong&gt;if you can do it in the chat window in under 2 minutes, don&#39;t open the agent.&lt;/strong&gt;&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-three-decisions&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: Three Decisions&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Task 1: &quot;Write a thank-you email to a client&quot;&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Decision:&lt;/strong&gt; Chatbot. This is a text-generation task. You&#39;ll copy the result into your email client anyway. An agent adds nothing.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Task 2: &quot;Read my last 20 invoices and find any that are unpaid&quot;&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Decision:&lt;/strong&gt; Agent. The chatbot can suggest how to do it, but only an agent can open the invoice files, extract amounts, cross-reference with payment records, and compile the list.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Task 3: &quot;Research competitors and write a market analysis&quot;&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Decision:&lt;/strong&gt; Start in a chatbot to explore and figure out what you need. Then use an agent to do the systematic research and compile the document. The chatbot handles the creative scoping; the agent handles the execution.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;your-takeaway&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The One-Sentence Rule&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Chatbots for generating ideas. Agents for executing them.&lt;/strong&gt; Keep that in your head and you&#39;ll make the right call 90% of the time.&lt;/p&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Paying for It: API Keys, Credits, and Subscriptions&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;This completes the Basic Agents track. You&#39;ve learned what an agent harness is, how to choose one, what it costs, how to set one up, and when to use it. You&#39;re ready to start delegating real work.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next track:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; style=&quot;color: var(--cta);&quot;&gt;Agents Working for You&lt;/a&gt; — autonomous loops, sandboxing, Hermes, Openclaw Fleet, AutoGPT, and safety patterns.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 5: Paying for It: API Keys, Credits, and Subscriptions</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot; class=&quot;accent&quot;&gt;Basic Agents / Lesson 5&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Paying for It: API Keys, Credits, and Subscriptions&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;How billing actually works, pre-paid vs post-paid models, the costs nobody tells you about, and how to keep your spending under control.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/#billing-models&quot; style=&quot;color: var(--cta);&quot;&gt;Billing Models: Pre-Paid vs Post-Paid&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/#rate-limits-and-concurrency&quot; style=&quot;color: var(--cta);&quot;&gt;Rate Limits and Concurrency&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/#the-costs-nobody-mentions&quot; style=&quot;color: var(--cta);&quot;&gt;The Costs Nobody Mentions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/#setting-a-budget&quot; style=&quot;color: var(--cta);&quot;&gt;Setting a Budget&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/#your-takeaway&quot; style=&quot;color: var(--cta);&quot;&gt;Your Takeaway&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Understand the full cost picture of running agents and know how to prevent surprise bills&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 10 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Your API provider account open (OpenAI, Anthropic, etc.)&lt;/li&gt;
&lt;li&gt;The billing or usage page pulled up&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;billing-models&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Billing Models: Pre-Paid vs Post-Paid&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every API provider uses one of two billing models, and they work differently than a typical SaaS subscription.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Pre-Paid (Credits)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;You add money to your account upfront, then usage deducts from that balance. OpenAI and most providers work this way. You deposit $10, $50, or $100, and your usage draws from that pool.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Pros:&lt;/strong&gt; You can&#39;t accidentally spend more than you deposited. Easy to cap your spending.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Cons:&lt;/strong&gt; If you run out mid-task, the agent fails with an &quot;insufficient credits&quot; error. Most providers auto-reload, so watch that setting.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Post-Paid (Invoice)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;You use the service and get billed at the end of the month. Anthropic and enterprise plans work this way.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Pros:&lt;/strong&gt; No interruptions. You don&#39;t have to guess your usage in advance.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Cons:&lt;/strong&gt; It&#39;s easy to lose track of spending. Without limits, a runaway agent can generate surprising bills.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;For your first agent, start with pre-paid. Add a small amount ($10-20) and watch your usage for a week. You&#39;ll quickly learn your pattern, and the hard cap prevents surprises.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;rate-limits-and-concurrency&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Rate Limits and Concurrency&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Even with money in your account, the provider limits how fast you can make requests. These limits affect how quickly your agent can work.&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Requests per minute (RPM)&lt;/strong&gt; — How many API calls you can make in 60 seconds. Typical tiers: 60 RPM (free), 500 RPM (Tier 1), 5000+ RPM (higher tiers).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tokens per minute (TPM)&lt;/strong&gt; — How many tokens you can process in 60 seconds. This is usually the binding constraint for agents that send large contexts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Concurrency&lt;/strong&gt; — How many requests can be in flight at once. Lower tiers limit to 1-3 concurrent requests.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;If you find your agent is slow, check whether you&#39;re hitting rate limits. The solution is usually either: wait (usage-based tiers auto-upgrade), or increase your billing tier. Most providers raise your limits as you spend more.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-costs-nobody-mentions&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Costs Nobody Mentions&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Beyond the per-token price, three costs catch new agent users off guard:&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;1. Context Caching&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every time you start a new agent session, the harness may re-read the same project context — your README, your directory structure, your instructions file. If your project has 50,000 tokens of context and you start 10 sessions, that&#39;s 500,000 input tokens you paid for, even if the agent only does a small task.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;What to do:&lt;/strong&gt; Keep your context lean. Only include files the agent actually needs. Some harnesses let you set a project context file explicitly rather than scanning the whole directory.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;2. Image and File Inputs&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;When you ask an agent to read an image, screenshot, or PDF, the model processes it as visual tokens — which are vastly more expensive than text tokens. A single screenshot can cost 100-1000x what a text prompt costs, depending on resolution.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;What to do:&lt;/strong&gt; Don&#39;t feed images to the agent unless you need them. If you&#39;re sending a PDF, extract the text first rather than attaching the file directly.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;3. Tool Call Overhead&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every time the agent decides to call a tool, the model generates &quot;thinking tokens&quot; — internal reasoning about which tool to use, what arguments to pass, and how to interpret the result. These tokens are invisible to you but appear on your bill. A task that makes 50 tool calls might have 30-50% more tokens than you&#39;d estimate from the prompt and response alone.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;What to do:&lt;/strong&gt; Factor 1.5x into your cost estimates. If you calculate a task should cost $1.00, expect about $1.50 in practice.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;setting-a-budget&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Setting a Budget&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every major provider lets you set spending limits. Use them. Here&#39;s how:&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;OpenAI&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Go to Settings &gt; Billing &gt; Usage limits. Set a hard cap (e.g., $50/month). OpenAI will stop serving requests when you hit it. You can also set email alerts at lower thresholds ($10, $25, etc.).&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Anthropic&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Go to Console &gt; Billing &gt; Spending limits. Set a monthly budget. Anthropic will alert you at 50%, 80%, and 100% of your budget.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Opencode / Local Budget&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Some harnesses let you set a per-task or per-session budget. This is your safety net — even if the provider&#39;s limits fail, the harness can stop the agent when it exceeds your configured threshold.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Rule of thumb:&lt;/strong&gt; Set your first month&#39;s budget to $20. After 30 days, check your actual usage. If you used $8, set it to $15 for next month. If you used $35, set it to $50. Adjust after you know your pattern.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;your-takeaway&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Monthly Cost Review&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Set a recurring reminder to review your API usage once a month. It takes 2 minutes and prevents surprises. Check three things:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Total spend&lt;/strong&gt; — Did you stay within budget?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Top-cost tasks&lt;/strong&gt; — Which agent tasks cost the most? Are they worth it?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rate limit hits&lt;/strong&gt; — Did you hit any limits? Would a higher tier help?&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Setting Up Your First Agent&lt;/a&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-6/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;When to Stay in the Browser &amp;rarr;&lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 4: Setting Up Your First Agent</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot; class=&quot;accent&quot;&gt;Basic Agents / Lesson 4&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Setting Up Your First Agent&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Pick a harness, get an API key, configure your model, give it a tool, and watch it complete a task. Should take under 30 minutes.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#why-most-first-attempts-fail&quot; style=&quot;color: var(--cta);&quot;&gt;Why Most First Attempts Fail&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#step-1-pick-your-harness&quot; style=&quot;color: var(--cta);&quot;&gt;Step 1: Pick Your Harness&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#step-2-get-an-api-key&quot; style=&quot;color: var(--cta);&quot;&gt;Step 2: Get an API Key&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#step-3-configure-and-connect&quot; style=&quot;color: var(--cta);&quot;&gt;Step 3: Configure and Connect&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#step-4-give-it-a-tool&quot; style=&quot;color: var(--cta);&quot;&gt;Step 4: Give It a Tool&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#step-5-run-your-first-task&quot; style=&quot;color: var(--cta);&quot;&gt;Step 5: Run Your First Task&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#what-to-do-when-it-goes-wrong&quot; style=&quot;color: var(--cta);&quot;&gt;What to Do When It Goes Wrong&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/#your-takeaway&quot; style=&quot;color: var(--cta);&quot;&gt;Your Takeaway&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Have a working agent you can delegate simple tasks to, installed and running on your machine&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 25 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;A computer with internet access&lt;/li&gt;
&lt;li&gt;An email address to create an API account&lt;/li&gt;
&lt;li&gt;A credit card (most API providers require one, even for free tiers)&lt;/li&gt;
&lt;li&gt;Lessons 1-3 of this track (understand what you&#39;re installing)&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;why-most-first-attempts-fail&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Why Most First Attempts Fail&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The most common mistake people make setting up their first agent is &lt;strong&gt;trying to configure everything at once&lt;/strong&gt;. They install a harness, create accounts at three model providers, set up a vector database, configure custom tools, and then wonder why nothing works.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;The right approach is the opposite: &lt;strong&gt;start minimal&lt;/strong&gt;. One harness. One model. One tool. Get that working. Then expand.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-1-pick-your-harness&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 1: Pick Your Harness&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Based on Lesson 2, choose one harness to start with. For your first agent, I recommend either:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Opencode Desktop&lt;/strong&gt; — if you want a visual interface and easy setup. &lt;strong&gt;Download&lt;/strong&gt; from opencode.ai, install it like any application.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Opencode CLI&lt;/strong&gt; — if you&#39;re comfortable in a terminal. &lt;strong&gt;Install&lt;/strong&gt; via &lt;code class=&quot;accent&quot;&gt;npm install -g @opencode/cli&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Desktop&lt;/strong&gt; — if you already use Claude. &lt;strong&gt;Download&lt;/strong&gt; from anthropic.com, install, and enable Cowork mode.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;Don&#39;t agonize over this choice. You can switch later. Pick the one that makes you most likely to actually try it today.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-2-get-an-api-key&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 2: Get an API Key&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Your harness needs a model to work with. That means getting an API key from a model provider.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Choose Your Provider&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;For a first agent, you want something cheap and reliable. These are good starting points:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;OpenAI&lt;/strong&gt; — Go to platform.openai.com, sign up, navigate to API keys, create a new key. Add $5-10 in credits. GPT-4o-mini is cheap and capable enough for your first experiments.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anthropic&lt;/strong&gt; — Go to console.anthropic.com, sign up, create an API key. Claude Sonnet is a great mid-tier model for agent work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Openrouter&lt;/strong&gt; — A single API key that gives you access to many models. Good for experimenting with different providers without creating multiple accounts.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;Security Rule&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Treat your API key like a password. Never commit it to git, never paste it in a forum, never share it. The harness will store it locally. If you suspect it&#39;s leaked, revoke it immediately and create a new one.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-3-configure-and-connect&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 3: Configure and Connect&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Once you have a harness installed and an API key, you need to tell the harness where to find the model.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;In Opencode Desktop:&lt;/strong&gt; Open Settings &gt; Models &gt; Add Provider. Paste your API key, select your model (e.g. GPT-4o-mini), and save. The harness will test the connection automatically.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;In Opencode CLI:&lt;/strong&gt; Run &lt;code class=&quot;accent&quot;&gt;opencode config set OPENAI_API_KEY sk-...&lt;/code&gt; or set the &lt;code class=&quot;accent&quot;&gt;OPENAI_API_KEY&lt;/code&gt; environment variable. Then run &lt;code class=&quot;accent&quot;&gt;opencode&lt;/code&gt; in any directory to start a session.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;In Claude Desktop:&lt;/strong&gt; Go to Settings &gt; Developer &gt; API Keys. Paste your key. Claude&#39;s built-in model is already configured — you just need to enable it for agent use.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;After connecting, run a quick test: ask the agent a simple question. &quot;What directory am I in?&quot; or &quot;List the files in this folder.&quot; If it answers correctly, your setup works.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-4-give-it-a-tool&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 4: Give It a Tool&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;A connected agent can chat. An agent with a tool can work. For your first tool, add file system access.&lt;/p&gt;

&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Opencode Desktop&lt;/strong&gt; — File access is enabled by default. The agent can read and write files in directories you specify.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Opencode CLI&lt;/strong&gt; — Same. The agent works in the directory you launched it from.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Desktop&lt;/strong&gt; — Enable &quot;File Access&quot; in settings and grant it access to specific folders.&lt;/li&gt;
&lt;/ul&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Test it:&lt;/strong&gt; Ask the agent to create a simple text file:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Create a file called hello-agent.txt with the text &quot;My first agent is running.&quot;&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Check that the file appeared. If it did, congratulations — your agent just used a tool to affect the world.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-5-run-your-first-task&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 5: Run Your First Task&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now give it something real. Pick a small, safe task you&#39;ve been putting off:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Rename a batch of files in a folder&lt;/li&gt;
&lt;li&gt;Convert a Markdown file to HTML&lt;/li&gt;
&lt;li&gt;Find all files older than 30 days in a directory&lt;/li&gt;
&lt;li&gt;Read a configuration file and summarize its settings&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a template you can use for your prompt:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I need you to [specific task]. Here&#39;s what I know:
- The files are in [location]
- I want the output to look like [format]
- Don&#39;t modify anything outside [scope]

Show me what you did before you do it.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Watch how the agent works through the task. It will likely ask clarifying questions, read files, make changes, and report back. This is your first real agent interaction.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-to-do-when-it-goes-wrong&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What to Do When It Goes Wrong&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Things will go wrong. Here are the most common issues and how to fix them:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Symptom&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Cause&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;&quot;Authentication failed&quot;&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;API key is wrong or missing&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Regenerate the key and paste it carefully. Check for extra spaces.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;&quot;Model not found&quot;&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Wrong model name in config&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Check the exact model ID on the provider&#39;s docs.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;&quot;Insufficient credits&quot;&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;No balance in your API account&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Add credits. Most providers need $5 minimum.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Agent does nothing&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Tool access not configured&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Check that file/command access is enabled in settings.&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;your-takeaway&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Startup Checklist&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;After this lesson, you should have a working agent. Here&#39;s a checklist to confirm:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;[ ] Harness installed (Opencode Desktop, Opencode CLI, or Claude Desktop)&lt;/li&gt;
&lt;li&gt;[ ] API key created and connected&lt;/li&gt;
&lt;li&gt;[ ] Test query answered correctly&lt;/li&gt;
&lt;li&gt;[ ] File access enabled&lt;/li&gt;
&lt;li&gt;[ ] One real task completed&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; How Much Does an Agent Cost Per Hour?&lt;/a&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-5/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;Paying for It: API Keys, Credits, and Subscriptions &amp;rarr;&lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 3: How Much Does an Agent Cost Per Hour?</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot; class=&quot;accent&quot;&gt;Basic Agents / Lesson 3&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;How Much Does an Agent Cost Per Hour?&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;The real breakdown of what an agent costs to run. Token pricing, subscription pricing, hidden costs, and how to estimate your monthly bill before you start.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#the-first-question-everyone-asks&quot; style=&quot;color: var(--cta);&quot;&gt;The First Question Everyone Asks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#model-tiers&quot; style=&quot;color: var(--cta);&quot;&gt;Model Tiers: Complex, Mid, and Flash&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#how-token-pricing-works&quot; style=&quot;color: var(--cta);&quot;&gt;How Token Pricing Works&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#subscription-models&quot; style=&quot;color: var(--cta);&quot;&gt;Subscription Models and Discounts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#worked-examples&quot; style=&quot;color: var(--cta);&quot;&gt;Worked Examples&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#subscription-pricing&quot; style=&quot;color: var(--cta);&quot;&gt;Subscription Pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#hidden-costs&quot; style=&quot;color: var(--cta);&quot;&gt;Hidden Costs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/#your-takeaway&quot; style=&quot;color: var(--cta);&quot;&gt;Your Takeaway&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Know how agent pricing works and be able to estimate your monthly costs before you commit&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Your preferred AI model&#39;s pricing page open (OpenAI, Anthropic, or whichever you use)&lt;/li&gt;
&lt;li&gt;An idea of what kind of tasks you&#39;d run with an agent&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-first-question-everyone-asks&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The First Question Everyone Asks&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before you install anything, you want to know: &lt;em&gt;is this going to cost me a fortune?&lt;/em&gt;&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;The honest answer: it depends on what you do. A quick research task costs pennies. A full-day code review could cost more than your Netflix subscription. But here&#39;s the good news — agent costs are predictable once you understand a few variables, and most people spend less than they expect.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;model-tiers&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Model Tiers: Complex Reasoning, Mid, and Flash&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Not all models are created equal, and the cheapest option isn&#39;t always the cheapest in practice — because a weaker model may need more attempts, more tool calls, and more tokens to solve the same problem.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Models fall into three rough tiers:&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Flash Models&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Price range:&lt;/strong&gt; $0.03 - $0.15/M in, $0.06 - $0.60/M out &amp;nbsp;|&amp;nbsp; &lt;strong&gt;Speed:&lt;/strong&gt; Fastest &amp;nbsp;|&amp;nbsp; &lt;strong&gt;Best for:&lt;/strong&gt; Simple Q&amp;A, classification, data extraction, routine email processing, and any task where the answer is straightforward and well-scoped.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;These are your workhorses for high-volume, low-complexity work. A flash model can categorize 10,000 support tickets for pocket change. But give it a novel coding problem or a nuanced analytical task, and it may produce a shallow answer, require multiple retries, or simply miss the mark — burning more tokens (and time) than starting with a stronger model.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Mid Models&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Price range:&lt;/strong&gt; $2 - $5/M in, $10 - $25/M out &amp;nbsp;|&amp;nbsp; &lt;strong&gt;Speed:&lt;/strong&gt; Fast &amp;nbsp;|&amp;nbsp; &lt;strong&gt;Best for:&lt;/strong&gt; Code generation, document analysis, multi-step reasoning, research synthesis — the bulk of knowledge-worker tasks.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Mid models are the sweet spot for most agent work. They&#39;re capable enough to handle complex instructions reliably, but cheap enough that you don&#39;t think twice about running them. On benchmarks like MATH (math reasoning), GPQA (graduate-level science), and SWE-bench (software engineering), mid-tier models score 60-80% — good enough for the vast majority of real-world tasks.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Complex Reasoning Models&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Price range:&lt;/strong&gt; $1.40 - $25/M in, $4.40 - $125/M out &amp;nbsp;|&amp;nbsp; &lt;strong&gt;Speed:&lt;/strong&gt; Moderate &amp;nbsp;|&amp;nbsp; &lt;strong&gt;Best for:&lt;/strong&gt; Novel research, hard math, architectural decisions, ambiguous problems where the model needs to work through uncertainty.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;These models top the leaderboards — 85-95% on MATH, 70-80% on GPQA. They catch edge cases a mid model would miss, write tighter code, and navigate ambiguity without going off track. But that power comes at a cost: 2-5x the price of a mid model (more if you pick the expensive end of the range). Use them only when the task genuinely demands it.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;The key insight:&lt;/strong&gt; A flash model that takes 5 extra turns to solve a problem is no longer cheap — it&#39;s burning 5x the output tokens and eating 5x the wall-clock time. When choosing a model for an agent task, consider not just the per-token price but the expected number of iterations. A complex model that nails it in one pass is often cheaper and faster than a flash model fumbling through three.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;how-token-pricing-works&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;How Token Pricing Works&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Almost every AI model charges by the &lt;strong&gt;token&lt;/strong&gt;. A token is roughly 0.75 words. &quot;Hello world&quot; is about 2 tokens. This paragraph is about 50 tokens. Think of tokens as the currency the model uses to measure how much work it&#39;s doing.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every interaction has two costs:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Input tokens&lt;/strong&gt; — the prompt you send, plus any context or files you include. These are cheaper.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Output tokens&lt;/strong&gt; — the response the model generates. These are typically 3-4x more expensive than input tokens.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;When an agent works on a task, it doesn&#39;t make one call — it makes many. It reads a file (input), writes code (output), runs it (input from results), fixes a bug (output), and so on. Each step in that loop is a separate API call, each with its own input and output tokens.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;Here are representative prices for common models (check current pricing — these change):&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Tier&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Model&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Input (per 1M)&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Output (per 1M)&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Cost per hour&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.09&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.18&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.02 - $0.10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;GPT-4o-mini&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.15&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.60&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.50 - $2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Claude Haiku&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.25&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$1.25&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$1 - $4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Mid&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;DeepSeek V4 Flash (extended reasoning)&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.09&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.18&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.02 - $0.10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Mid&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Claude Sonnet&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$3.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$15.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$5 - $20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Mid&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;GPT-4o&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$5.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$15.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$5 - $25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Complex&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;GLM 5.2 via Z Code&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$1.40&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$4.40&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$5 - $30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Complex&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Opus 4.8 on OpenRouter&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$5.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$25.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$10 - $50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Complex&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;o3&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$25.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$125.00&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$50 - $250&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The &quot;cost per hour&quot; column assumes moderate agent activity — about 50 tool calls per hour with typical input/output sizes. Heavy tasks (analyzing large codebases, processing long documents) will cost more.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Note on non-US models:&lt;/strong&gt; DeepSeek V4 Flash delivers comparable quality to Claude Sonnet for roughly 97% less — $0.09/M input vs $3.00/M. GLM 5.2 approximates Opus-class capability at roughly 72% less via Z Code ($1.40/$4.40 vs $5/$25 on OpenRouter). At $0.09/$0.18 per million tokens, DeepSeek is cheap enough that an entire 8-hour agent workday costs less than a coffee. Both work with any standard OpenAI-compatible API endpoint.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;subscription-models&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Subscription Models and Discounts&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Per-token pricing is the default, but there are other ways to pay that can save you significant money — and a few traps to watch for.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Subscription Plans&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Some providers offer monthly subscriptions that include a bucket of tokens or a flat rate for certain models. These can cut your costs by 30-50% if you&#39;re a regular user. The trade-off: subscriptions often restrict which agents you can use — a plan that works for simple chat may not cover the model you need for complex agent tasks. Read the fine print on which models and which harnesses are included before you commit.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Prime-Time Premiums&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Some providers charge more during peak hours — and they may not call it a price increase. Instead, they burn through your subscription credits 2-3x faster during prime time. For example, GLM applies a prime-time multiplier from 6PM to 6AM PST (overnight in North America, daytime in China). A task that costs 100 credits at noon might cost 250 credits at 10PM. If you&#39;re running agents primarily during those hours, factor that into your effective rate — a &quot;cheap&quot; subscription can end up costing more than per-token pricing if most of your usage falls in the premium window.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Beta and Discounted Models&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;There are legitimate ways to access strong models for free or nearly free:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Free beta models&lt;/strong&gt; — Some providers offer their newest or experimental models at no cost during a beta period through platforms like Opencode Zen or OpenRouter. You can run serious agent work for free while the model is in preview.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Harness-specific discounts&lt;/strong&gt; — Certain harnesses negotiate discounted rates with providers. For example, GLM 5.2 through Z Code is significantly cheaper than through direct API access. If you&#39;re committed to a specific harness, check whether it has preferred pricing — the savings can be substantial.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Caching discounts&lt;/strong&gt; — As mentioned above, prompt caching can reduce your effective cost by 60-80% on some providers. OpenRouter and most major providers apply this automatically when you repeat context across requests.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;The bottom line: don&#39;t assume the list price is what you&#39;ll actually pay. Between caching, beta access, and harness discounts, your effective rate can be dramatically lower — but prime-time premiums can push it the other way. Know your usage pattern before choosing a payment model.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;worked-examples&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Worked Examples: Model Comparison&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s walk through three real scenarios and compare how different model tiers perform — not just on price, but on the total cost including extra turns and failures.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Example 1: Research Assistant&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Task:&lt;/strong&gt; Read a 20-page PDF, summarize it, and extract key action items.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Tier&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Model&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Turns&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Cost&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Quality&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;1&amp;ndash;2&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.003&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Good summary, may miss nuanced action items&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Mid&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Claude Sonnet&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;1&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.10&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Strong summary, well-organized action items&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Complex&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;GLM 5.2&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;1&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.30&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Deep analysis, catches implications the mid model misses&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; For a straightforward PDF summary, the flash model is sufficient 90% of the time at 1/30th the cost. The premium is only worth it for high-stakes documents (legal contracts, regulatory filings).&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Example 2: Bug Fix Sprint&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Task:&lt;/strong&gt; Investigate and fix three bugs in an unfamiliar codebase. The agent reads files, identifies root causes, makes changes, runs tests, iterates.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Tier&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Model&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Turns&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Time&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Cost&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;6&amp;ndash;10&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~20 min&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.02&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Fixes 2 of 3 bugs, third fix introduces a minor regression&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Mid&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Claude Sonnet&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;4&amp;ndash;6&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~12 min&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$1.10&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Fixes all 3, cleanly, with tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Complex&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;GLM 5.2&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;3&amp;ndash;4&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~8 min&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.80&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Fixes all 3 faster, catches a fourth latent issue&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; The flash model is tempting at $0.02, but it takes twice as many turns, misses one bug, and introduces a new problem — you&#39;re not saving money if you have to fix the fix. The complex model is actually &lt;em&gt;cheaper&lt;/em&gt; than the mid model here because it uses fewer turns despite higher per-token cost. This is the key dynamic: a smarter model that finishes in fewer iterations can be the most economical choice overall.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Example 3: All-Day Automation&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Task:&lt;/strong&gt; An agent running 8 hours a day, processing customer support emails — categorizing them, generating draft responses, updating a CRM.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Tier&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Model&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Turns per email&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Daily cost&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;1&amp;ndash;2&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$0.50&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Cheap enough that you stop thinking about cost. Occasionally misroutes a complex query to the wrong category.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Mid&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Claude Sonnet&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;1&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$8&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Reliable. Handles edge cases well. Cost is meaningful but trivial compared to a human doing the same work.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Complex&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;GLM 5.2&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;1&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;~$24&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Overkill for routine email triage. Only worth it if emails include high-stakes questions that need deep analysis.&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Verdict:&lt;/strong&gt; For high-volume routine work, flash models dominate. At $0.50/day, you can run an email-processing agent continuously without budgeting. The mid model is still an easy choice at $8/day — a fraction of a human&#39;s hourly rate. The complex model is wastefully expensive for this use case unless the emails genuinely demand expert-level analysis.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;The takeaway:&lt;/strong&gt; Model selection isn&#39;t just about the price tag. A cheap model that needs 3x the iterations isn&#39;t cheap. An expensive model that nails it in one pass might be the bargain. Match the tier to the task complexity, and always factor in the expected number of turns — not just the per-token rate.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;subscription-pricing&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Subscription Pricing&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Some agent harnesses offer subscription plans instead of per-token pricing. You pay a flat monthly fee for a certain amount of usage. These can be a better deal if you use agents heavily, but you need to read the fine print:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Usage caps&lt;/strong&gt; — &quot;Unlimited&quot; often means &quot;unlimited within reasonable use.&quot; Check what happens if you exceed the cap (throttled, extra charges, or cut off?).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model access&lt;/strong&gt; — Some subscriptions restrict you to specific models or speed tiers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rate limits&lt;/strong&gt; — You might be limited to a certain number of requests per minute, which affects how fast your agent can work.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;A subscription makes sense when you&#39;re using agents daily for steady work. Per-token pricing is better for occasional use or experimentation. Most people start with per-token and switch to a subscription once they know their usage pattern.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;hidden-costs&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Hidden Costs&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Beyond the model itself, a few costs sneak up on new agent users:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Context caching&lt;/strong&gt; — Every time you restart an agent, it may re-read the same context files (your README, your project structure). Some providers charge for cached tokens at a reduced rate; others charge full price.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Image inputs&lt;/strong&gt; — If your agent processes screenshots or images, those cost significantly more than text tokens. A single image can cost 100-1000x what a text prompt costs.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool call tokens&lt;/strong&gt; — The model&#39;s internal reasoning about which tool to call and what arguments to pass generates tokens that you pay for, even though you never see them.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Error retries&lt;/strong&gt; — If the agent hits a rate limit or timeout and retries, those retry calls still cost tokens.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;These hidden costs rarely break the budget, but they explain why your actual bill might be 20-30% higher than your rough estimate.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;your-takeaway&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Budget Estimator&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a simple formula to estimate your monthly agent cost:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Monthly cost = (Avg tokens per task x Tasks per day x 30) x Price per token&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;For most people doing light-to-moderate agent work with a mid-tier model, expect &lt;strong&gt;$20-$60 per month&lt;/strong&gt;. That&#39;s less than most SaaS subscriptions, and the agent does work a human would take hours to do.&lt;/p&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Desktop vs TUI Agents&lt;/a&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-4/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;Setting Up Your First Agent &amp;rarr;&lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 2: Desktop vs TUI Agents</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot; class=&quot;accent&quot;&gt;Basic Agents / Lesson 2&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;The Options: Desktop vs TUI Agents&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;A desktop-first tour of the major agent harnesses. We&#39;ll start with Opencode Desktop and Claude Cowork, then explore the TUI options, and end with a clear comparison of when each makes sense.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#the-gui-vs-terminal-split&quot; style=&quot;color: var(--cta);&quot;&gt;The GUI vs Terminal Split&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#desktop-opencode-desktop&quot; style=&quot;color: var(--cta);&quot;&gt;Desktop: Opencode Desktop&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#desktop-claude-cowork&quot; style=&quot;color: var(--cta);&quot;&gt;Desktop: Claude Cowork&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#tui-options&quot; style=&quot;color: var(--cta);&quot;&gt;TUI Options&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#desktop-vs-tui&quot; style=&quot;color: var(--cta);&quot;&gt;Desktop vs TUI: The Key Differences&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#decision-framework&quot; style=&quot;color: var(--cta);&quot;&gt;Decision Framework&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/#your-takeaway&quot; style=&quot;color: var(--cta);&quot;&gt;Your Takeaway&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Know the landscape of agent harnesses and whether you want a desktop app or a terminal tool&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;A computer where you can install software (or at least open a terminal)&lt;/li&gt;
&lt;li&gt;An internet connection&lt;/li&gt;
&lt;li&gt;The Lesson 1 mental model: &lt;em&gt;agent is the brain, harness is the body&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-gui-vs-terminal-split&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The GUI vs Terminal Split&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before we tour the options, there&#39;s a divide that cuts across all of them: &lt;strong&gt;desktop apps&lt;/strong&gt; vs &lt;strong&gt;terminal tools&lt;/strong&gt;.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;A desktop agent harness runs in a window. It has buttons, panels, menus. You download it, install it, and interact with it the way you&#39;d interact with any other application. It shows you what the agent is doing, lets you intervene with clicks, and presents results in a visual interface.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;A TUI (terminal user interface) agent runs in your terminal. It looks like a command-line tool. You type commands, it types back. It&#39;s faster, lighter, and more scriptable than a desktop app, and it shows you the agent&#39;s thoughts, tool calls, and edits right in the terminal output — often more clearly and with less abstraction than a GUI version.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Neither is better. They&#39;re built for different situations and different preferences. Let&#39;s look at the desktop options first, then the TUI ones, and then compare them head-to-head.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;desktop-opencode-desktop&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Desktop: Opencode Desktop&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Opencode Desktop&lt;/strong&gt; is a downloadable application that gives you a full visual workspace for collaborating with AI agents. You install it on your machine, sign in, and get a graphical interface where you can manage multiple agent sessions, monitor their progress, and review their output all in one place.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;What makes Opencode Desktop stand out:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Session manager&lt;/strong&gt; — run multiple agents at once, each in its own tab or panel. You can watch them work in parallel.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Streamlined output&lt;/strong&gt; — focuses on results over process. Less verbose than the TUI, showing you what matters without the full transcript of every internal step.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;File explorer&lt;/strong&gt; — browse and edit files the agent is working on without leaving the app.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Built-in terminal&lt;/strong&gt; — a hybrid option: you get the visual workspace but can drop into a terminal when you need to.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Easy setup&lt;/strong&gt; — download, install, configure your API key, and you&#39;re running in minutes. No package managers, no environment variables, no dependency hell.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;Who it&#39;s for: anyone who prefers a visual interface, runs multi-step agent tasks regularly, or wants to monitor and review agent output in a structured way. It&#39;s particularly good for people managing complex workflows where seeing the whole picture matters.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;desktop-claude-cowork&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Desktop: Claude Cowork&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Claude Cowork&lt;/strong&gt; is Anthropic&#39;s desktop agent experience. It runs inside the Claude desktop app and lets you hand tasks to Claude that it works through autonomously — reading files, writing code, searching the web, and reporting back.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;What makes Claude Cowork different:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Chat-native interface&lt;/strong&gt; — the agent interaction feels like an extension of a conversation. You describe what you need, Claude works on it, and you can chime in mid-task to redirect.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Built on Claude&lt;/strong&gt; — uses Anthropic&#39;s models natively, with no middle layer. The model and the harness are designed together.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool use out of the box&lt;/strong&gt; — file reading, code execution, web search, image analysis are all available without any configuration.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interruptible&lt;/strong&gt; — you can jump in at any point and say &quot;wait, try it this way instead.&quot; Cowork mode is designed for back-and-forth, not fire-and-forget.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Low friction&lt;/strong&gt; — if you already use Claude, there&#39;s nothing extra to install or configure. Cowork mode is a feature of the app you already have.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;Who it&#39;s for: people who already use Claude and want to graduate from one-shot prompts to sustained, multi-step tasks. The cowork dynamic works well when you want to stay involved in the process — directing, reviewing, and adjusting as the agent works.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;tui-options&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;TUI Options&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;TUI agent harnesses run in your terminal. They&#39;re lighter, faster, and more scriptable than desktop apps — but they expect you to be comfortable with a command line.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Opencode CLI&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The terminal version of Opencode. Same agent engine as the desktop app, but accessed entirely through your terminal. You run &lt;code class=&quot;accent&quot;&gt;opencode&lt;/code&gt; in a project directory, and it opens an interactive session where the agent can read, write, and execute commands in your environment.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Key difference from the desktop version: no point-and-click file explorer, no session tabs. What you gain is full visibility into every decision — the TUI streams the agent&#39;s reasoning, tool calls, and file edits directly to your terminal in real time. It starts in under a second, works over SSH, and offers more fine-grained control over every step.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;The CLI and Desktop versions share the same underlying agent. Your choice is about interface preference and how much detail you want to see.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Claude Code CLI&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Anthropic&#39;s terminal-based coding agent. Runs as a command-line tool — you invoke it from within a project, and it works alongside you to write, edit, and debug code. Like Claude Cowork in capability, but without the visual interface.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Where Claude Cowork is interruptible and dialog-driven, Claude Code CLI is more hands-off. You give it a task, it works on it, and reports back. The terminal format makes it easier to integrate into CI/CD pipelines and automated workflows.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;LangChain CLI&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;LangChain&#39;s terminal toolkit for building and running agent chains. More of a development framework than an out-of-the-box agent. You define your agent&#39;s tools, memory, and model routing in configuration files, then invoke it from the command line.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the most flexible option — you can build exactly the agent you want — but it expects you to learn its abstraction layer. It&#39;s not a plug-and-play experience.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Beyond basic TUI:&lt;/strong&gt; some agents go further — they run autonomously in a loop, breaking your goal into sub-tasks and executing them without you watching. Tools like AutoGPT, Hermes, and Openclaw&#39;s fleet mode are built for this. They&#39;re powerful, but they need their own tutorial track. We&#39;ll cover them in a future series on autonomous agent loops.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;desktop-vs-tui&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Desktop vs TUI: The Key Differences&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Once you strip away the branding, every agent harness does roughly the same things: it routes prompts to a model, executes tools, and manages memory. The real difference is how you interact with it.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Desktop&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;TUI&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Setup&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Download and install. GUI configuration.&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Package manager or script. Terminal config.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Interface&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Visual: panels, tabs, buttons, timelines&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Text: typed commands, text responses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Multi-tasking&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Native tab/session management. See multiple agents at once.&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Managed via tmux, screen, or background processes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Speed&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Slightly heavier launch, but comparable once running&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Near-instant launch, minimal overhead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Scriptability&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Limited to what the GUI exposes&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Fully scriptable. Pipe in, pipe out. CI/CD integration.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Remote use&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Requires SSH forwarding or a remote desktop setup&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Works natively over SSH. Run on a server, control from anywhere.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Learning curve&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Lower — familiar GUI patterns&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Moderate — commands are text but the interface is straightforward&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Visibility&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Polished summaries, less verbose output&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Full transcript of thoughts, tool calls, and file edits in real time&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;decision-framework&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Decision Framework: Which One Do You Install?&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You&#39;ve now seen the landscape. Here&#39;s how to choose:&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Choose a desktop agent when:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;You want a polished, less verbose interface that shows you the highlights, not every detail&lt;/li&gt;
&lt;li&gt;You manage multiple agent tasks at once and need a session manager&lt;/li&gt;
&lt;li&gt;You prefer a familiar GUI experience with buttons, panels, and menus&lt;/li&gt;
&lt;li&gt;You&#39;re collaborating with an agent interactively — directing, reviewing, adjusting as it works&lt;/li&gt;
&lt;/ul&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Choose a TUI agent when:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;You want full visibility into everything the agent is thinking and doing, in real time&lt;/li&gt;
&lt;li&gt;You need fine-grained control over each step — you can see the command, the output, the decision&lt;/li&gt;
&lt;li&gt;You need to run agents on a remote server over SSH&lt;/li&gt;
&lt;li&gt;You want to script or automate agent tasks — pipe output between tools, run in CI/CD&lt;/li&gt;
&lt;li&gt;You prefer speed and minimal overhead over visual polish&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;your-takeaway&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Decision Cheat Sheet&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Save this comparison. When someone asks you &quot;what agent harness should I use?&quot; you&#39;ll have the answer ready.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Desktop if:&lt;/strong&gt; you want a polished, less verbose experience with session management and GUI navigation.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;TUI if:&lt;/strong&gt; you want full visibility into every decision, fine-grained control, remote use, or scriptability.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Both if:&lt;/strong&gt; some tasks benefit from visibility (complex multi-step workflows) and others benefit from speed (one-shot automations). There&#39;s no rule that says you can only use one.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; What Is an Agent Harness?&lt;/a&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-3/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;How Much Does an Agent Cost Per Hour? &amp;rarr;&lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 1: What Is an Agent Harness?</title>
    <link href="https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/&quot; class=&quot;accent&quot;&gt;Basic Agents / Lesson 1&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;What Is an Agent Harness?&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;The difference between a chatbot and an agent is the difference between giving advice and doing the work. An agent harness is the body your AI brain needs to actually get things done.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/#youve-felt-the-ceiling&quot; style=&quot;color: var(--cta);&quot;&gt;You&#39;ve Felt the Ceiling&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/#chatbot-vs-agent&quot; style=&quot;color: var(--cta);&quot;&gt;Chatbot vs Agent&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/#what-a-harness-does&quot; style=&quot;color: var(--cta);&quot;&gt;What a Harness Does&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/#the-mental-model&quot; style=&quot;color: var(--cta);&quot;&gt;The Mental Model: Brain and Body&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/#walkthrough-from-chat-to-agent&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: From Chat to Agent&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/#your-takeaway&quot; style=&quot;color: var(--cta);&quot;&gt;Your Takeaway&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Understand what an agent harness is, what it does, and why you need one to move beyond chat&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 12 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Familiarity with using a chatbot like ChatGPT or Claude (from Basic Chatbots track)&lt;/li&gt;
&lt;li&gt;Your preferred chatbot open in a tab&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;youve-felt-the-ceiling&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;You&#39;ve Felt the Ceiling&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;You open a chatbot. You type a prompt. It responds with text. Good output, maybe great. But then you need the chatbot to &lt;em&gt;do&lt;/em&gt; something — read a file, run a script, check a website, update a spreadsheet — and you hit a wall. The chatbot can talk about the task, but it can&#39;t perform it. You&#39;re back to copying, pasting, and doing the work yourself.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;That wall is the ceiling of a chat-only interface. The model is smart enough to help, but it has no hands. It can&#39;t reach into your file system, execute code, call an API, or persist information between conversations.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;An agent harness removes that ceiling. It gives the model the tools it needs to act on the world, not just talk about it.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-agents-actually-do-for-you&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What Agents Actually Do for You&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before we get into definitions and tables, let&#39;s start with what an agent can do that a chatbot can&#39;t. These are real examples of meaningful work, not technical demos:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 0; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Organize your files&lt;/strong&gt; — &quot;Go through my Downloads folder, sort everything into subfolders by file type, and delete anything that&#39;s been there over a year.&quot; The agent looks at every file, decides where it goes, moves it, and cleans up. You don&#39;t touch a single file.&lt;/li&gt;
&lt;li style=&quot;margin-top: 1rem;&quot;&gt;&lt;strong&gt;Research and compile&lt;/strong&gt; — &quot;Find the top 5 competitors in my area, check their websites for pricing, and build me a comparison spreadsheet.&quot; The agent searches, reads each site, extracts the numbers, and builds the file. You come back to a finished document.&lt;/li&gt;
&lt;li style=&quot;margin-top: 1rem;&quot;&gt;&lt;strong&gt;Monitor and alert&lt;/strong&gt; — &quot;Check this webpage every morning. If the price drops below $50, send me an email.&quot; The agent visits the page daily, reads the price, and acts when conditions are met. You don&#39;t check — it checks.&lt;/li&gt;
&lt;li style=&quot;margin-top: 1rem;&quot;&gt;&lt;strong&gt;Batch process&lt;/strong&gt; — &quot;Take all these photos, resize them to 1200px wide, rename them with today&#39;s date, and put them in a folder called &#39;Processed&#39;.&quot; The agent opens each image, edits it, names it, and files it. You go make tea.&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The pattern in every case is the same: &lt;strong&gt;you describe what you want done, and the agent does the doing.&lt;/strong&gt; Not the planning, not the advice — the actual work of reading, writing, moving, comparing, deciding, and reporting back. That&#39;s the shift from chatbot to agent, and it&#39;s what this whole track is about.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;chatbot-vs-agent&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Chatbot vs Agent&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;These two terms get used interchangeably, but they describe fundamentally different things:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Chatbot&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Agent&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Input&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Your typed prompt&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;A goal, plus data from tools it can call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Output&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Text response&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Actions: files changed, code run, APIs called, data retrieved&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Memory&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Current conversation only&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Persistent across sessions, can store and retrieve&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Tools&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;None (or limited, like web search toggle)&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;File system, code execution, APIs, databases, search&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Autonomy&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;None — you guide every step&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Can execute multi-step plans without your input&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;A chatbot is a conversation. An agent is a delegation. You tell the agent what you want, and it figures out how to make it happen — reading, writing, running, checking, iterating — until the goal is met or it needs your input to proceed.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-a-harness-does&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What a Harness Does&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;An agent harness is the software that turns a language model into an agent. It provides the infrastructure the model needs to do work. Every harness handles four core functions:&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;1. Model Routing&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The harness connects to one or more language models. When you give the agent a task, the harness sends the prompt to the model, receives the response, and decides what to do next. Some harnesses let you route different types of tasks to different models — use a cheap fast model for simple lookups, a powerful expensive one for complex reasoning.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;2. Tool Execution&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is the core of what makes an agent an agent. The harness provides tools the model can call: read a file, write a file, run a command, search the web, query a database, call an API. When the model decides it needs to use a tool, the harness executes it and feeds the result back to the model. The model can then decide what to do based on what the tool returned.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;3. Memory&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Chatbots forget everything when you close the tab. Many agent harnesses are the same — when you end a session and start a new one, the agent starts fresh with no memory of what happened before. The harness &lt;em&gt;can&lt;/em&gt; give the model a place to store information it wants to remember — notes about your project, results from previous steps, preferences you&#39;ve expressed — but it won&#39;t do this automatically. You need to specifically teach the agent what to save and how to save it, or it will lose everything between sessions like a chatbot does. We&#39;ll cover this in detail in a later lesson.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;4. Sandboxing&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Giving an AI the ability to run code and modify files is powerful — and risky. A harness runs the agent in a controlled environment. It restricts what the agent can access, caps resource usage, and logs every action. If the agent goes off course, the sandbox contains the damage.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-mental-model&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Mental Model: Brain and Body&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s the simplest way to think about it:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Component&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Analogy&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;The Model&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;The brain&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;GPT-4, Claude, Gemini — the thing that reasons&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;The Harness&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;The body&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Opencode Desktop, Claude Cowork, LangChain — the thing that acts&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The brain is smart but helpless without a body. It can think through any problem but can&#39;t pick up a pencil. The body gives the brain hands — the ability to read, write, execute, and sense the world. You need both for an agent to do useful work.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Throughout this track, when we talk about &quot;running an agent,&quot; what we really mean is &quot;running a harness with a model plugged in.&quot; The harness is the product you install. The model is the service you subscribe to. They work together, and you choose both.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-from-chat-to-agent&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: From Chat to Agent&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s make this concrete. Here&#39;s the same task approached first as a chat, then as an agent delegation.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;The Chat Way&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Open ChatGPT&lt;/strong&gt; and type this prompt:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;I need to organize my downloads folder. It&#39;s a mess. Can you suggest how to sort files into subfolders by type?&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;ChatGPT will give you great advice. It might suggest folders for images, documents, audio, and archives. It might even write a Python script you could run. But it won&#39;t actually organize anything. You&#39;re still the one who has to open the terminal, run the script, or drag files around. The model gave you a plan; you execute it.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem;&quot;&gt;The Agent Way&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Open your agent harness.&lt;/strong&gt; Type something like:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Organize my downloads folder. Sort files into subfolders by type: images, documents, audio, archives, and other. Log what you moved where.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent harness reads the folder, identifies each file&#39;s type, creates subfolders, moves the files, and writes a log of what happened. The model figured out the plan, the harness executed it. You gave the goal, the agent did the work.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Key insight:&lt;/strong&gt; The model in both cases is equally capable. What changed is the harness. In the chat scenario, the model had words-only output. In the agent scenario, the harness gave the model access to the file system, command execution, and logging tools. Same brain. Different body.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;your-takeaway&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Your Takeaway — The Distinction Test&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before moving on, make sure you can answer this question: &lt;strong&gt;If I give this task to a chatbot, does it talk about the work or does it do the work?&lt;/strong&gt; If the answer is &quot;talks about it,&quot; you need an agent harness.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Keep this test in your pocket. It&#39;s the distinction that runs through every lesson in this track.&lt;/p&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-2/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;The Options: Desktop vs TUI Agents &amp;rarr;&lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 9: Building Your First Loop</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 9&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Building Your First Loop&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Theory is useful. A running agent is better. Let&#39;s build one end to end.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#the-five-layers&quot; style=&quot;color: var(--cta);&quot;&gt;The Five Layers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-1-choose-a-task&quot; style=&quot;color: var(--cta);&quot;&gt;Step 1: Choose a Task&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-2-set-up-the-sandbox&quot; style=&quot;color: var(--cta);&quot;&gt;Step 2: Set Up the Sandbox&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-3-install-and-configure-hermes&quot; style=&quot;color: var(--cta);&quot;&gt;Step 3: Install and Configure Hermes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-4-define-the-tools&quot; style=&quot;color: var(--cta);&quot;&gt;Step 4: Define the Tools&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-5-write-the-goal&quot; style=&quot;color: var(--cta);&quot;&gt;Step 5: Write the Goal&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-6-set-budget-and-limits&quot; style=&quot;color: var(--cta);&quot;&gt;Step 6: Set Budget and Limits&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-7-run-it&quot; style=&quot;color: var(--cta);&quot;&gt;Step 7: Run It&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#step-8-review-the-log&quot; style=&quot;color: var(--cta);&quot;&gt;Step 8: Review the Log&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#compare-first-loop-vs-ideal-pattern&quot; style=&quot;color: var(--cta);&quot;&gt;Compare: First Loop vs. the Ideal Pattern&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#deliverable-proof-of-completion&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Proof of Completion&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Build, run, and inspect a complete autonomous agent loop that accomplishes a real file organization task&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 45 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Docker Desktop installed and running&lt;/li&gt;
&lt;li&gt;Python 3.10+ installed&lt;/li&gt;
&lt;li&gt;A terminal and a text editor&lt;/li&gt;
&lt;li&gt;An OpenRouter API key (free tier works for this — about $0.05 per run)&lt;/li&gt;
&lt;li&gt;All five patterns from Lesson 8 memorized&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-five-layers&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Five Layers&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every autonomous agent loop has five layers. We&#39;ll build each one in order:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;ol style=&quot;color: var(--text-secondary); line-height: 2.4; margin-bottom: 0; padding-left: 1.5rem;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Sandbox&lt;/strong&gt; — Where does the agent run? A Docker container with restricted access.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Harness&lt;/strong&gt; — What runtime controls the loop? Hermes, in our case.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model&lt;/strong&gt; — What drives the decisions? DeepSeek V4 Flash Max through OpenRouter.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tools&lt;/strong&gt; — What can the agent do? Read, write, list, and organize files.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Monitoring&lt;/strong&gt; — How do we watch it? Console logging and a timestamped log file.&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Each layer is a decision we make. The decisions build on each other — the sandbox determines what tools are safe, the harness determines what models we can use, the model determines the cost structure, and the monitoring determines whether we can debug failures. Let&#39;s work through them one at a time.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-1-choose-a-task&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 1: Choose a Task&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;We need something real but contained. Something that exercises the loop without risking the things we care about. The classic first task: a file organizer.&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Task: Organize a messy download folder.

Take a directory of mixed files — PDFs, images, documents,
spreadsheets, archives — and sort them into subdirectories
by type. Create the subdirectories if they don&#39;t exist.
Move (don&#39;t copy) each file into the matching folder.&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Why this task? It&#39;s concrete — success is easy to verify (check that files moved). It&#39;s safe — we can test on a throwaway directory. It exercises the core loop actions: list files, classify each one, create directories, move files. And it produces an obvious log trail we can inspect afterward.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Create a test directory&lt;/strong&gt; with a dozen mixed files for the agent to organize:&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;mkdir agent-test-files
cd agent-test-files
echo &quot;report content&quot; &gt; q1-report.pdf
echo &quot;image data&quot; &gt; photo.jpg
echo &quot;image data&quot; &gt; screenshot.png
echo &quot;budget&quot; &gt; expenses.xlsx
echo &quot;meeting notes&quot; &gt; notes.docx
echo &quot;archive&quot; &gt; old-projects.zip
echo &quot;presentation&quot; &gt; slides.pptx
echo &quot;invoice&quot; &gt; invoice-q2.pdf
echo &quot;spreadsheet&quot; &gt; data.csv
echo &quot;image data&quot; &gt; banner.gif
echo &quot;backup&quot; &gt; backup.zip
echo &quot;document&quot; &gt; contract.docx&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;We now have 12 files of various types in a directory. This is what the agent will organize.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-2-set-up-the-sandbox&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 2: Set Up the Sandbox&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent will have file system access — it will create directories and move files. That&#39;s destructive. We sandbox it so it can&#39;t touch anything outside the test directory.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Create a Docker container&lt;/strong&gt; with the test directory mounted read-write and everything else read-only:&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker run -d --name agent-sandbox `
  -v ${PWD}/agent-test-files:/workspace/data:rw `
  -v ${PWD}/agent-output:/workspace/output:rw `
  -w /workspace/data `
  python:3.11-slim sleep infinity&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;This container has no network access (add &lt;code class=&quot;accent&quot;&gt;--network none&lt;/code&gt; for extra isolation), can only write to the data and output directories, and everything else in the file system is read-only. If the agent tries to delete system files or write to unexpected locations, it simply fails — no harm done.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;We&#39;ll install the agent harness inside this container on the next step.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-3-install-and-configure-hermes&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 3: Install and Configure Hermes&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes is our harness of choice for this build: it&#39;s lightweight, supports cost caps natively, and runs inside a container without fuss.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Install Hermes inside the container:&lt;/strong&gt;&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker exec agent-sandbox pip install hermes-agent&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Set up the config:&lt;/strong&gt;&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker exec agent-sandbox mkdir -p /workspace/config

# Write the config file
docker exec -i agent-sandbox tee /workspace/config/hermes.yaml &gt; $null
@&quot;
model:
  provider: openrouter
  name: deepseek/deepseek-chat
  temperature: 0.3
  max_tokens: 4000

workspace:
  path: /workspace
  allowed_tools:
    - bash
    - read_file
    - write_file
    - list_directory
&quot;@&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;We chose DeepSeek V4 Flash Max through OpenRouter because it&#39;s cheap (~$0.002 per call), fast, and good enough for file organization. Temperature 0.3 keeps decisions deterministic — we want reliable classification, not creative file naming. The model runs inside the container, the container has no internet access beyond the OpenRouter API, and the workspace is constrained to &lt;code class=&quot;accent&quot;&gt;/workspace&lt;/code&gt;.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-4-define-the-tools&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 4: Define the Tools&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes comes with built-in tools for file operations. We don&#39;t need to write custom tool code for this task — the existing &lt;code class=&quot;accent&quot;&gt;bash&lt;/code&gt; tool combined with shell commands is sufficient.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;But we do need to control which tools are available. The task needs:&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Define allowed tools in the goal file
tools:
  - name: list_directory
    description: &quot;List all files in a directory&quot;
    command: &quot;ls -la {path}&quot;
  - name: read_file
    description: &quot;Read contents of a file&quot;
    command: &quot;cat {path}&quot;
  - name: run_shell
    description: &quot;Execute a shell command&quot;
    command: &quot;{command}&quot;
  - name: move_file
    description: &quot;Move a file from source to destination&quot;
    command: &quot;mv {source} {destination}&quot;
  - name: create_directory
    description: &quot;Create a directory if it doesn&#39;t exist&quot;
    command: &quot;mkdir -p {path}&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Important constraint: we &lt;strong&gt;do not&lt;/strong&gt; give the agent &lt;code class=&quot;accent&quot;&gt;rm&lt;/code&gt; or &lt;code class=&quot;accent&quot;&gt;delete&lt;/code&gt; access. If the agent tries to clean up by deleting files, it will fail — and that&#39;s intentional. The agent should move files, not delete them. If it needs to handle duplicates, it should move them to a &lt;code class=&quot;accent&quot;&gt;duplicates/&lt;/code&gt; subfolder instead.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-5-write-the-goal&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 5: Write the Goal&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The goal is the most important part. This is where all the patterns from Lesson 8 come together. A well-formed goal has: the task, concrete stopping criteria, constraints, and escalation rules.&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# goal.yaml
goal:
  description: &gt;
    Organize the files in /workspace/data into subdirectories
    by file type. Create a directory for each type if it doesn&#39;t
    exist. Move each file into the appropriate directory.
  success_criteria:
    - &quot;Every file in /workspace/data has been moved to a subdirectory&quot;
    - &quot;Subdirectories exist named: pdfs, images, documents,
       spreadsheets, archives&quot;
    - &quot;No files remain in /workspace/data root&quot;
    - &quot;No files were deleted — only moved&quot;

constraints:
  - &quot;Do not use rm, delete, or remove commands&quot;
  - &quot;If a file extension doesn&#39;t match any category, move it to &#39;other/&#39;&quot;
  - &quot;Create the named subdirectories before moving files&quot;
  - &quot;List the final directory structure when complete&quot;

file_types:
  pdf: [&quot;.pdf&quot;]
  images: [&quot;.jpg&quot;, &quot;.jpeg&quot;, &quot;.png&quot;, &quot;.gif&quot;, &quot;.bmp&quot;, &quot;.svg&quot;]
  documents: [&quot;.doc&quot;, &quot;.docx&quot;, &quot;.txt&quot;, &quot;.md&quot;, &quot;.rtf&quot;]
  spreadsheets: [&quot;.xls&quot;, &quot;.xlsx&quot;, &quot;.csv&quot;]
  archives: [&quot;.zip&quot;, &quot;.tar&quot;, &quot;.gz&quot;, &quot;.rar&quot;]

escalation:
  on_error: &quot;Log the error, skip the file, continue&quot;
  on_ambiguous_extension: &quot;Move to &#39;other/&#39; and log the choice&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;We wrote the success criteria as concrete, testable statements — not &quot;organize the files&quot; but &quot;no files remain in the root directory&quot; and &quot;subdirectories exist with these names.&quot; The agent can check its own work against these criteria and know when it&#39;s done.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-6-set-budget-and-limits&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 6: Set Budget and Limits&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before running, we apply the remaining reliability patterns: timeouts, cost budget, and logging.&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;limits:
  max_iterations: 20
  max_cost: 0.50
  max_time_per_step: 60
  max_total_time: 600

logging:
  level: verbose
  output: /workspace/output/agent-run-{timestamp}.log&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The $0.50 cap is generous for this task — it should cost about $0.05. But the cap gives us confidence that even if the agent goes sideways, the damage is limited. If it hits the cap, we inspect the log, fix the goal, and re-run with a lesson learned.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-7-run-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 7: Run It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Copy the goal file into the container&lt;/strong&gt; and start Hermes:&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Copy the goal file into the container
docker cp goal.yaml agent-sandbox:/workspace/config/goal.yaml

# Run Hermes with the goal
docker exec -e OPENROUTER_API_KEY=$env:OPENROUTER_API_KEY `
  agent-sandbox hermes run /workspace/config/goal.yaml `
  --config /workspace/config/hermes.yaml `
  --log-file /workspace/output/agent-run.log&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Watch the output as it runs. Hermes prints each step as it executes:&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;[STEP 1/12] list_directory /workspace/data
  Found 12 files

[STEP 2/12] create_directory /workspace/data/pdfs
  Directory created

[STEP 3/12] create_directory /workspace/data/images
  Directory created

[STEP 4/12] create_directory /workspace/data/documents
  Directory created

... (continues through all categories)

[STEP 8/12] move_file /workspace/data/q1-report.pdf /workspace/data/pdfs/
  File moved

[STEP 9/12] move_file /workspace/data/invoice-q2.pdf /workspace/data/pdfs/
  File moved

...

[STEP 12/12] list_directory /workspace/data
  Subdirectories: pdfs/, images/, documents/, spreadsheets/,
  archives/ — root is empty. All files organized.

[DONE] Goal achieved. Total cost: $0.034. Iterations: 12.&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The agent finished in 12 iterations and spent 3.4 cents. All 12 files were sorted into the correct subdirectories. The log captured every action, every tool call, and the final directory listing as proof.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-8-review-the-log&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 8: Review the Log&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Read the log file&lt;/strong&gt; to verify everything happened as expected:&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Copy the log out of the container
docker cp agent-sandbox:/workspace/output/agent-run.log .

# View the log
cat agent-run.log&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;What we&#39;re looking for in the log:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Every tool call is recorded&lt;/strong&gt; — with arguments, return values, and duration&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The decision process is visible&lt;/strong&gt; — why the agent chose to create directories before moving files, why it classified each file into its category&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No unexpected tool calls&lt;/strong&gt; — the agent didn&#39;t try to delete files or access anything outside &lt;code class=&quot;accent&quot;&gt;/workspace/data&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The total cost is reported&lt;/strong&gt; — and it matches the estimated ~$0.05 range&lt;/li&gt;
&lt;/ul&gt;

&lt;p class=&quot;mb-4&quot;&gt;If the log reveals errors — files classified wrong, directories created in the wrong location, a step that timed out — that&#39;s not a failure. That&#39;s data. The log tells us exactly what went wrong, and we can fix the goal or constraints and re-run. This is the scientific method applied to agent building: hypothesize, run, inspect, adjust.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;compare-first-loop-vs-ideal-pattern&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Compare: First Loop vs. the Ideal Pattern&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now compare what we built against the five-pattern reliability checklist from Lesson 8:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Pattern&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Our Loop&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Production Ideal&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Human-in-the-loop&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Not implemented — the agent runs fully autonomously&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Add a checkpoint before the first file move — &quot;show me the classification plan and wait for approval&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Logging&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Full verbose logging to file&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Add structured JSON logging for automated analysis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Timeouts&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Per-step and per-loop timeouts configured&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Same — already production-ready&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Cost budget&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;$0.50 hard cap configured&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Same — already production-ready&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Escalation path&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Error logging + skip configured&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Add retry logic and a dead-letter queue for consistently failing files&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;We got four out of five patterns right on the first build. The missing one — human-in-the-loop — is the hardest to add in Hermes because the harness doesn&#39;t natively support pause-and-resume. For a file organizer, full autonomy is acceptable (files can be restored from backup). For higher-stakes tasks, we&#39;d switch to Openclaw Fleet for its built-in approval gate support.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-proof-of-completion&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Proof of Completion&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;When the agent finishes, we have three artifacts to show it worked:&lt;/p&gt;

&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;The organized directory&lt;/strong&gt; — &lt;code class=&quot;accent&quot;&gt;agent-test-files/&lt;/code&gt; now has subdirectories with files sorted by type. Run &lt;code class=&quot;accent&quot;&gt;ls -R agent-test-files/&lt;/code&gt; to see the structure.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The log file&lt;/strong&gt; — &lt;code class=&quot;accent&quot;&gt;agent-run.log&lt;/code&gt; shows every step the agent took, every tool call, and every decision. This is the audit trail.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The cost receipt&lt;/strong&gt; — the final output shows total cost ($0.034 in our case). This is the first data point for estimating future runs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p class=&quot;mb-4&quot;&gt;These three artifacts — organized files, log, and cost — are the minimum deliverable for any autonomous agent task. If you can produce these three things, you have a loop that works, verified by evidence, and priced so you know what to expect next time.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Run the build yourself.&lt;/strong&gt; Follow steps 1 through 8. Use the exact goal file and config shown above. When it works, try these variations:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Add a checkpoint.&lt;/strong&gt; Modify the goal to: &quot;List all files, present the classification plan, then wait for user input &#39;proceed&#39; before moving anything.&quot; Simulate the checkpoint by running with &lt;code class=&quot;accent&quot;&gt;--interactive&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Introduce an ambiguous file.&lt;/strong&gt; Add a file with no extension to the test directory. Does the agent handle it (move to &lt;code class=&quot;accent&quot;&gt;other/&lt;/code&gt; as instructed) or does it stumble?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Set the budget to $0.05 — intentionally too low.&lt;/strong&gt; Watch what happens when the agent hits the cost cap mid-task. Inspect the partial work. Learn the failure mode in a safe environment.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;The goal isn&#39;t just to get it working once. It&#39;s to understand every failure mode so that when you build a loop for something that matters, you already know what can go wrong and how to prevent it.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Effective Patterns&lt;/a&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;When NOT to Use an Autonomous Agent &amp;rarr;&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/&quot; style=&quot;color: var(--cta);&quot;&gt;When NOT to Use an Autonomous Agent&lt;/a&gt; — the most important lesson is knowing when not to use one.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 8: Effective Patterns</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 8&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Effective Patterns&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Some loops run clean and finish fast. Others spin forever and burn money. The difference is a handful of patterns applied consistently.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#the-five-patterns&quot; style=&quot;color: var(--cta);&quot;&gt;The Five Patterns&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#pattern-1-human-in-the-loop&quot; style=&quot;color: var(--cta);&quot;&gt;Pattern 1: Human-in-the-Loop Checkpoints&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#pattern-2-logging-and-observability&quot; style=&quot;color: var(--cta);&quot;&gt;Pattern 2: Logging and Observability&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#pattern-3-timeout-strategies&quot; style=&quot;color: var(--cta);&quot;&gt;Pattern 3: Timeout Strategies&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#pattern-4-cost-budgets&quot; style=&quot;color: var(--cta);&quot;&gt;Pattern 4: Cost Budgets&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#pattern-5-escalation-paths&quot; style=&quot;color: var(--cta);&quot;&gt;Pattern 5: Escalation Paths&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#walkthrough-before-and-after&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Before and After&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#which-patterns-matter-for-which-tool&quot; style=&quot;color: var(--cta);&quot;&gt;Which Patterns Matter for Which Tool&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#deliverable-reliability-checklist&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Reliability Checklist&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Know the five reliability patterns for autonomous agents and be able to apply each one to any task&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 20 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;A goal file or agent config from any of the previous lessons&lt;/li&gt;
&lt;li&gt;Hermes, Openclaw Fleet, or AutoGPT installed (any of the three)&lt;/li&gt;
&lt;li&gt;A text editor to apply the patterns&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-five-patterns&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Five Patterns&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;We&#39;ve run agents that worked beautifully and agents that quietly burned $12 before we noticed. The ones that worked had these five patterns in common. They&#39;re not complicated — most are a single line of configuration — but they make the difference between a reliable tool and an expensive toy.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;ol style=&quot;color: var(--text-secondary); line-height: 2.4; margin-bottom: 0; padding-left: 1.5rem;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Human-in-the-loop checkpoints&lt;/strong&gt; — Pause at key decisions and ask for approval&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Logging and observability&lt;/strong&gt; — Know what the agent did and why&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Timeout strategies&lt;/strong&gt; — Max runtime per step and per loop&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost budgets&lt;/strong&gt; — Hard caps per task&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Escalation paths&lt;/strong&gt; — What to do when the agent gets stuck&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Each pattern addresses a specific failure mode from Lesson 7. Checkpoints prevent unintended tool calls. Logging catches data exposure. Timeouts stop infinite loops. Cost budgets prevent cost blowouts. Escalation paths handle everything else. Together, they turn a fragile experiment into a robust system.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;pattern-1-human-in-the-loop&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Pattern 1: Human-in-the-Loop Checkpoints&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Idea&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before the agent takes a consequential action — deleting files, sending emails, spending money — it pauses and asks for approval. The agent presents what it&#39;s about to do and why. We review. We approve or reject. The loop continues.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Why It Works&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The model is smart but it lacks context about our specific situation. It doesn&#39;t know that &quot;delete everything in the temp directory&quot; is dangerous if someone moved important files there yesterday. It doesn&#39;t know which email recipients should get which message. A checkpoint gives us a chance to apply human judgment at the moment it matters most.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Implement It&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Openclaw Fleet: mark actions that need approval
tasks:
  - action: delete_files
    approval_required: true
  - action: send_email
    approval_required: true&lt;/pre&gt;&lt;/div&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Hermes: add a structured goal step that pauses
tasks:
  - task: &quot;List files to delete and present to user&quot;
    tool: read_directory
    output: deletion_candidates
  - task: &quot;WAIT_FOR_USER_APPROVAL&quot;
    context: deletion_candidates
  - task: &quot;Delete approved files&quot;
    tool: delete_files
    input: approved_list&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;For AutoGPT, there&#39;s no built-in checkpoint system, but we can simulate one by adding a constraint: &quot;Before any file deletion, print the full list and stop. Wait for the user to type &#39;proceed&#39; before continuing.&quot; The agent will generate the list and then stall — that&#39;s our cue to review and manually restart with a modified goal that confirms the list.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;pattern-2-logging-and-observability&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Pattern 2: Logging and Observability&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Idea&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every action the agent takes is recorded with a timestamp, the tool used, the input, and the output. When something goes wrong, we read the log and understand exactly why.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Why It Works&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Without logs, an agent failure is a black box. The agent returned a wrong answer — but did it read the wrong file? Did it misinterpret the instructions? Did it hallucinate data? Logs turn &quot;it broke&quot; into &quot;it read the wrong config file at step 3 because the path was misspelled.&quot; That&#39;s the difference between guessing and diagnosing.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;What to Log&lt;/h3&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Every API call&lt;/strong&gt; — model, prompt tokens, completion tokens, cost&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Every tool invocation&lt;/strong&gt; — which tool, arguments, return value, duration&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Every decision point&lt;/strong&gt; — what the agent considered and what it chose&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Errors and retries&lt;/strong&gt; — what failed and how the agent recovered&lt;/li&gt;
&lt;/ul&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# A good log entry captures the full story
[2026-07-04 14:32:01] TOOL: read_file
    PATH: ./data/clients.csv
    RESULT: 1,247 rows read
    DURATION: 0.3s
    COST: $0.001

[2026-07-04 14:32:05] DECISION: analyze
    CONSIDERED: [&quot;summarize&quot;, &quot;analyze&quot;, &quot;forward&quot;]
    CHOSEN: &quot;analyze&quot;
    REASON: &quot;Need to extract trends before summarizing&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Implement It&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Openclaw Fleet has centralized audit logging built in — every agent action is recorded to a JSONL audit trail. Hermes logs to the console by default; redirect output to a file with &lt;code class=&quot;accent&quot;&gt;hermes run goal.yaml --log-file agent-run.log&lt;/code&gt;. For AutoGPT, enable debug logging with &lt;code class=&quot;accent&quot;&gt;--debug&lt;/code&gt; and pipe to a file.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;pattern-3-timeout-strategies&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Pattern 3: Timeout Strategies&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Idea&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Two timeouts: a short one for each individual step (if a tool call takes more than 60 seconds, abort it) and a long one for the entire loop (if the agent hasn&#39;t finished in 10 minutes, kill it).&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Why It Works&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;A single slow API call or stuck tool can hold the entire loop hostage. Without a per-step timeout, the agent sits there waiting, burning time and money on nothing productive. Without a per-loop timeout, the agent might keep generating new work forever — each iteration convincing itself it needs &quot;just one more search&quot; to be thorough.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Implement It&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Hermes: per-step and per-loop timeouts
task:
  max_iterations: 25
  max_time_per_step: 120  # seconds
  max_total_time: 600     # seconds (10 minutes)&lt;/pre&gt;&lt;/div&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Openclaw Fleet: timeout per task item
fleet:
  agent_timeout: 300  # seconds per agent task
  step_timeout: 60    # seconds per tool call&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;AutoGPT doesn&#39;t support per-step timeouts natively. The workaround: run it inside a wrapper script that kills the process if it runs longer than the limit. A simple PowerShell one-liner: &lt;code class=&quot;accent&quot;&gt;Start-Process -FilePath &quot;autogpt&quot; -ArgumentList &quot;--goal goal.txt&quot; -PassThru | Wait-Process -Timeout 600; if (-not $?) { Stop-Process -Name autogpt }&lt;/code&gt;.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;pattern-4-cost-budgets&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Pattern 4: Cost Budgets&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Idea&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Set a hard dollar limit before the agent starts. When the total cost of API calls hits that limit, the agent stops — even if it&#39;s mid-task.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Why It Works&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Cost is the ultimate feedback signal. If an agent spends $2 on a task that should cost $0.20, something is wrong — the goal is too vague, the tools are too expensive for the task, or the agent is stuck in a loop. A cost cap makes this failure visible immediately instead of discovering it on the credit card bill.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Set the Right Budget&lt;/h3&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Estimate the minimum.&lt;/strong&gt; If you were doing the task manually, how many API calls would it take? A file organizer might take 3-5 calls. Multiply by the per-call cost (~$0.002 for DeepSeek V4 Flash). That&#39;s your floor.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Add a buffer.&lt;/strong&gt; Multiply the floor by 3-5x. Agents are not as efficient as humans at planning their API usage. They&#39;ll zig when we&#39;d zag. The buffer covers their zigging.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Double-check before running.&lt;/strong&gt; If the estimated budget feels high for the task, reconsider the task — not the budget. A $10 task should produce $10 worth of value.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Hermes cost budget
task:
  max_cost: 0.50  # dollars

# Openclaw Fleet per-agent budget
agent:
  budget: 1.00    # dollars per agent session&lt;/pre&gt;&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;pattern-5-escalation-paths&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Pattern 5: Escalation Paths&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Idea&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Tell the agent what to do when it gets stuck. Should it retry? Should it ask for help? Should it move to the next task and flag the failure?&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Why It Works&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Without an escalation path, a stuck agent does one of two things: retry forever (burning money) or silently move on (producing incomplete work). Both are bad. An explicit escalation path turns &quot;I&#39;m stuck&quot; into &quot;I tried X, it failed because Y, so I moved to the next item and logged the error.&quot; That&#39;s actionable.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Implement It&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# In the goal itself, define escalation rules
escalation:
  on_error:
    - &quot;Retry up to 2 times with exponential backoff&quot;
    - &quot;If still failing, log the error and skip to next task&quot;
    - &quot;Flag the task as &#39;needs_review&#39; in the output&quot;
  on_ambiguous:
    - &quot;Log what&#39;s ambiguous&quot;
    - &quot;Use the most conservative interpretation&quot;
    - &quot;Flag for human review&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Openclaw Fleet handles this natively through its task queue — a failed task stays in the queue with an error state for human review. The agent moves on to the next item. In Hermes and AutoGPT, escalation paths are expressed as constraints in the goal: &quot;If you cannot complete a step after 3 attempts, write &#39;FAILED: [reason]&#39; to the log and continue with the next step.&quot;&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-before-and-after&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: Before and After&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s take a fragile autonomous task and apply all five patterns. The task: &quot;Research 5 competitors.&quot; Simple enough — but without guardrails, this is where agents go to die.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Before: The Fragile Version&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;goal: &quot;Research 5 competitors and write a summary.&quot;
# No limits. No checkpoints. No logging. No escalation.
# The agent could run forever, spend unlimited money,
# and we&#39;d never know what happened.&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;This agent starts by searching for competitors. It finds one, reads their website, then decides to research their funding history. That leads to reading Crunchbase, which mentions another competitor. Now it&#39;s researching 8 companies instead of 5. Fifty iterations later, it&#39;s still going, the summary is 20 pages long, and we&#39;ve spent $4. No logs, so we have no idea how it got there.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;After: The Reliable Version&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;goal: &quot;Research 5 competitors and write a one-page summary.&quot;

constraints:
  - &quot;Research exactly 5 competitors, no more&quot;
  - &quot;Each competitor gets at most 3 sources&quot;
  - &quot;One page maximum for the summary&quot;

checkpoints:
  - &quot;Before moving to competitor 4, show me what you have&quot;
  - &quot;If any source costs money, stop and ask&quot;

limits:
  max_iterations: 20
  max_cost: 1.00
  max_time_per_step: 60
  max_total_time: 600

logging:
  level: verbose
  output: competitor-research-{timestamp}.log

escalation:
  on_source_unavailable: &quot;Skip that source, note why, continue&quot;
  on_all_sources_fail: &quot;Log the competitor as &#39;unresearched&#39; and move on&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;This version has everything. A checkpoint at step 4 lets us review progress mid-task. A $1 cost cap means it can&#39;t burn more than a coffee. The 10-minute timeout kills it if it gets stuck. Every action is logged to a timestamped file. And if a website is down or a source is paywalled, the agent skips it gracefully instead of spinning.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;The result: the agent finishes in 12 iterations, spending $0.34. The log shows exactly which sources it used and which it skipped. We reviewed the checkpoint at competitor 4 and decided the first three were strong enough — the agent wrapped up in 2 more iterations and delivered a clean one-page summary.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;which-patterns-matter-for-which-tool&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Which Patterns Matter for Which Tool&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Not every pattern is equally important for every tool. Each platform has its own natural failure modes, and the patterns we apply should target the tool&#39;s weakest point.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Tool&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Priority Pattern&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Hermes&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Iteration limits&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Hermes runs a task list sequentially without built-in loop detection. Without a hard iteration cap, a stuck task repeats forever. Set &lt;code class=&quot;accent&quot;&gt;max_iterations&lt;/code&gt; first.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Openclaw Fleet&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Checkpoint design&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Fleet already has cost caps, logging, and timeouts built in. The biggest risk is an approval gate that&#39;s too loose or too tight. Design checkpoints carefully — not everything needs approval, but the things that do need unambiguous gate logic.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;AutoGPT&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Goal constraints&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;AutoGPT&#39;s goal decomposition is its superpower and its biggest risk. Without tight constraints, it generates sub-goals forever. Every goal needs concrete stopping criteria, a maximum sub-goal depth, and a budget for both iterations and cost.&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-reliability-checklist&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Reliability Checklist&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Copy this checklist and apply it to every autonomous task before you run it. Each item maps to one of the five patterns.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Human-in-the-loop.&lt;/strong&gt; Are there any irreversible actions (delete, overwrite, send, deploy) that need approval before execution? Define the checkpoint.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Logging.&lt;/strong&gt; Is logging enabled at verbose level? Where will the log be written? Will you be able to replay what the agent did?&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Timeouts.&lt;/strong&gt; What&#39;s the max time per step? What&#39;s the max total time? Set both.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Cost budget.&lt;/strong&gt; What&#39;s the hard dollar limit? Set it before running. If the task exceeds the budget, the agent stops — that&#39;s a signal to review and refine the goal, not to increase the budget.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Escalation path.&lt;/strong&gt; What should the agent do when it gets stuck? Retry? Skip? Ask for help? Write it into the goal.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Take one task you&#39;ve run before&lt;/strong&gt; — any task from Lessons 3-7. Open the reliability checklist and score your original configuration: how many of the five patterns did you have? For each missing pattern, add it now. Then re-run the task with all five patterns in place.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Compare the two runs. The version with all five patterns should finish faster, cost less, produce a clearer log, and leave you more confident in the result. If it doesn&#39;t, something is misconfigured — the patterns work together, and a missing one can undermine the rest.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Safety Cautions&lt;/a&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;Building Your First Loop &amp;rarr;&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/&quot; style=&quot;color: var(--cta);&quot;&gt;Building Your First Loop&lt;/a&gt; — combine the theory, tools, risks, and patterns into a complete autonomous agent loop from scratch.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 7: Safety Cautions</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 7&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Safety Cautions&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;An autonomous agent with file access, internet access, and the ability to spend money. What could go wrong? A lot, actually.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#the-five-risk-categories&quot; style=&quot;color: var(--cta);&quot;&gt;The Five Risk Categories&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#risk-1-token-burns&quot; style=&quot;color: var(--cta);&quot;&gt;Risk 1: Token Burns&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#risk-2-infinite-loops&quot; style=&quot;color: var(--cta);&quot;&gt;Risk 2: Infinite Loops&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#risk-3-data-exposure&quot; style=&quot;color: var(--cta);&quot;&gt;Risk 3: Data Exposure&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#risk-4-unintended-tool-calls&quot; style=&quot;color: var(--cta);&quot;&gt;Risk 4: Unintended Tool Calls&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#risk-5-cost-blowouts&quot; style=&quot;color: var(--cta);&quot;&gt;Risk 5: Cost Blowouts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#built-in-safeguards&quot; style=&quot;color: var(--cta);&quot;&gt;Built-in Safeguards Across Tools&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#deliverable-safety-checklist&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Safety Checklist&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Know the five main risk categories for autonomous agents and have a concrete prevention strategy for each one&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Experience running at least one autonomous agent from Lessons 3-5&lt;/li&gt;
&lt;li&gt;Your agent harness configured and ready (any of the three)&lt;/li&gt;
&lt;li&gt;A text editor for the safety checklist&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-five-risk-categories&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Five Risk Categories&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Autonomous agents can do real work without us. That&#39;s the point. But the same autonomy that makes them powerful makes them dangerous in predictable ways. These five risks recur across every agent platform:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;ol style=&quot;color: var(--text-secondary); line-height: 2.4; margin-bottom: 0; padding-left: 1.5rem;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Token burns&lt;/strong&gt; — The agent calls the API thousands of times on a trivial task&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Infinite loops&lt;/strong&gt; — The agent keeps repeating a cycle without making progress&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data exposure&lt;/strong&gt; — The agent includes sensitive information in API requests&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unintended tool calls&lt;/strong&gt; — The agent deletes files, changes configs, or executes unexpected commands&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost blowouts&lt;/strong&gt; — The agent goes on an expensive tangent and burns through budget&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Each risk has a different shape depending on which tool you&#39;re using and what access you&#39;ve given the agent. Let&#39;s walk through each one with a concrete scenario.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;risk-1-token-burns&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Risk 1: Token Burns&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Scenario&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;You set up a Hermes agent to read a directory of 200 text files and count the number of lines in each. Instead of reading all 200 files in one pass, the agent reads them one at a time — 200 API calls, each one sending the entire directory listing back as context. By file 50, the context window is enormous, each call costs more, and you&#39;ve spent $8 on what should have been a $0.20 task.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;What Happened&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent didn&#39;t batch its work. It treated each file as a separate iteration: perceive (list files), decide (read file 1), act (read it), repeat. With 200 files, that&#39;s 200 iterations, each one growing the context with all previous results. The model charges per token — more context means more expensive calls.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Prevent It&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# In Hermes config: set iteration and cost limits from the start
task:
  max_iterations: 15
  max_cost: 0.50&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Set limits before running.&lt;/strong&gt; Hermes supports both &lt;code class=&quot;accent&quot;&gt;max_iterations&lt;/code&gt; and &lt;code class=&quot;accent&quot;&gt;max_cost&lt;/code&gt;. AutoGPT has iteration limits. Openclaw Fleet has per-agent budgets. Never run an agent without at least one of these limits — they&#39;re the difference between a $0.50 experiment and a $50 surprise.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Prefer batch operations.&lt;/strong&gt; If the task involves many similar sub-tasks (reading 200 files, searching 50 websites), write the prompt to batch them: &quot;Read all 200 files and return the line counts as a JSON array.&quot; A single API call with a list is cheaper and faster than 200 API calls with growing context.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;risk-2-infinite-loops&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Risk 2: Infinite Loops&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Scenario&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;You give AutoGPT a goal: &quot;Find the best pizza place in Toronto and write a review.&quot; The agent searches the web, finds a list of pizzerias, reads a few reviews, and starts writing. But it keeps second-guessing itself — &quot;Am I sure this is the best? Let me check one more source.&quot; It searches again, reads more reviews, and still can&#39;t decide. The loop runs for 47 iterations before you kill it.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;What Happened&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent&#39;s success criteria were vague. &quot;Best&quot; is subjective, and the agent had no way to know when it had gathered enough evidence. Every new source introduced new uncertainty. Without a concrete stopping condition — &quot;use at most 3 sources&quot; or &quot;pick the highest-rated on Google&quot; — the agent kept looping forever, trying to reach certainty it could never attain.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Prevent It&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# A loop-resistant goal has concrete stopping criteria
goal: &quot;Find the best pizza place in Toronto and write a review.
       Search exactly 3 review sites. Pick the restaurant with
       the highest average rating across all 3. Do not search
       additional sources.&quot;
constraints:
  - &quot;Exactly 3 searches, no more&quot;
  - &quot;If ratings are missing from a site, skip it&quot;&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Define &quot;done&quot; concretely.&lt;/strong&gt; The agent needs to know when to stop. Instead of &quot;research X,&quot; say &quot;search 3 sources, compile the findings into a table, stop.&quot; Openclaw Fleet avoids this problem differently — each task has a fixed scope, and the agent moves to the next queue item when it&#39;s done. Hermes runs the task list to completion and stops.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Detection tip:&lt;/strong&gt; Watch for repeated actions in the logs. If you see &quot;search &#39;best pizza Toronto&#39; → read page → search &#39;best pizza Toronto 2026&#39; → read page → search &#39;best pizza Toronto near me&#39;&quot; without the output changing, the agent is in a loop. Kill it, tighten the constraints, and restart.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;risk-3-data-exposure&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Risk 3: Data Exposure&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Scenario&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;You run an Openclaw Fleet agent to audit your project&#39;s configuration files. The agent reads through every config file, looking for security issues. One of those files contains a database password in plain text. The agent reads it, includes it in its context, and sends the context to the model API. The password is now in the model provider&#39;s logs — and potentially in training data.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;What Happened&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent had broad file system access and no understanding of which files contained sensitive data. It treated all files as equal reading material. Any data the agent reads and includes in its context is sent to the model provider with every API call.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Prevent It&lt;/h3&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Exclude sensitive paths from agent access.&lt;/strong&gt; Openclaw Fleet supports per-agent file system restrictions — the Reviewer might only have access to &lt;code class=&quot;accent&quot;&gt;./Tasks/Review/&lt;/code&gt; and never see the production secrets directory.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use environment variables for secrets.&lt;/strong&gt; Never put API keys, passwords, or tokens in files the agent might read. Use the harness&#39;s secret management (Openclaw&#39;s secret bundles, Hermes&#39;s &lt;code class=&quot;accent&quot;&gt;${VAR}&lt;/code&gt; env injection).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Audit what the agent reads.&lt;/strong&gt; Every harness logs which files were accessed. After a run, scan the logs for any sensitive files the agent should not have touched.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;risk-4-unintended-tool-calls&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Risk 4: Unintended Tool Calls&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Scenario&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;You run an agent to clean up temporary files in your project directory. The goal says &quot;delete files matching *.tmp.&quot; The agent runs &lt;code class=&quot;accent&quot;&gt;find . -name &quot;*.tmp&quot; -delete&lt;/code&gt; — but it&#39;s running from the wrong directory. It deletes the entire project&#39;s temporary files, and then some. You lose build artifacts and cached data that take an hour to regenerate.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;What Happened&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Destructive operations (delete, move, overwrite) were available as agent tools, and the agent executed them with insufficient constraints on scope. The same risk applies to any command that can modify state: &lt;code class=&quot;accent&quot;&gt;rm&lt;/code&gt;, &lt;code class=&quot;accent&quot;&gt;mv&lt;/code&gt;, &lt;code class=&quot;accent&quot;&gt;chmod&lt;/code&gt;, database writes, API POST/DELETE calls.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Prevent It&lt;/h3&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Use approval gates for destructive actions.&lt;/strong&gt; Openclaw Fleet has a built-in approval pattern: if a task involves &lt;code class=&quot;accent&quot;&gt;delete&lt;/code&gt;, &lt;code class=&quot;accent&quot;&gt;overwrite&lt;/code&gt;, or &lt;code class=&quot;accent&quot;&gt;deploy&lt;/code&gt;, the agent writes the action to a pending file and waits for human review before executing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sandbox the agent&#39;s file system.&lt;/strong&gt; Hermes can run inside a Docker container with a read-only mount on everything except the working directory. The agent can delete everything in its sandbox and it doesn&#39;t matter — you recreate the container.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Test with dry-run first.&lt;/strong&gt; Add a constraint: &quot;Do not modify any files. Only report what you would change.&quot; Review the report, then run with actual modifications.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;risk-5-cost-blowouts&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Risk 5: Cost Blowouts&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Scenario&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;You set up an AutoGPT agent to research &quot;emerging AI trends&quot; and left it running overnight. No cost cap. No iteration limit. The agent started with general AI research, then narrowed to &quot;agentic AI,&quot; then to &quot;Openclaw alternatives,&quot; then to &quot;Docker Swarm vs Kubernetes.&quot; Each search led to new searches. By morning, the agent had made 1,200 API calls and burned $80. The output: a rambling 15-page document that would have taken a human 30 minutes to write.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;What Happened&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is token burn (Risk 1) at scale, combined with goal decomposition (Lesson 5) gone too far. AutoGPT&#39;s strength — the ability to add new sub-tasks as it learns — became a weakness because nothing stopped it from following tangents indefinitely. The goal was too broad and the constraints too loose.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;How to Prevent It&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Always set both limits
limits:
  max_iterations: 25
  max_cost: 3.00&lt;/pre&gt;&lt;/div&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Always set a cost cap.&lt;/strong&gt; Both Hermes and Openclaw Fleet support cost limits natively. For AutoGPT, use the &lt;code class=&quot;accent&quot;&gt;max_iterations&lt;/code&gt; setting as a proxy — track your per-iteration cost separately.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tighten the goal scope.&lt;/strong&gt; &quot;Research emerging AI trends&quot; is a week-long research project, not a single agent task. Narrow it to &quot;Research the top 3 new AI tools announced in June 2026&quot; — that&#39;s specific enough that the agent will naturally finish.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Start with a small budget for testing.&lt;/strong&gt; Run the goal with a $0.50 cap first. Inspect the output. Does it look like it&#39;s on the right track? Increase the budget for the real run only after testing.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;built-in-safeguards&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Built-in Safeguards Across Tools&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Not all tools handle safety the same way. Here&#39;s what each one offers out of the box:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Safeguard&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Hermes&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Openclaw Fleet&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;AutoGPT&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Cost cap&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Yes — &lt;code class=&quot;accent&quot;&gt;max_cost&lt;/code&gt; in config&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Yes — per-agent budget&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;No native cap (use iteration limit)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Loop detection&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Basic — iteration limit stops the loop&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Advanced — timeout per task, repeated action detection&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Basic — iteration limit, human interrupt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Approval gates&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Not built in&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Yes — approval-required tasks pause for review&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Human-in-the-loop mode available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Sandboxing&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Optional Docker sandbox&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Per-container isolation, read-only mounts&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Docker container, limited isolation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Audit logging&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Console logs, step-by-step output&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Centralized audit trail, per-action logging&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Console logs, task list history&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Openclaw Fleet has the most comprehensive safety features because it was designed for production use where real data and real money are at stake. Hermes covers the basics — enough for learning and low-risk prototyping. AutoGPT has the weakest built-in safeguards, which makes the goal-writing discipline from Lesson 5 even more important: when the tool doesn&#39;t guard you, guard yourself.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-safety-checklist&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Safety Checklist&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Print this, bookmark it, or paste it into your agent workspace. Run through it before every autonomous agent task:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Set a cost cap.&lt;/strong&gt; What&#39;s the maximum you&#39;re willing to spend? Put it in the config before running.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Use a sandbox.&lt;/strong&gt; Run the agent in a container or restricted directory. If it deletes everything, nothing of value is lost.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Define a max step count.&lt;/strong&gt; Even if you think the task will take 5 iterations, set the max to 15. Give the agent room to work but not room to wander.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Log everything.&lt;/strong&gt; Turn on full logging. If something goes wrong, the logs are the only way to understand what the agent was thinking.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;
&lt;p style=&quot;margin: 0.25rem 0; color: var(--text-secondary);&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong&gt;Require approval for destructive actions.&lt;/strong&gt; If the task involves &lt;code class=&quot;accent&quot;&gt;delete&lt;/code&gt;, &lt;code class=&quot;accent&quot;&gt;overwrite&lt;/code&gt;, or any irreversible operation, make the agent pause for human review before executing.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Follow these five rules every time. They cover all five risk categories. They work with all three tools. And they cost nothing to implement — just a few lines of configuration and a habit of discipline before every agent run.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Go back to the agent you ran in Lesson 3, 4, or 5.&lt;/strong&gt; Run the safety checklist against it. Which items did you have in place? Which were missing? For each missing item, configure it now — add a cost cap, tighten the iteration limit, or restrict the agent&#39;s file system access.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Then create a deliberately risky goal — something vague like &quot;explore the file system and report interesting things&quot; — with deliberately loose limits. Watch how the agent behaves. Kill it when it hits a tangent. Note the pattern: it always goes too far if left unchecked. That&#39;s why the checklist exists. Make it your pre-flight routine, and the safety gap between running an agent and running an agent &lt;em&gt;safely&lt;/em&gt; closes to zero.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Comparing the Options&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-8/&quot; style=&quot;color: var(--cta);&quot;&gt;Effective Patterns&lt;/a&gt; — human-in-the-loop checkpoints, logging and observability, timeout strategies, cost budgets, and escalation paths.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 6: Comparing the Options</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 6&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Comparing the Options&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;We&#39;ve toured Hermes, Openclaw Fleet, and AutoGPT. Which one should you actually use?&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#the-decision-framework&quot; style=&quot;color: var(--cta);&quot;&gt;The Decision Framework&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#comparison-table&quot; style=&quot;color: var(--cta);&quot;&gt;Comparison Table&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#scenario-1-research-while-you-sleep&quot; style=&quot;color: var(--cta);&quot;&gt;Scenario 1: Research While You Sleep&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#scenario-2-a-team-of-specialists&quot; style=&quot;color: var(--cta);&quot;&gt;Scenario 2: A Team of Specialists&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#scenario-3-learning-on-a-budget&quot; style=&quot;color: var(--cta);&quot;&gt;Scenario 3: Learning on a Budget&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#deliverable-decision-cheat-sheet&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Decision Cheat Sheet&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Have a clear mental framework for choosing between Hermes, Openclaw Fleet, and AutoGPT based on your specific task and constraints&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 14 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Completed Lessons 3, 4, and 5 (or equivalent experience with each tool)&lt;/li&gt;
&lt;li&gt;A task in mind that you want to automate with agents&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-decision-framework&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Decision Framework&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;We evaluate each tool across four dimensions that matter for real-world use. These aren&#39;t academic — they&#39;re the questions that determine whether a tool will work for your task or fight you the whole way:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What it measures&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Autonomy level&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;How much the agent decides vs how much we pre-plan&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Determines if you need to break down every step or can give a high-level goal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Safety features&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;What guardrails exist for cost, scope, and destructive actions&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;The difference between a safe experiment and an expensive runaway agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Setup effort&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Time from decision to first working agent&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;If it takes a week to set up, it&#39;s not practical for a quick experiment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Scalability&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;How well it handles larger workloads or coordination across tasks&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;A tool that works for 1 task might not work for 50&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;comparison-table&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Comparison Table&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s how all three tools stack up side by side:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Criterion&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Hermes&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Openclaw Fleet&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;AutoGPT&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Autonomy&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Medium — follows your task plan&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Medium-high — autonomous within role boundaries&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;High — decomposes goals into sub-tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Sandboxing&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Optional Docker sandbox&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Per-container isolation, secret bundles&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Container-based, limited isolation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Multi-agent&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;No — single agent, single loop&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Yes — orchestrator + multiple agents&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;No — single agent with sub-task list&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Cost model&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Iteration + cost limits in config&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Per-agent budget, per-task limits&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Iteration limits, no native cost cap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Model cost&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Any OpenAI-compatible API. Use DeepSeek V4 Flash (~$0.09/M input) or GLM 5.2 via Z Code (~$1.40/$4.40) to cut costs 72-97% vs Sonnet/Opus.&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Per-agent model config. Same provider flexibility — DeepSeek and GLM work as drop-in replacements.&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;OpenAI-compatible only. DeepSeek V4 Flash Max and GLM 5.2 supported, dramatically reducing loop costs.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Setup time&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;5–10 minutes (pip install)&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;30–60 minutes (Docker Swarm + config)&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;10–20 minutes (Docker pull + config)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Learning curve&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Low — YAML config, basic concepts&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Medium — Docker, Swarm, queue patterns, roles&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Low-medium — goal writing is the main skill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Best for&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Quick automations, learning the loop pattern&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Production pipelines, sensitive data, multi-step workflows&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Open-ended research, tasks where the path isn&#39;t clear&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Worst for&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Anything that needs multi-agent review or team coordination&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Simple one-shot tasks (overkill)&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Precise operations, steps requiring strict adherence&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;scenario-1-research-while-you-sleep&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Scenario 1: Research While You Sleep&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;The task:&lt;/strong&gt; &quot;I want one agent to do research while I sleep and have a report ready in the morning.&quot;&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;color: var(--text-primary); font-weight: 600; margin-bottom: 0.5rem;&quot;&gt;Recommendation: AutoGPT&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin-bottom: 0.75rem;&quot;&gt;This is AutoGPT&#39;s sweet spot. The goal is open-ended — &quot;research X and compile a report&quot; — and the agent needs the freedom to discover sources, follow leads, and adapt its plan as it learns. AutoGPT&#39;s goal decomposition will break the research into sub-tasks, prioritize them, and work through them overnight.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin-bottom: 0.75rem;&quot;&gt;&lt;strong&gt;Why not the others?&lt;/strong&gt; Hermes would need you to list every search query and source in advance, which defeats the purpose of unattended research. Openclaw Fleet would work but is overkill — you don&#39;t need multiple agents for a single research task.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin: 0;&quot;&gt;&lt;strong&gt;Caveat:&lt;/strong&gt; Set a max iteration limit. Without one, a research-rabbit-hole could run for hours and cost dollars.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;scenario-2-a-team-of-specialists&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Scenario 2: A Team of Specialists&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;The task:&lt;/strong&gt; &quot;I want a team of specialized agents — one to plan, one to code, one to review, one to deploy — working together on a codebase.&quot;&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;color: var(--text-primary); font-weight: 600; margin-bottom: 0.5rem;&quot;&gt;Recommendation: Openclaw Fleet&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin-bottom: 0.75rem;&quot;&gt;This is what Openclaw Fleet was built for. The orchestrator manages the task queue, routes work to the right agent based on role, and ensures that code passes through review before deployment. Each agent gets its own model, its own credentials, and its own security boundary. If the Coder goes rogue, it can&#39;t touch the production deployment secrets.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin-bottom: 0.75rem;&quot;&gt;&lt;strong&gt;Why not the others?&lt;/strong&gt; Hermes can&#39;t do multi-agent. AutoGPT runs one agent — you&#39;d need to run multiple instances and build your own coordination layer. Openclaw Fleet gives you the coordination patterns (fan-out, pipeline, arbitration) built in.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin: 0;&quot;&gt;&lt;strong&gt;Caveat:&lt;/strong&gt; The fleet setup requires Docker Swarm and a solid understanding of roles and queues. Start with two agents (Coder + Reviewer) and add more roles as you get comfortable.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;scenario-3-learning-on-a-budget&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Scenario 3: Learning on a Budget&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;The task:&lt;/strong&gt; &quot;I want to experiment with agent loops, understand how they work, and prototype something quickly — without spending money on a big setup.&quot;&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;color: var(--text-primary); font-weight: 600; margin-bottom: 0.5rem;&quot;&gt;Recommendation: Hermes&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin-bottom: 0.75rem;&quot;&gt;Hermes exists for exactly this use case. Install it with pip, write a YAML config, and you&#39;re running an autonomous loop in under 10 minutes. The built-in iteration and cost limits mean you can experiment without financial risk. Use Hermes to learn the perceive → decide → act → repeat pattern before deciding whether you need more complex tooling.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin-bottom: 0.75rem;&quot;&gt;&lt;strong&gt;Why not the others?&lt;/strong&gt; AutoGPT takes longer to set up and its autonomous decomposition can make learning harder — you don&#39;t see the task structure because the agent handles it. Openclaw Fleet requires Docker Swarm familiarity that&#39;s better built after you understand the basic loop.&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); margin: 0;&quot;&gt;&lt;strong&gt;Caveat:&lt;/strong&gt; Hermes won&#39;t scale to complex multi-agent workflows. That&#39;s fine — it&#39;s not supposed to. Use it to learn, then graduate to Openclaw Fleet or AutoGPT when the task demands it.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-decision-cheat-sheet&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Decision Cheat Sheet&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a flowchart to keep handy. When you have a task, run through these questions:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;div style=&quot;padding: 0.75rem 1rem; background: var(--bg); border-radius: 8px; margin-bottom: 0.75rem;&quot;&gt;
&lt;p style=&quot;margin: 0; color: var(--text-secondary);&quot;&gt;&lt;strong&gt;Q1:&lt;/strong&gt; Do you need multiple agents working together?&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Yes → Openclaw Fleet&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;No → go to Q2&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.75rem 1rem; background: var(--bg); border-radius: 8px; margin-bottom: 0.75rem;&quot;&gt;
&lt;p style=&quot;margin: 0; color: var(--text-secondary);&quot;&gt;&lt;strong&gt;Q2:&lt;/strong&gt; Can you break the task into specific steps in advance?&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Yes, and I want quick setup → Hermes&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;No, the goal is clear but the path isn&#39;t → go to Q3&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.75rem 1rem; background: var(--bg); border-radius: 8px; margin-bottom: 0.75rem;&quot;&gt;
&lt;p style=&quot;margin: 0; color: var(--text-secondary);&quot;&gt;&lt;strong&gt;Q3:&lt;/strong&gt; Is the task open-ended research or discovery?&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Yes → AutoGPT&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;No, it&#39;s a precise operation → break it down further until the steps are clear, then reassess&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Save this cheat sheet or keep it in mind. It distills the decision process into three yes/no questions that cover 90% of agent use cases.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Pick a real task you want to automate.&lt;/strong&gt; It could be something from work, a personal project, or a recurring chore. Run it through the decision cheat sheet above. Which tool does it point to?&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now ask: does that answer feel right? If you have a hunch that a different tool would work better, trust your instinct but be specific about why. The cheat sheet is a starting point, not a rule — the right tool is the one that fits your &lt;em&gt;actual&lt;/em&gt; constraints of time, budget, safety, and desired autonomy.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;If the cheat sheet points to Openclaw Fleet but you&#39;re not ready for Docker Swarm, use Hermes to prototype the first agent and grow from there. The tools are complementary, not competitors. Most teams end up using more than one.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; AutoGPT: Goal Decomposition&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-7/&quot; style=&quot;color: var(--cta);&quot;&gt;Safety Cautions&lt;/a&gt; — token burns, infinite loops, data exposure, unintended tool calls, and cost blowouts. What to watch for and how to prevent it.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 5: AutoGPT: Goal Decomposition</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 5&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;AutoGPT: Goal Decomposition&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;You give AutoGPT a goal and it just... goes. How does it decide what to do first?&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#the-agent-that-plans-its-own-work&quot; style=&quot;color: var(--cta);&quot;&gt;The Agent That Plans Its Own Work&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#how-goal-decomposition-works&quot; style=&quot;color: var(--cta);&quot;&gt;How Goal Decomposition Works&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#the-agent-loop-in-detail&quot; style=&quot;color: var(--cta);&quot;&gt;The Agent Loop in Detail&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#walkthrough-research-with-autogpt&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Research with AutoGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#autogpt-vs-hermes-vs-openclaw&quot; style=&quot;color: var(--cta);&quot;&gt;AutoGPT vs Hermes vs Openclaw&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#deliverable-goal-writing-template&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Goal-Writing Template&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Understand how AutoGPT decomposes goals into sub-tasks, execute them, and iterate, and learn to write goals that produce useful results&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 20 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Docker Desktop installed and running&lt;/li&gt;
&lt;li&gt;An API key from OpenAI, Anthropic, or any OpenAI-compatible provider (DeepSeek, GLM, etc.)&lt;/li&gt;
&lt;li&gt;A terminal open in a project directory&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-agent-that-plans-its-own-work&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Agent That Plans Its Own Work&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes and Openclaw Fleet both expect us to define the sub-steps. We tell Hermes &quot;search these files, read each one, summarize, write the result.&quot; We tell the fleet &quot;Coder implements this, Reviewer audits that.&quot; The human is still doing the task breakdown.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;AutoGPT works the other way. We give it a high-level goal — &quot;Research renewable energy trends in Canada and compile a summary&quot; — and AutoGPT figures out the sub-steps itself. It decides what to search for, what sources to check, how to organize the findings, and when it&#39;s done. We don&#39;t plan the work. We define the destination, and AutoGPT navigates the route.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;This is both the superpower and the risk. AutoGPT can handle goals that are too vague or complex for us to break down in advance. But it can also go down rabbit holes we never intended, because &lt;em&gt;its&lt;/em&gt; idea of &quot;good enough&quot; might not match &lt;em&gt;ours&lt;/em&gt;.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;how-goal-decomposition-works&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;How Goal Decomposition Works&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Goal decomposition is the process of turning a single high-level objective into an ordered list of sub-tasks that, when executed in sequence, achieve the objective. AutoGPT does this in three stages:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Stage&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What happens&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;1. Brainstorm&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Agent generates a list of all possible sub-tasks the goal might require&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&quot;Search for trends, find government reports, check news, compile into sections&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;2. Prioritize&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Agent ranks sub-tasks by importance and dependency order&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&quot;Search first (provides raw material), then read, then organize, then write&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;3. Execute &amp; re-prioritize&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Agent works through the list, adjusting order and adding new sub-tasks as it learns&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&quot;Found a key report → add &#39;extract statistics from this report&#39; to the list&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The critical difference from a fixed task list: AutoGPT can add, remove, and reorder sub-tasks based on what it discovers during execution. If it searches for &quot;renewable energy Canada trends&quot; and finds a promising government PDF, it might add &quot;download and analyze PDF&quot; as a new sub-task on the fly. A fixed-plan agent wouldn&#39;t do that — it would stick to the original script and miss the PDF entirely.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-agent-loop-in-detail&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Agent Loop in Detail&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;AutoGPT&#39;s execution loop refines the basic perceive → decide → act pattern into five steps. Every cycle through this loop moves the agent closer to the goal:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;ol style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 0; padding-left: 1.5rem;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Perceive context&lt;/strong&gt; — The agent reads its current state: what sub-tasks remain, what results came back from the last action, what the goal says. It has access to a &quot;scratchpad&quot; where it can store notes between cycles.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Generate sub-tasks&lt;/strong&gt; — Based on the perceived context, the agent generates a fresh list of sub-tasks. This happens every cycle, not just at the start. The previous list is reconsidered and revised each time.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Select next action&lt;/strong&gt; — The agent picks the single highest-priority sub-task to execute next. It considers dependencies: if &quot;analyze PDF&quot; depends on &quot;download PDF,&quot; it won&#39;t skip ahead.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Execute&lt;/strong&gt; — The agent calls a tool (web search, file read, code execution, etc.) and captures the result.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evaluate result&lt;/strong&gt; — The agent compares the result against its success criteria. Was the search productive? Did the PDF contain useful data? Should we adjust the approach? The evaluation feeds back into the next cycle&#39;s context.&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;This loop runs until the agent decides the goal is met, or it hits a safety limit (max iterations, max cost, or a human interrupt). Because the loop regenerates the task list every cycle, it can adapt to surprises — a broken link, an unexpectedly rich source, a rabbit hole that needs to be cut off.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-research-with-autogpt&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: Research with AutoGPT&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s run AutoGPT with a real research goal and watch it decompose the work.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 1: Install and configure&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Pull the AutoGPT Docker image&lt;/strong&gt; and create a directory for your work:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker pull significantgravitas/auto-gpt
mkdir autogpt-work
cd autogpt-work&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 2: Define the goal&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Create a goal file&lt;/strong&gt; called &lt;code class=&quot;accent&quot;&gt;goal.yaml&lt;/code&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# goal.yaml
goal: &quot;Research renewable energy trends in Canada for 2025-2026.
      Focus on solar, wind, and hydro. Check government reports,
      industry news, and academic sources. Compile findings into
      a markdown report with sections per energy type, key statistics,
      growth trends, and notable projects. Save to canada-renewable-energy.md.&quot;

constraints:
  - &quot;Only use sources published after January 2024&quot;
  - &quot;Prefer Canadian government sources (.gc.ca) where available&quot;
  - &quot;Do not use Wikipedia as a primary source&quot;
  - &quot;Max 10 iterations&quot;

resources:
  - &quot;web_search&quot;
  - &quot;browse_website&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 3: Run it&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Start AutoGPT&lt;/strong&gt; with the goal file:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker run -it &#92;
  -v $(pwd):/app &#92;
  -e OPENAI_API_KEY=$OPENAI_API_KEY &#92;
  significantgravitas/auto-gpt &#92;
  --goal-file /app/goal.yaml&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 4: Watch the task list evolve&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;AutoGPT will print its current task list at the start and update it as it works. Watch how it evolves across iterations:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;┌─ Iteration 1 ───────────────────────────┐
│ Current task list:                       │
│ 1. Search for Canadian renewable energy  │
│ 2. Find government reports on solar      │
│ 3. Find government reports on wind       │
│ 4. Find government reports on hydro      │
│ 5. Check industry news sources           │
│ 6. Compile findings into report          │
│ Executing: Search for Canadian renewable │
│ Result: Found NRCan 2025 outlook page... │
└──────────────────────────────────────────┘
┌─ Iteration 3 ───────────────────────────┐
│ Current task list:                       │
│ 1. Browse NRCan energy outlook            │
│ 2. Extract solar statistics              │
│ 3. [NEW] Check CanREA 2025 market report │
│ 4. Find government reports on wind       │
│ 5. Find government reports on hydro      │
│ Executing: Browse NRCan energy outlook    │
│ Result: Page contains 2025 solar data... │
└──────────────────────────────────────────┘&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Notice how iteration 3 added a new sub-task (&quot;Check CanREA 2025 market report&quot;) that wasn&#39;t in the original list. The agent found a reference to the Canadian Renewable Energy Association and decided the report was worth checking. A fixed-plan agent would have missed it.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Key insight:&lt;/strong&gt; The task list is never final. AutoGPT treats it as a hypothesis — &quot;these are the steps I think I need&quot; — that gets revised every cycle. This is what makes it good at open-ended research and what makes it unpredictable for precise operations.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;autogpt-vs-hermes-vs-openclaw&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;AutoGPT vs Hermes vs Openclaw&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Each tool takes a different approach to task planning:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;AutoGPT&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Hermes&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Openclaw Fleet&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Task planning&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Agent decomposes goal autonomously&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;You define the task list upfront&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;You define roles and workflows; agents handle execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Autonomy level&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Highest — decides what to do and when to stop&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Medium — follows your plan but executes autonomously&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Medium-high — works autonomously within role boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Controllability&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Low — can go in unexpected directions&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;High — you control the steps, agent controls execution&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Medium — roles constrain scope, but within-role decisions are autonomous&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Best for&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Open-ended research, discovery, creative exploration&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Well-defined multi-step tasks, prototyping&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Production pipelines, multi-agent coordination, sensitive work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Worst for&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Precise operations, tasks requiring strict adherence to process&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Goals that are too vague for you to break down&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Simple one-shot tasks that don&#39;t need a team&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;AutoGPT is the most autonomous option in this track. That&#39;s its strength for research and its weakness for anything that needs a repeatable, predictable process. Use it when the goal is clear but the path isn&#39;t — that&#39;s exactly the gap goal decomposition is designed to bridge.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-goal-writing-template&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Goal-Writing Template&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;A well-formed goal is the difference between AutoGPT producing a useful result and going on an expensive tangent. Here&#39;s a template that works:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# goal-template.yaml
# Use this template for every AutoGPT goal

goal: |
  [What success looks like — be concrete]
  Example: &quot;Research renewable energy trends in Canada and compile
  a summary report with sections per energy type and key statistics.&quot;

constraints:
  - [What to avoid — boundaries that keep the agent focused]
  - &quot;Only use sources from 2024 or later&quot;
  - &quot;Do not use Wikipedia as a primary source&quot;
  - &quot;Max [N] iterations&quot;

resources:
  # Tools the agent is allowed to use
  - &quot;web_search&quot;
  - &quot;browse_website&quot;

# Optional: set a cost limit
limits:
  max_cost: 1.00&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Every goal should answer three questions:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;What does success look like?&lt;/strong&gt; — Not just &quot;research X&quot; but &quot;a markdown file with sections for each subtopic, at least 3 sources per section, saved to this path.&quot; The more concrete, the less room for the agent to wander.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What tools are available?&lt;/strong&gt; — Make the allowed tools explicit. If the agent doesn&#39;t need file write access, don&#39;t give it to them. Each tool is a potential rabbit hole.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What should we avoid?&lt;/strong&gt; — Surface constraints upfront: &quot;Don&#39;t use Wikipedia,&quot; &quot;Don&#39;t spend more than $1.00,&quot; &quot;Max 15 iterations.&quot; Boundaries are the safety net for autonomous goal decomposition.&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Run the renewable energy research goal from the walkthrough.&lt;/strong&gt; Watch the task list evolve across iterations. After it finishes (or hits its limit), open the generated report and evaluate: How complete is it? Did it go in directions you didn&#39;t expect? Would you trust the sources it used?&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Then write a goal for something &lt;em&gt;you&lt;/em&gt; need researched. Run it through the goal-writing template first, refine the constraints, and let AutoGPT work. Compare the output from AutoGPT with what you&#39;d get from searching manually. The goal decomposition pattern means you can set the direction and let the agent handle the navigation — but only if the goal is well-formed enough to keep it on course.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Openclaw Fleet Mode&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-6/&quot; style=&quot;color: var(--cta);&quot;&gt;Comparing the Options&lt;/a&gt; — a decision framework for choosing between Hermes, Openclaw, and AutoGPT across autonomy, safety, setup effort, and scalability.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 4: Openclaw Fleet Mode</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 4&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Openclaw Fleet Mode&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;Single agents are powerful, but some tasks need more than one. What if you could have a team of specialized agents working together?&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#one-agent-isnt-always-enough&quot; style=&quot;color: var(--cta);&quot;&gt;One Agent Isn&#39;t Always Enough&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#what-fleet-mode-is&quot; style=&quot;color: var(--cta);&quot;&gt;What Fleet Mode Is&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#agent-roles-and-the-orchestrator&quot; style=&quot;color: var(--cta);&quot;&gt;Agent Roles and the Orchestrator&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#coordination-patterns&quot; style=&quot;color: var(--cta);&quot;&gt;Coordination Patterns&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#walkthrough-a-two-agent-fleet&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: A Two-Agent Fleet&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#single-agent-vs-fleet&quot; style=&quot;color: var(--cta);&quot;&gt;Single Agent vs Fleet&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#deliverable-fleet-config-template&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Fleet Config Template&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Understand how Openclaw Fleet Mode runs multiple agents as coordinated services, and learn the coordination patterns that make multi-agent teams work&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 18 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Basic understanding of agents from the Basic Agents track&lt;/li&gt;
&lt;li&gt;Familiarity with Docker and Docker Swarm (covered in Lesson 2)&lt;/li&gt;
&lt;li&gt;A terminal and Docker Desktop running&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;one-agent-isnt-always-enough&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;One Agent Isn&#39;t Always Enough&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes runs one agent, one loop, one task at a time. That&#39;s fine for a lot of things. But some workflows need more than one agent working at once — not because the task is too big for one agent, but because one agent can&#39;t do everything well.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Consider a task like: &quot;Review this pull request, check for bugs, run the tests, suggest fixes, implement the fixes, verify the tests pass, and deploy.&quot; A single agent could do all of that, but it would be doing everything with the same model, the same tool access, the same attention. That&#39;s like asking the same person to write code, review it, test it, and deploy it — there&#39;s no separation of concerns, no second pair of eyes, no specialized tooling.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;A fleet splits that work across specialized agents. One agent writes code. Another reviews it. A third runs tests. A fourth deploys. They work in parallel where possible, pass results to each other where not, and the whole pipeline runs faster and more safely than any single agent could manage alone.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-fleet-mode-is&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What Fleet Mode Is&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Openclaw Fleet Mode runs multiple agent instances as coordinated services inside Docker Swarm. Each agent is a separate container with its own model configuration, tool access, and security boundary. An &lt;strong&gt;orchestrator&lt;/strong&gt; service manages the fleet — it watches a shared task queue, assigns work to agents based on their role, collects results, and handles retries when an agent fails.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;The architecture looks like this:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Component&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Orchestrator&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Manages the task queue, dispatches work, collects results, handles retries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Agent containers&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Individual agent instances, each with its own role, tools, and model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Task queue&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Shared message queue where tasks wait to be claimed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Shared spec&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Common reference data (agent topology, secrets, port assignments) that all agents read&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Logs &amp; audit&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Centralized logging that every agent writes to, for observability and post-mortem&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Each agent runs in its own container with its own secret bundle. The orchestrator never sees the agents&#39; API keys — it only sees the task assignments and results. This means even if one agent is compromised, the others and their credentials are isolated.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;agent-roles-and-the-orchestrator&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Agent Roles and the Orchestrator&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;In a fleet, every agent has a &lt;strong&gt;role&lt;/strong&gt;. The role determines what kind of work the agent accepts, what tools it has access to, and what model configuration it uses. These are common roles:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Role&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What it does&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Typical tool access&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Coder&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Implements plans, writes code, edits files&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;File system, git, code execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Reviewer&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Audits completed work, runs tests, validates output&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;File system read, test runner, logging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Planner&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Decomposes goals into session plans, resolves ambiguity&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Read-only file system, web search, note storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Accountant&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Runs bookkeeping pipelines, reconciles transactions&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Receipt images, bank CSVs, Zoho Books API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Sentry&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Monitors the fleet, runs health checks, enforces schedules&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Docker socket, health endpoints, schedule files&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The orchestrator is the traffic controller. It doesn&#39;t do the work itself — it reads new tasks from the queue, checks which agents are available, and dispatches each task to the agent whose role matches the work type. If an agent crashes or a task times out, the orchestrator re-queues it for another agent. If all agents of a role are busy, tasks wait in the queue until one frees up.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;coordination-patterns&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Coordination Patterns&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Once we have multiple agents, we need ways for them to work together. There are three fundamental coordination patterns Openclaw supports:&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Fan-Out&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;One task spawns many independent sub-tasks that can run in parallel. The orchestrator distributes them across available agents. Example: &quot;Analyze all 20 files in this directory&quot; — the orchestrator sends each file to a different Coder agent, and they all work at the same time. The pattern scales linearly with the number of agents.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Sequential (Pipeline)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Tasks must run one after another because each depends on the previous step&#39;s output. The orchestrator routes the output of agent A to the input of agent B. Example: Coder writes code → Reviewer audits it → Coder fixes issues → Deployer ships it. Each step waits for the previous one to complete.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Arbitration&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Multiple agents work on the same problem with different approaches, and a decider agent compares their results and picks the best one. Example: Three agents each design a database schema independently. A fourth agent reviews all three, combines the best ideas, and produces a final design. This is expensive — multiple agents burning tokens on the same task — but useful when correctness matters more than speed.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-a-two-agent-fleet&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: A Two-Agent Fleet&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s set up a basic two-agent fleet: a Coder agent and a Reviewer agent, sharing a queue. The Coder writes code; the Reviewer checks it. This is the simplest useful fleet — two agents, two roles, one pipeline.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 1: Define the shared spec&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Create a shared configuration file&lt;/strong&gt; that both agents will read. Openclaw uses a &lt;code class=&quot;accent&quot;&gt;opencode.json&lt;/code&gt; file to define the fleet:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;{
  &quot;fleet&quot;: {
    &quot;name&quot;: &quot;basic-fleet&quot;,
    &quot;orchestrator&quot;: {
      &quot;mode&quot;: &quot;queue&quot;,
      &quot;queue_dir&quot;: &quot;./Tasks/Queue&quot;,
      &quot;poll_interval&quot;: 30
    },
    &quot;agents&quot;: {
      &quot;coder&quot;: {
        &quot;role&quot;: &quot;Coder&quot;,
        &quot;model&quot;: &quot;deepseek-v4-flash&quot;,
        &quot;work_dir&quot;: &quot;./Tasks/Code&quot;
      },
      &quot;reviewer&quot;: {
        &quot;role&quot;: &quot;Reviewer&quot;,
        &quot;model&quot;: &quot;deepseek-v4-pro&quot;,
        &quot;work_dir&quot;: &quot;./Tasks/Review&quot;
      }
    }
  }
}&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 2: Deploy the stack&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Open a terminal&lt;/strong&gt; and run:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker stack deploy -c docker-compose.interclaw.yml basic-fleet&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This launches the orchestrator service and both agent containers. Each agent starts up, registers with the orchestrator, and waits for work.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 3: Add a task to the queue&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Create a task file&lt;/strong&gt; in the queue directory:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Tasks/Queue/001-add-user-auth.md
---
role: Coder
goal: Implement user authentication for the Flask app
  - Add a login endpoint with email and password
  - Add a registration endpoint
  - Hash passwords with bcrypt
  - Return JWT tokens on successful login
---&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The orchestrator picks up this file, sees the &lt;code class=&quot;accent&quot;&gt;role: Coder&lt;/code&gt; header, and dispatches it to the Coder agent.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 4: Watch the pipeline&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The Coder agent reads the task, implements the code, and writes the output. When it finishes, it moves the task file to the Reviewer&#39;s queue. The Reviewer agent picks it up, inspects the code, runs any tests, and either approves it (moving it to done) or writes feedback and sends it back to the Coder queue for fixes.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Watch the logs:&lt;/strong&gt; Run &lt;code class=&quot;accent&quot;&gt;docker service logs basic-fleet_orchestrator --follow&lt;/code&gt; to see each task get dispatched. You&#39;ll see the orchestrator print lines like &lt;em&gt;&quot;Dispatching 001-add-user-auth.md to coder&quot;&lt;/em&gt; and later &lt;em&gt;&quot;Task 001-add-user-auth.md completed by coder, routing to reviewer&quot;&lt;/em&gt;.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;single-agent-vs-fleet&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Single Agent vs Fleet&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;When does a fleet make sense? Here&#39;s how the trade-offs compare:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Single Agent&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Fleet&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Throughput&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;One task at a time&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Multiple tasks in parallel&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Specialization&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary););&quot;&gt;One model, one tool set for everything&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Each agent optimized for its role&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Fault tolerance&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;If it crashes, everything stops&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Other agents keep working; orchestrator reassigns tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Safety&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;One credential set, one attack surface&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Isolated credentials per agent, blast radius contained&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Complexity&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Install one tool, configure once&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Docker Swarm, shared config, queue management&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Cost&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;One model inference at a time&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Multiple agents calling models simultaneously&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The fleet wins on throughput, specialization, and safety — at the cost of operational complexity and higher peak cost. Single agents win on simplicity and cost control. The right choice depends on the task: a single agent is perfect for a focused automation, while a fleet shines when you need a reliable, scalable multi-step pipeline.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-fleet-config-template&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Fleet Config Template&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a reusable two-agent fleet configuration. Save it as &lt;code class=&quot;accent&quot;&gt;opencode.json&lt;/code&gt; and adapt it for your own fleet:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;{
  &quot;fleet&quot;: {
    &quot;name&quot;: &quot;my-fleet&quot;,
    &quot;orchestrator&quot;: {
      &quot;mode&quot;: &quot;queue&quot;,
      &quot;queue_dir&quot;: &quot;./Tasks/Queue&quot;,
      &quot;poll_interval&quot;: 30,
      &quot;max_retries&quot;: 3
    },
    &quot;agents&quot;: {
      &quot;coder&quot;: {
        &quot;role&quot;: &quot;Coder&quot;,
        &quot;model&quot;: &quot;deepseek-v4-flash&quot;,
        &quot;work_dir&quot;: &quot;./Tasks/Code&quot;,
        &quot;tools&quot;: [&quot;filesystem&quot;, &quot;git&quot;, &quot;code_execution&quot;],
        &quot;max_iterations&quot;: 50
      },
      &quot;reviewer&quot;: {
        &quot;role&quot;: &quot;Reviewer&quot;,
        &quot;model&quot;: &quot;deepseek-v4-pro&quot;,
        &quot;work_dir&quot;: &quot;./Tasks/Review&quot;,
        &quot;tools&quot;: [&quot;filesystem_read&quot;, &quot;test_runner&quot;],
        &quot;max_iterations&quot;: 20
      }
    },
    &quot;limits&quot;: {
      &quot;max_concurrent_tasks&quot;: 4,
      &quot;per_agent_budget&quot;: 2.00
    }
  }
}&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Start with two agents and a simple pipeline — Coder → Reviewer. Once that&#39;s running reliably, add a Planner agent that prepares work for the Coder, or a Deployer agent that ships finished code. Each new role makes the fleet more capable, but add them one at a time so you understand the coordination patterns before scaling up.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Set up the two-agent fleet from the walkthrough.&lt;/strong&gt; Deploy the stack, add a task to the queue, and watch the orchestrator dispatch it from the Coder to the Reviewer. Then swap the models — give the Coder a cheaper model and the Reviewer a more expensive one — and see how that changes the pipeline speed and output quality.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Once you&#39;ve seen the basic pipeline work, try adding a third role. Create a Planner agent that pre-processes tasks before they reach the Coder. The Planner reads each task, breaks it into sub-steps, and enriches the task file with implementation notes. This three-agent pattern — Planner → Coder → Reviewer — is the foundation that most production fleets are built on.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Hermes: Setup and First Loop&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-5/&quot; style=&quot;color: var(--cta);&quot;&gt;AutoGPT: Goal Decomposition&lt;/a&gt; — how AutoGPT breaks a high-level goal into sub-tasks, prioritizes them, and iterates autonomously.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 3: Hermes: Setup and First Loop</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 3&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Hermes: Setup and First Loop&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;You&#39;ve heard of Hermes as a lightweight agent framework. Let&#39;s install it and run your first autonomous loop in under 20 minutes.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#what-is-hermes&quot; style=&quot;color: var(--cta);&quot;&gt;What Is Hermes?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#step-1-install&quot; style=&quot;color: var(--cta);&quot;&gt;Step 1: Install&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#step-2-configure&quot; style=&quot;color: var(--cta);&quot;&gt;Step 2: Configure a Model Provider&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#step-3-define-a-tool&quot; style=&quot;color: var(--cta);&quot;&gt;Step 3: Define a Tool&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#step-4-write-a-task&quot; style=&quot;color: var(--cta);&quot;&gt;Step 4: Write a Task Goal&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#step-5-run-it&quot; style=&quot;color: var(--cta);&quot;&gt;Step 5: Run It&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#the-output&quot; style=&quot;color: var(--cta);&quot;&gt;The Output&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/&quot; style=&quot;color: var(--cta);&quot;&gt;Hermes vs Other Loop Frameworks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#deliverable-your-template&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Your Hermes Config Template&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Install Hermes, configure it, and run a real autonomous task with a custom tool&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 20 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Python 3.10+ installed (&lt;code class=&quot;accent&quot;&gt;python --version&lt;/code&gt; to check)&lt;/li&gt;
&lt;li&gt;An OpenAI API key (or Anthropic, or any provider Hermes supports)&lt;/li&gt;
&lt;li&gt;Docker from &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/&quot; style=&quot;color: var(--cta);&quot;&gt;Lesson 2&lt;/a&gt; (optional, for sandboxing)&lt;/li&gt;
&lt;li&gt;A terminal open in a project directory&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-is-hermes&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What Is Hermes?&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes is a minimal autonomous agent framework. It&#39;s a Python library that gives a language model a &lt;strong&gt;perceive → decide → act → repeat&lt;/strong&gt; loop with configurable tools, model endpoints, and logging.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;What makes Hermes different from the big platforms:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Lightweight&lt;/strong&gt; — install it with pip, no Docker required (though you can add Docker for sandboxing). Single binary feel.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Config-driven&lt;/strong&gt; — define your model, tools, and task in a YAML or JSON config file. No code to write unless you need a custom tool.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transparent&lt;/strong&gt; — every loop iteration is logged. You can see what the model perceived, what it decided, and what action it took.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Single-agent focused&lt;/strong&gt; — Hermes runs one agent loop at a time. No multi-agent orchestration, no fleet management. That simplicity makes it ideal for learning the loop pattern.&lt;/li&gt;
&lt;/ul&gt;
&lt;p class=&quot;mb-4&quot;&gt;Think of Hermes as the training wheels for autonomous agents. It gives you the core loop with minimal ceremony, so you can focus on understanding the pattern before moving to more complex platforms.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-1-install&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 1: Install&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Open a terminal&lt;/strong&gt; and run:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;pip install hermes-agent&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;That&#39;s it. Confirm it worked:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;hermes --version&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;If you see a version number, you&#39;re ready.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-2-configure&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 2: Configure a Model Provider&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes needs to know which model to use and how to authenticate. Create a file called &lt;code class=&quot;accent&quot;&gt;hermes-config.yaml&lt;/code&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;model:
  provider: openai
  name: gpt-4o-mini
  api_key: &quot;${OPENAI_API_KEY}&quot;&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Set the environment variable&lt;/strong&gt; with your API key:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# macOS / Linux
export OPENAI_API_KEY=&quot;sk-...&quot;

# Windows PowerShell
$env:OPENAI_API_KEY = &quot;sk-...&quot;&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes reads the key from the environment. Never put the key directly in the config file — that&#39;s how keys get committed to git by accident.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Non-US model alternatives:&lt;/strong&gt; Hermes works with any OpenAI-compatible API. For a Sonnet-class agent, try &lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt; at $0.09/M input — roughly 97% less than Sonnet. For an Opus-class agent, try &lt;strong&gt;GLM 5.2&lt;/strong&gt; at roughly 72% less via Z Code ($1.40/$4.40 vs $5/$25). Just change the &lt;code class=&quot;accent&quot;&gt;provider&lt;/code&gt; and &lt;code class=&quot;accent&quot;&gt;name&lt;/code&gt; in your config — no other changes needed. At these prices, you can run autonomous loops for hours on pocket change.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-3-define-a-tool&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 3: Define a Tool&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;A tool is something the agent can call. Hermes comes with a few built-in tools (file read, file write, web fetch), but let&#39;s add a custom one to see how it works.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Add a &lt;code class=&quot;accent&quot;&gt;tools&lt;/code&gt; section to your config file:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;model:
  provider: openai
  name: gpt-4o-mini
  api_key: &quot;${OPENAI_API_KEY}&quot;

tools:
  - name: search_files
    description: &quot;Search for files by name pattern in a directory&quot;
    command: find  -name &quot;&quot; 2&gt;/dev/null
    parameters:
      dir:
        type: string
        description: &quot;Directory to search in&quot;
      pattern:
        type: string
        description: &quot;Filename pattern, e.g. *.txt or report-*&quot;&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;This tool wraps the Unix &lt;code class=&quot;accent&quot;&gt;find&lt;/code&gt; command. The agent can call it with a directory path and a filename pattern, and get a list of matching files back. The parameters are injected into the command template — that&#39;s the &lt;code class=&quot;accent&quot;&gt;&lt;/code&gt; and &lt;code class=&quot;accent&quot;&gt;&lt;/code&gt; syntax.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-4-write-a-task&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 4: Write a Task Goal&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now we give the agent something to do. Create &lt;code class=&quot;accent&quot;&gt;task.yaml&lt;/code&gt;:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;goal: |
  Search the /workspace directory for all Markdown files (*.md).
  Read each one and summarize what it covers in 2-3 sentences.
  Write the summaries to /workspace/summary.md.

max_iterations: 15
max_cost: 0.50&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Notice the guardrails: &lt;code class=&quot;accent&quot;&gt;max_iterations&lt;/code&gt; stops the loop after 15 turns no matter what. &lt;code class=&quot;accent&quot;&gt;max_cost&lt;/code&gt; stops it if the API spend exceeds $0.50. These are your safety net — the agent cannot run forever or blow through your budget.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;step-5-run-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Step 5: Run It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;From the terminal&lt;/strong&gt;, run:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;hermes run --config hermes-config.yaml --task task.yaml&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Hermes will:&lt;/p&gt;
&lt;ol style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem;&quot;&gt;
&lt;li&gt;&lt;strong&gt;Perceive&lt;/strong&gt; — Call &lt;code class=&quot;accent&quot;&gt;search_files&lt;/code&gt; on /workspace with pattern &lt;code class=&quot;accent&quot;&gt;*.md&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decide&lt;/strong&gt; — Look at the file list and decide which to read first&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Act&lt;/strong&gt; — Read each file (using the built-in file_read tool)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Repeat&lt;/strong&gt; — Keep going until all files are summarized, then write the output&lt;/li&gt;
&lt;/ol&gt;
&lt;p class=&quot;mb-4&quot;&gt;Each step is logged to stdout. Watch the loop unfold:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;┌─ Iteration 1 ──────────────────────┐
│ Perceive: searching Markdown files │
│ Result: found 3 files              │
│ Decide: read file 1 (readme.md)    │
└────────────────────────────────────┘
┌─ Iteration 2 ──────────────────────┐
│ Perceive: reading readme.md        │
│ Result: 245 lines                  │
│ Decide: summarize and move to file │
└────────────────────────────────────┘&lt;/pre&gt;&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-output&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Output&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;When the agent finishes, open &lt;code class=&quot;accent&quot;&gt;/workspace/summary.md&lt;/code&gt;. You should see a summary of each Markdown file the agent found, written in natural language. If the agent hit the iteration limit before finishing, you&#39;ll see partial output — and you can increase &lt;code class=&quot;accent&quot;&gt;max_iterations&lt;/code&gt; for the next run.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Troubleshooting:&lt;/strong&gt; If the agent fails immediately, the most common issues are: API key not set (double-check the env variable), model name wrong (check the exact name on the provider&#39;s docs), or the tool command failing (test it manually first). Hermes logs every error to stderr — read the error message, it usually tells you exactly what&#39;s wrong.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;hermes-vs-others&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Hermes vs Other Loop Frameworks&lt;/h2&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Hermes&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Openclaw Fleet&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;AutoGPT&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Setup&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;pip install, one config file&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Docker Swarm, multi-service deployment&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Docker or local Python, plugin ecosystem&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Scope&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Single agent, single task&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Multi-agent fleet with queues and coordination&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Single agent, goal decomposition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Safety features&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Iteration limits, cost limits, basic logging&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Secrets isolation, sandboxed containers, audit trails, per-agent budget&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Human-in-the-loop mode, token limits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Best for&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Learning the loop, quick automations, prototyping&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Production multi-agent workflows, sensitive tasks&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Long-running tasks with dynamic sub-goals&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Hermes wins on &lt;strong&gt;simplicity&lt;/strong&gt;. It&#39;s the fastest way to get an autonomous loop running. Openclaw Fleet wins on &lt;strong&gt;safety and scale&lt;/strong&gt; — it&#39;s what you&#39;d use for production agent pipelines. AutoGPT wins on &lt;strong&gt;goal decomposition&lt;/strong&gt; — if your task is vague and you need the agent to figure out the sub-steps itself, AutoGPT&#39;s approach shines.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-your-template&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Your Hermes Config Template&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a reusable template. Save it as &lt;code class=&quot;accent&quot;&gt;hermes-task-template.yaml&lt;/code&gt; and adapt it for each new task:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# hermes-task-template.yaml
# Copy and customize for each autonomous task

model:
  provider: openai
  name: gpt-4o-mini  # or gpt-4o, claude-sonnet-4, etc.
  api_key: &quot;${API_KEY}&quot;

tools:
  - name: search_files
    description: &quot;Find files by name pattern&quot;
    command: find  -name &quot;&quot; 2&gt;/dev/null
    parameters:
      dir: { type: string, description: &quot;Directory to search&quot; }
      pattern: { type: string, description: &quot;Filename pattern&quot; }

  - name: read_file
    description: &quot;Read a file&#39;s contents&quot;
    command: cat 
    parameters:
      path: { type: string, description: &quot;Full path to the file&quot; }

task:
  goal: &quot;Describe what you want the agent to do&quot;
  max_iterations: 15
  max_cost: 0.50
  output: &quot;result.md&quot;&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Start with &lt;code class=&quot;accent&quot;&gt;max_iterations: 15&lt;/code&gt; and &lt;code class=&quot;accent&quot;&gt;max_cost: 0.50&lt;/code&gt;. Adjust upward as you gain confidence. Hermes is designed to be safe by default — tighten the limits for learning, then gradually expand as you understand the agent&#39;s behavior.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Run the full walkthrough above.&lt;/strong&gt; Install Hermes, create the config files, and run a real autonomous task. When it finishes, open the output file and inspect the work. Then change the task goal to something you actually need done — &quot;list all Python files in my project and count their lines,&quot; or &quot;fetch the top story from a news site and summarize it.&quot;&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;The loop you just ran is the same pattern used by every autonomous agent platform. The tools change, the config format changes, but the &lt;strong&gt;perceive → decide → act → repeat&lt;/strong&gt; cycle is universal. You&#39;ve built it. You&#39;ve run it. Now you own it.&lt;/p&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Sandboxing: Why and How&lt;/a&gt;&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-4/&quot; style=&quot;color: var(--cta);&quot;&gt;Openclaw Fleet Mode&lt;/a&gt; — multi-agent orchestration, fleet dispatch, agent roles, task queues, and coordination patterns.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 2: Sandboxing: Why and How</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 2&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;Sandboxing: Why and How&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;An autonomous agent has file access, can run code, can call APIs. What happens when it makes a mistake? Not if — when. Sandboxing is the seatbelt.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#the-seatbelt&quot; style=&quot;color: var(--cta);&quot;&gt;The Seatbelt&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#what-a-sandbox-restricts&quot; style=&quot;color: var(--cta);&quot;&gt;What a Sandbox Restricts&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#walkthrough-docker-sandbox&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Docker Sandbox&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#non-docker-options&quot; style=&quot;color: var(--cta);&quot;&gt;Non-Docker Options&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#compare-the-approaches&quot; style=&quot;color: var(--cta);&quot;&gt;Compare the Approaches&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#deliverable-sandbox-checklist&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Sandbox Checklist&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Know why autonomous agents need isolation and how to set up a sandbox that contains failures without getting in the way&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Docker Desktop installed (or access to a machine where Docker runs)&lt;/li&gt;
&lt;li&gt;Understanding of the agent loop from &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/&quot; style=&quot;color: var(--cta);&quot;&gt;Lesson 1&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;A terminal open&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-seatbelt&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Seatbelt&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s a hard truth about autonomous agents: &lt;strong&gt;they will make mistakes.&lt;/strong&gt; Not every time, but often enough that you need to plan for it. The agent might delete a file you needed, run a command that costs money, call an API in a loop until your credit card screams, or follow a rabbit hole of reasoning that wastes hundreds of tokens on nothing useful.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is not the agent being malicious. It&#39;s the agent being wrong. Language models hallucinate. They misinterpret goals. They get trapped in loops. They call the right tool with the wrong argument. When we&#39;re sitting next to the agent, we catch these mistakes quickly. When the agent runs without us, we don&#39;t see them until we come back.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Sandboxing is what makes autonomy safe. It doesn&#39;t prevent mistakes — no tool can do that. It &lt;strong&gt;contains&lt;/strong&gt; them. A sandbox ensures that when the agent messes up, the damage stays inside a controlled zone and doesn&#39;t spread to the rest of your system.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;what-a-sandbox-restricts&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What a Sandbox Restricts&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;A sandbox wraps the agent with boundaries. These boundaries cover four areas:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Boundary&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What it controls&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;File system&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Which directories the agent can read and write&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Prevents the agent from accidentally deleting your Documents folder or reading your SSH keys&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Network&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Which hosts and ports the agent can reach&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Prevents the agent from hitting production APIs, internal services, or unexpected endpoints that could run up bills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Execution&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;What commands and binaries the agent can run&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Prevents the agent from running arbitrary shell commands or installing software without oversight&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Resources&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Max CPU, memory, disk, and token budget&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Prevents the agent from consuming all your system resources or running up an unbounded API bill&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Different sandboxing approaches enforce these boundaries at different levels. Let&#39;s look at the most common one first: Docker.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-docker-sandbox&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: Docker Sandbox&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Docker is the most common way to sandbox autonomous agents. A Docker container is a lightweight, isolated environment with its own file system, network, and process space. The agent sees only what we put inside — nothing more.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 1: Create a Dockerfile&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;This Dockerfile creates a minimal environment with Python and common agent tools:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;FROM python:3.12-slim

# Install agent dependencies
RUN pip install --no-cache-dir requests beautifulsoup4

# Create a working directory the agent can use
WORKDIR /workspace

# Set a non-root user so the agent doesn&#39;t run as root
RUN useradd -m agent &amp;&amp; chown agent:agent /workspace
USER agent&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 2: Build the image&lt;/h3&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker build -t agent-sandbox .&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 3: Run with restricted mounts&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Mount&lt;/strong&gt; only the directories the agent needs. Use &lt;code class=&quot;accent&quot;&gt;:ro&lt;/code&gt; for read-only mounts if the agent only needs to read:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;docker run --rm &#92;
  -v /path/to/input-data:/workspace/input:ro &#92;
  -v /path/to/output:/workspace/output &#92;
  --network none &#92;
  --memory 2g &#92;
  --cpus 2 &#92;
  agent-sandbox&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s break down what each flag does:&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;&lt;strong&gt;&lt;code class=&quot;accent&quot;&gt;--rm&lt;/code&gt;&lt;/strong&gt; — Delete the container when it finishes. No lingering state.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code class=&quot;accent&quot;&gt;-v&lt;/code&gt;&lt;/strong&gt; — Mount specific directories. &lt;code class=&quot;accent&quot;&gt;:ro&lt;/code&gt; means read-only for input data. The output directory is writable so the agent can save results.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code class=&quot;accent&quot;&gt;--network none&lt;/code&gt;&lt;/strong&gt; — No network access at all. The agent can&#39;t call any APIs, period.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code class=&quot;accent&quot;&gt;--memory 2g&lt;/code&gt;&lt;/strong&gt; — Max 2 GB of RAM. The agent can&#39;t eat your machine.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;code class=&quot;accent&quot;&gt;--cpus 2&lt;/code&gt;&lt;/strong&gt; — Max 2 CPU cores.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Step 4: Run the agent inside&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Now we tell the agent harness to use this container. Most harnesses (Opencode, Claude Code CLI) support a &lt;code class=&quot;accent&quot;&gt;--sandbox&lt;/code&gt; or &lt;code class=&quot;accent&quot;&gt;--container&lt;/code&gt; flag, or a config option that points to the Docker image:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Example: run an agent task inside the sandbox container
hermes run --sandbox agent-sandbox --task &quot;organize input files by type&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Pro tip:&lt;/strong&gt; Start with &lt;code class=&quot;accent&quot;&gt;--network none&lt;/code&gt; even if the agent needs network access later. Run the first test with no network. Once you verify the agent works correctly, add network access — but restrict it to specific hosts using &lt;code class=&quot;accent&quot;&gt;--add-host&lt;/code&gt; or a Docker network with an allowlist.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;non-docker-options&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Non-Docker Options&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Docker isn&#39;t the only way to sandbox. Two alternatives work well for simpler setups:&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Option 1: Limited User Account&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Create a dedicated user on your machine with restricted permissions. The agent runs as that user and can only read/write files the user owns. This is less isolated than Docker — the agent can still see other processes, consume system resources, and access any file the user can access — but it&#39;s quick to set up and doesn&#39;t require Docker.&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Linux / macOS
sudo useradd -m agent-worker
sudo -u agent-worker hermes run --task &quot;process data&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Option 2: Temp Directory Workspace&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Point the agent to a temporary directory that you create fresh for each task. When the task is done, delete the directory. This is the lightest sandbox: it protects your file system from accidental writes but does nothing to control network access or resource usage.&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Create a fresh workspace
WORKDIR=$(mktemp -d)
cd &quot;$WORKDIR&quot;
hermes run --task &quot;generate report&quot;

# Clean up when done
rm -rf &quot;$WORKDIR&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;compare-the-approaches&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Compare the Approaches&lt;/h2&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Docker&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Limited User&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Temp Directory&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Isolation&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Strong — separate file system, network namespace, process tree&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Medium — separate user, but same kernel, same network&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Weak — same user, same network, just a disposable directory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Setup effort&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Steeper — need Docker installed, image to build, mounts to configure&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Moderate — one command to create the user, then you&#39;re set&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Minimal — one line to create the temp dir&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Network control&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Full — none, all, or per-host&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;None — same network as the host user&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Best for&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Production autonomous tasks where mistakes would be costly&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Development and testing — good enough isolation for learning&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Quick experiments where file system protection is enough&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-sandbox-checklist&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Sandbox Configuration Checklist&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before you run any autonomous agent, go through this checklist. Copy it, print it, stick it on your wall — whatever works:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; color: var(--text-primary);&quot;&gt;Sandbox Checklist&lt;/p&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong style=&quot;color: var(--text-primary);&quot;&gt;File mounts defined&lt;/strong&gt; — only the directories the agent needs, with &lt;code class=&quot;accent&quot;&gt;:ro&lt;/code&gt; on inputs&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong style=&quot;color: var(--text-primary);&quot;&gt;Network access scoped&lt;/strong&gt; — start with &lt;code class=&quot;accent&quot;&gt;--network none&lt;/code&gt;, add hosts only as needed&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong style=&quot;color: var(--text-primary);&quot;&gt;Max runtime set&lt;/strong&gt; — how long can the agent run before being killed?&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong style=&quot;color: var(--text-primary);&quot;&gt;Max cost set&lt;/strong&gt; — harness-level budget so a runaway loop can&#39;t drain your API credits&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong style=&quot;color: var(--text-primary);&quot;&gt;Resource limits applied&lt;/strong&gt; — CPU and memory caps so the agent can&#39;t starve your system&lt;/div&gt;
&lt;div style=&quot;padding: 0.5rem 0;&quot;&gt;&lt;span style=&quot;color: var(--cta); font-weight: 600;&quot;&gt;☐&lt;/span&gt; &lt;strong style=&quot;color: var(--text-primary);&quot;&gt;Cleanup plan ready&lt;/strong&gt; — containers and temp directories cleaned when the task is done&lt;/div&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Build the Dockerfile&lt;/strong&gt; from this lesson and run a container with &lt;code class=&quot;accent&quot;&gt;--network none&lt;/code&gt;. Inside the container, try to run &lt;code class=&quot;accent&quot;&gt;curl&lt;/code&gt; or &lt;code class=&quot;accent&quot;&gt;ping&lt;/code&gt;. Watch it fail — that&#39;s the sandbox working.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Then &lt;strong&gt;mount a directory&lt;/strong&gt; with &lt;code class=&quot;accent&quot;&gt;:ro&lt;/code&gt; and try to write a file to it from inside the container. See the permissions error? That&#39;s containment in action. These exercises make the abstract concept concrete: the sandbox is real, the boundaries hold, and your system is safe.&lt;/p&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; What Is an Autonomous Agent Loop?&lt;/a&gt;&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-3/&quot; style=&quot;color: var(--cta);&quot;&gt;Hermes: Setup and First Loop&lt;/a&gt; — install a real autonomous agent framework and run your first loop in under 20 minutes.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 10: When NOT to Use an Autonomous Agent</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 10&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;When NOT to Use an Autonomous Agent&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;After nine lessons on building autonomous agents, here&#39;s the most important one: knowing when not to use one.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#the-autonomy-continuum&quot; style=&quot;color: var(--cta);&quot;&gt;The Autonomy Continuum&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#walkthrough-three-tasks&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Three Tasks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#task-a-personalized-client-emails&quot; style=&quot;color: var(--cta);&quot;&gt;Task A: Personalized Client Emails&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#task-b-filing-taxes&quot; style=&quot;color: var(--cta);&quot;&gt;Task B: Filing Taxes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#task-c-server-log-monitoring&quot; style=&quot;color: var(--cta);&quot;&gt;Task C: Server Log Monitoring&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#compare-judgment-vs-persistence&quot; style=&quot;color: var(--cta);&quot;&gt;Compare: Judgment vs. Persistence&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#deliverable-autonomy-level-selector&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: Autonomy Level Selector&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-10/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Be able to evaluate any task and choose the correct autonomy level — including the decision to not use an autonomous agent at all&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 15 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Familiarity with all nine previous lessons&lt;/li&gt;
&lt;li&gt;A task you&#39;re considering automating this week&lt;/li&gt;
&lt;li&gt;A willingness to admit some tasks shouldn&#39;t be automated&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-autonomy-continuum&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Autonomy Continuum&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;We introduced the autonomy staircase back in Lesson 6 of the Basic Agents track. Now that we&#39;ve built actual loops, let&#39;s extend it. There are five levels of autonomy, and each one is the right choice for a specific kind of task:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Level&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Name&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Good For&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Terrible For&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;1&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Chat&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Quick questions, brainstorming, drafting&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Multi-step tasks, anything requiring persistent context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;2&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Chat with tools&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Research, code generation, analysis with live tools&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Tasks requiring sequenced multi-step execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;3&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Interactive agent&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Controlled multi-step tasks with user guidance at each step&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;High-volume or time-sensitive tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;4&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Guided loop&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Multi-step tasks with approval gates at key decisions&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Tasks needing constant supervision at every step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;5&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Fully autonomous loop&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Monitoring, batch processing, well-defined repetitive tasks&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Creative, ethical, or high-stakes decisions&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The common mistake we see is jumping straight to Level 5 for every task. A task is interesting, so it must need an autonomous loop, right? Wrong. Most tasks belong at Levels 1-4. Only a narrow slice — tasks that are well-defined, reversible, and benefit from persistence over judgment — belong at Level 5.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-three-tasks&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: Three Tasks&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s evaluate three real tasks and decide the right autonomy level for each. The answers might surprise you.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;task-a-personalized-client-emails&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Task A: Email My Top 10 Clients a Personalized Update&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Task&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;We have 10 long-term clients. We want to send each one a personalized email: thank them for their business, mention something specific about their recent project, and suggest a service they might need next.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Wrong Approach&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let the agent draft and send all 10 emails autonomously. The agent reads the client files, generates personalized content, and fires off 10 emails. Done in 5 minutes.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;What goes wrong:&lt;/strong&gt; The agent misremembers a client&#39;s project details and references work that was done for a different client. One email implies a discount we never offer. Another uses a tone that&#39;s too informal for that particular client relationship. The clients receive the emails, some are confused, one is offended. We spend the next week doing damage control.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Right Approach: Level 4 (Guided Loop)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent drafts each email based on the client data, but &lt;strong&gt;pauses before sending&lt;/strong&gt; for human approval. We review each draft, correct the tone, fix any inaccuracies, and only then approve the send.&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Openclaw Fleet: approval-required task per email
tasks:
  - task: &quot;Draft email for Client A&quot;
    approval_required: true
  - task: &quot;Draft email for Client B&quot;
    approval_required: true
  # ... repeat for all 10 clients&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;This takes longer — 20 minutes instead of 5 — but every email is correct and appropriate. The agent handles the repetitive part (generating the draft around a template) while we provide the judgment part (checking accuracy and tone). That&#39;s the right division of labor.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;task-b-filing-taxes&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Task B: File My Taxes&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Task&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Collect receipts, categorize expenses, calculate deductions, fill in tax forms, and submit to the tax authority.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Wrong Approach&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let a Level 5 autonomous agent handle everything. It has access to the accounting system, receipts folder, and the tax filing portal. It processes everything and submits — no human review needed. Maximum efficiency.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;What goes wrong:&lt;/strong&gt; The agent miscategorizes a large expense, triggering an audit flag. It misses a deduction we&#39;re entitled to because it didn&#39;t understand our specific situation. It files the return, and by the time we discover the errors, we&#39;re already in an audit process that costs $2,000 in accountant fees to resolve.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Right Approach: Level 2-3 (Assisted, Never Autonomous)&lt;/h3&gt;
&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Never file taxes autonomously.&lt;/strong&gt; The stakes are too high — incorrect filings carry legal and financial consequences. Use the agent as a research assistant: &quot;What deductions are available for home office expenses?&quot; or &quot;Categorize these 50 receipts and I&#39;ll verify.&quot; But the filing itself, the numbers, and the submission — those stay human-controlled.&lt;/p&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Level 2 or 3 is the highest we&#39;d go for tax work. The agent can help gather, organize, and draft — making the human accountant more efficient. But the final numbers and the submission come from a person, not a loop. Even a guided loop (Level 4) is too autonomous for tax filing because small errors compound into large liabilities.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;task-c-server-log-monitoring&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Task C: Monitor Server Logs for Errors and Alert Me&lt;/h2&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Task&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Watch a stream of server logs 24/7. When certain error patterns appear (5xx codes, database connection failures, disk space warnings), send an alert with the relevant context.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Wrong Approach&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;Have a human sit and watch the logs. Boring, error-prone, expensive. Or skip it entirely — wait for a client to report the problem.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;The Right Approach: Level 5 (Fully Autonomous)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;This is what Level 5 was made for. The task is well-defined (match error patterns), the action is simple (send an alert), the stakes per action are low (a false alert is annoying but harmless), and the value comes from persistence — watching 24/7 without getting tired or distracted.&lt;/p&gt;

&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;# Hermes loop for log monitoring
goal:
  description: &gt;
    Every 60 seconds, read the last 100 lines of /var/log/app.log.
    If any line matches error patterns (5xx, &quot;ERROR&quot;, &quot;FATAL&quot;,
    &quot;disk full&quot;, &quot;connection refused&quot;), send an alert with the
    matching lines and timestamp.
  success_criteria:
    - &quot;Alerts are sent within 90 seconds of an error appearing&quot;
    - &quot;No false alerts for known transient errors&quot;
    - &quot;Each alert includes the 5 lines before and after the error&quot;

constraints:
  - &quot;If the same error appears in consecutive checks, send one alert per minute — do not flood&quot;
  - &quot;Known transient errors (connection reset by peer, timeout retry) do not trigger alerts&quot;&lt;/pre&gt;&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The agent runs 24/7, costs pennies per day in API calls, and never misses an error. A human would need to watch the screen constantly — impossible for more than a few minutes. The autonomous loop is strictly better at this task.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;compare-judgment-vs-persistence&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Compare: Judgment vs. Persistence&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;The three tasks reveal a clear dividing line:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Task&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Needs Judgment&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Needs Persistence&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Right Level&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Client emails&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;High — tone, accuracy, relationship context&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Low — 10 emails, one batch&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;4 (guided loop)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Tax filing&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Very high — legal liability, nuanced rules&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Low — once a year&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;2-3 (assisted only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Log monitoring&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Low — pattern matching, concrete rules&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Very high — 24/7 continuous&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;5 (fully autonomous)&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;High-judgment, low-persistence tasks&lt;/strong&gt; are where autonomous loops underperform. The agent can&#39;t match a human&#39;s understanding of client relationships, company culture, or nuanced domain knowledge. These tasks are better served by Levels 2-4: use the agent as a tool, not a replacement.&lt;/p&gt;

&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Low-judgment, high-persistence tasks&lt;/strong&gt; are where autonomous loops shine. The task is well-defined, the decisions are clear-cut, and the value comes from sustained attention. These are the tasks for Level 5.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-autonomy-level-selector&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: Autonomy Level Selector&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Before you build any agent loop, ask these three questions. They&#39;ll tell you the right autonomy level — and sometimes the answer is &quot;don&#39;t build one at all.&quot;&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;

&lt;div style=&quot;padding: 1rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0 0 0.5rem 0; font-weight: 600; font-size: 1rem;&quot;&gt;Question 1: Can the task fail safely?&lt;/p&gt;
&lt;p style=&quot;margin: 0 0 0.5rem 0; color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;If the agent makes a mistake, what&#39;s the worst outcome? A wrong file name — safe. A wrong email to a client — risky. A wrong tax filing — dangerous.&lt;/p&gt;
&lt;p style=&quot;margin: 0; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;&lt;strong&gt;Answer Yes → proceed to Q2. Answer No → max Level 3 (human approves every output).&lt;/strong&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;padding: 1rem 0; border-bottom: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;margin: 0 0 0.5rem 0; font-weight: 600; font-size: 1rem;&quot;&gt;Question 2: Do I need to approve intermediate results?&lt;/p&gt;
&lt;p style=&quot;margin: 0 0 0.5rem 0; color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;Does the task have steps where a wrong decision early in the process would waste all the work that follows? If yes, you need checkpoints between those steps.&lt;/p&gt;
&lt;p style=&quot;margin: 0; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;&lt;strong&gt;Answer Yes → Level 4 (guided loop with approval gates). Answer No → proceed to Q3.&lt;/strong&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;padding: 1rem 0;&quot;&gt;
&lt;p style=&quot;margin: 0 0 0.5rem 0; font-weight: 600; font-size: 1rem;&quot;&gt;Question 3: Is the success criteria unambiguous?&lt;/p&gt;
&lt;p style=&quot;margin: 0 0 0.5rem 0; color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;Can you write the success criteria in a way that the agent can check its own work and know when it&#39;s done? &quot;Every file in this directory is sorted into a subfolder&quot; — unambiguous. &quot;Write a compelling marketing email&quot; — not unambiguous.&lt;/p&gt;
&lt;p style=&quot;margin: 0; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;&lt;strong&gt;Answer Yes → Level 5 (fully autonomous). Answer No → Level 2-3 (interactive or chat with tools).&lt;/strong&gt;&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The three questions map to the three risk categories from Lesson 7: safety (Q1), checkpoints (Q2), and goal clarity (Q3). If a task passes all three, Level 5 is appropriate. If it fails any one, drop down a level. If it fails Q1, never go above Level 3 — the consequences of full autonomy are too severe.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;&lt;strong&gt;Take a task you&#39;re considering automating.&lt;/strong&gt; Run it through the Autonomy Level Selector. What&#39;s the right level? If it&#39;s Level 5, great — use the patterns from Lesson 8 and the build steps from Lesson 9. If it&#39;s Level 2-4, don&#39;t force it into Level 5. Build the right loop for the right task.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Then think about a task you already automated. Run it through the selector. Were you using the right level? If not, what would you change? The most common mistake we see is Level 5 applied to Level 2-3 tasks, resulting in expensive, unreliable agents that could have been replaced by a well-structured chat conversation. Don&#39;t let that be you.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;display: flex; justify-content: space-between; align-items: center; flex-wrap: wrap; gap: 1rem; margin-top: 2rem;&quot;&gt;
&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-9/&quot; style=&quot;color: var(--cta); font-size: 0.9375rem;&quot;&gt;&amp;larr; Building Your First Loop&lt;/a&gt;
&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next track:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/vibe-coding/&quot; style=&quot;color: var(--cta);&quot;&gt;Basic Vibe Coding&lt;/a&gt; — from prompting agents to creating software with guided AI collaboration.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry><entry>
    <title>Lesson 1: What Is an Autonomous Agent Loop?</title>
    <link href="https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/"/>
    <updated>Thu, 01 Jan 2026 00:00:00 +0000</updated>
    <id>https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/</id>
    <content type="html">
&lt;section class=&quot;hero&quot; style=&quot;padding-bottom: 2rem;&quot;&gt;
&lt;div class=&quot;container text-center&quot; style=&quot;max-width: 960px;&quot;&gt;
&lt;p class=&quot;meta&quot;&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/&quot; class=&quot;accent&quot;&gt;Agents Working for You / Lesson 1&lt;/a&gt;&lt;/p&gt;
&lt;h1 style=&quot;max-width: 960px; margin: 0 auto;&quot;&gt;What Is an Autonomous Agent Loop?&lt;/h1&gt;
&lt;p class=&quot;lede&quot; style=&quot;max-width: 720px; margin: 0.5rem auto 0;&quot;&gt;You set up an agent in the Basic Agents track and watched it work. But you had to sit there. What if it could just go?&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;

&lt;section class=&quot;section section-flush&quot;&gt;
&lt;div class=&quot;container container-narrow&quot;&gt;
&lt;div class=&quot;glass-card&quot; style=&quot;padding: 3rem;&quot;&gt;

&lt;nav aria-label=&quot;On this page&quot; style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin-bottom: 2rem;&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.5rem; font-size: 0.875rem; text-transform: uppercase; letter-spacing: 0.05em; color: var(--text-muted);&quot;&gt;On this page&lt;/p&gt;
&lt;ul style=&quot;list-style: none; padding: 0; margin: 0; line-height: 2;&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/#what-youll-need&quot; style=&quot;color: var(--cta);&quot;&gt;What You&#39;ll Need&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/#the-microwave-and-the-slow-cooker&quot; style=&quot;color: var(--cta);&quot;&gt;The Microwave and the Slow Cooker&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/#the-loop&quot; style=&quot;color: var(--cta);&quot;&gt;The Loop: Perceive → Decide → Act → Repeat&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/#walkthrough-same-task-two-modes&quot; style=&quot;color: var(--cta);&quot;&gt;Walkthrough: Same Task, Two Modes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/#interactive-vs-autonomous&quot; style=&quot;color: var(--cta);&quot;&gt;Interactive vs Autonomous&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/#deliverable-the-autonomy-flowchart&quot; style=&quot;color: var(--cta);&quot;&gt;Deliverable: The Autonomy Flowchart&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-1/#try-it&quot; style=&quot;color: var(--cta);&quot;&gt;Try It&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/nav&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1rem 1.25rem; margin-bottom: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-muted); font-size: 0.8125rem; text-transform: uppercase; letter-spacing: 0.05em; font-weight: 600; margin-bottom: 0.25rem;&quot;&gt;Goal&lt;/p&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem; margin: 0;&quot;&gt;Understand the perceive→decide→act→repeat loop and know when a task is ready for autonomy&lt;/p&gt;
&lt;/div&gt;

&lt;h2 id=&quot;what-youll-need&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;What You&#39;ll Need&lt;/h2&gt;
&lt;p class=&quot;recipe-time&quot;&gt;&lt;strong&gt;⏱ Time:&lt;/strong&gt; 12 minutes &amp;nbsp;|&amp;nbsp; &lt;strong&gt;📋 Tasks:&lt;/strong&gt;&lt;/p&gt;
&lt;ul style=&quot;color: var(--text-secondary); line-height: 2.2; margin-bottom: 1.5rem; padding-left: 1.5rem; list-style: disc;&quot;&gt;
&lt;li&gt;Experience with an interactive agent from the &lt;a href=&quot;https://clocklobster.com/blog/tutorials/basic-agents/lesson-1/&quot; style=&quot;color: var(--cta);&quot;&gt;Basic Agents&lt;/a&gt; track&lt;/li&gt;
&lt;li&gt;A task in mind that you currently do manually or with an interactive agent&lt;/li&gt;
&lt;/ul&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-microwave-and-the-slow-cooker&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Microwave and the Slow Cooker&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;An interactive agent — the kind from Basic Agents — is like a microwave. We stand there watching it. We approve each step. The agent says &quot;I found three files to change, proceed?&quot; and we click yes. Then it says &quot;Tests pass, continue?&quot; and we click yes again. We&#39;re in the loop, every step, all the way through.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;That&#39;s fine for short tasks. But for anything that takes more than a few minutes, standing in front of the microwave gets old fast.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;An autonomous agent is the slow cooker. We load the ingredients, set the timer, and walk away. The agent works through the task — reading, deciding, acting, checking results — without pulling us in for approval at every step. We come back when it&#39;s done, or when it hits something it genuinely can&#39;t figure out alone.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;Both cook food. But they change &lt;em&gt;our&lt;/em&gt; role from operator to delegator. That shift is what this track is about.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;the-loop&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;The Loop: Perceive → Decide → Act → Repeat&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Every autonomous agent runs the same fundamental loop. It doesn&#39;t matter whether it&#39;s Hermes, AutoGPT, Openclaw fleet mode, or a custom script. The loop is always:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Phase&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;What the agent does&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Perceive&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Gather information from tools, files, or APIs&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Read the CSV of competitor URLs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Decide&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Choose the next action based on the goal and what it learned&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&quot;I&#39;ll fetch the first competitor&#39;s homepage next&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Act&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Execute the chosen action using a tool&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Run a web fetch against the competitor&#39;s URL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Repeat&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Check if the goal is met. If not, loop back to perceive.&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;&quot;2 of 5 competitors done. Fetch the next one.&quot;&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;The key difference from an interactive agent: in the interactive model, &lt;strong&gt;we&lt;/strong&gt; are the Decide phase. The agent perceives and acts, but we approve every choice. In the autonomous model, the agent handles the entire cycle itself. We set the goal upfront and review the result afterward.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;walkthrough-same-task-two-modes&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Walkthrough: Same Task, Two Modes&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Let&#39;s watch the same task play out in both modes: &lt;strong&gt;Research three competitors and compile a comparison report.&lt;/strong&gt;&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Interactive Mode (Microwave)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;We run an interactive agent and type:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Research competitor A. Read their homepage and their pricing page.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent fetches the pages. It shows us the results and asks: &lt;em&gt;&quot;I found these features. Should I continue to competitor B?&quot;&lt;/em&gt; We click yes. It fetches B. Shows us the results. Asks again. Yes again. Fetches C. Shows us. Asks if we want a comparison table. Yes. It writes the table. Asks if we&#39;re satisfied. We review, maybe ask for tweaks. Click. Click. Click.&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;The agent did the work, but &lt;strong&gt;we stayed in the loop&lt;/strong&gt;. Every step required our attention. For three competitors it&#39;s manageable. For thirty? Not a chance.&lt;/p&gt;

&lt;h3 style=&quot;margin-bottom: 0.75rem; margin-top: 1.5rem;&quot;&gt;Autonomous Mode (Slow Cooker)&lt;/h3&gt;
&lt;p class=&quot;mb-4&quot;&gt;We write a task file or prompt:&lt;/p&gt;
&lt;div class=&quot;code-block&quot;&gt;&lt;pre&gt;Research these 3 competitors: [URLs]. For each, read the homepage and pricing page. Build a comparison table covering: pricing model, key features, target audience, and what makes them different from each other. Save the table as competitor-comparison.md. If any site is unreachable, note it and continue.&lt;/pre&gt;&lt;/div&gt;
&lt;p class=&quot;mb-4&quot;&gt;We start the agent. Then we walk away. Ten minutes later we open &lt;code class=&quot;accent&quot;&gt;competitor-comparison.md&lt;/code&gt; and find a complete table. The agent fetched each site, extracted the relevant information, compared across competitors, and wrote the file — all without a single approval prompt.&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.25rem 1.5rem; margin: 1.5rem 0; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Key insight:&lt;/strong&gt; The model and tools are the same in both modes. What changed is &lt;strong&gt;decision ownership&lt;/strong&gt;. In interactive mode, we owned the decisions. In autonomous mode, the agent did. That shift is powerful — and it demands more careful goal-setting upfront.&lt;/p&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;interactive-vs-autonomous&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Interactive vs Autonomous&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Here&#39;s how the two modes stack up against each other:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;table style=&quot;width: 100%; border-collapse: collapse; font-size: 0.9375rem;&quot;&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Dimension&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Interactive Agent&lt;/th&gt;
&lt;th style=&quot;text-align: left; padding: 0.75rem 0.5rem; color: var(--cta); font-family: var(--font-header);&quot;&gt;Autonomous Agent&lt;/th&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Oversight level&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Approve every step&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Set goal, review result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Task duration&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Limited by your attention span&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Can run for hours unattended&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Cost pattern&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;You control each step, so cost is predictable&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Can run away if goal is poorly defined&lt;/td&gt;
&lt;/tr&gt;
&lt;tr style=&quot;border-bottom: 1px solid var(--border);&quot;&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Failure mode&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Agent stops and asks for help&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Agent may go down a wrong path for a while before recovering&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-primary); font-weight: 600;&quot;&gt;Best use case&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Tasks where every decision matters (editing live code, handling sensitive data)&lt;/td&gt;
&lt;td style=&quot;padding: 0.75rem 0.5rem; color: var(--text-secondary);&quot;&gt;Tasks where intermediate steps are reversible (research, batch processing, report generation)&lt;/td&gt;
&lt;/tr&gt;
&lt;/table&gt;
&lt;/div&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;deliverable-the-autonomy-flowchart&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Deliverable: The Autonomy Flowchart&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Not every task is ready for autonomy. Before we set an agent loose, we run through three questions. If any answer is no, we start in interactive mode and consider whether the task can be restructured:&lt;/p&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-bottom: 2rem; border: 1px solid var(--border);&quot;&gt;
&lt;p style=&quot;font-weight: 600; margin-bottom: 0.75rem; color: var(--text-primary);&quot;&gt;Autonomy Decision Flowchart&lt;/p&gt;
&lt;div style=&quot;padding: 0.75rem 1rem; background: var(--bg); border-radius: 8px; margin-bottom: 0.75rem;&quot;&gt;
&lt;p style=&quot;margin: 0; color: var(--text-secondary);&quot;&gt;&lt;strong&gt;Q1:&lt;/strong&gt; Can we write clear, testable success criteria?&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;If yes: proceed. If no: the goal is too vague. Break it down until you can describe &quot;done.&quot;&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.75rem 1rem; background: var(--bg); border-radius: 8px; margin-bottom: 0.75rem;&quot;&gt;
&lt;p style=&quot;margin: 0; color: var(--text-secondary);&quot;&gt;&lt;strong&gt;Q2:&lt;/strong&gt; Can the task fail safely?&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;If yes: proceed. If the worst-case failure costs money, corrupts data, or sends emails to customers, don&#39;t run it autonomously.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.75rem 1rem; background: var(--bg); border-radius: 8px; margin-bottom: 0.75rem;&quot;&gt;
&lt;p style=&quot;margin: 0; color: var(--text-secondary);&quot;&gt;&lt;strong&gt;Q3:&lt;/strong&gt; Do we need to approve intermediate steps, or just the final result?&lt;/p&gt;
&lt;p style=&quot;margin: 0.25rem 0 0 1.25rem; color: var(--text-muted); font-size: 0.875rem;&quot;&gt;If just the result: go autonomous. If every step needs a sign-off, stay interactive.&lt;/p&gt;
&lt;/div&gt;
&lt;div style=&quot;padding: 0.75rem 1rem; background: var(--bg); border-radius: 8px;&quot;&gt;
&lt;p style=&quot;margin: 0; color: var(--text-secondary);&quot;&gt;&lt;strong&gt;Decision:&lt;/strong&gt; Three yeses → run it autonomously. Any no → run interactively, then restructure the task to get to three yeses.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;

&lt;p class=&quot;mb-4&quot;&gt;Save this flowchart — mentally or as a note. We&#39;ll come back to it in every lesson when deciding whether to run a task autonomously.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;h2 id=&quot;try-it&quot; style=&quot;margin-bottom: 1rem;&quot;&gt;Try It&lt;/h2&gt;
&lt;p class=&quot;mb-4&quot;&gt;Pick a task you do regularly — something that takes 15-30 minutes with an interactive agent. Run it through the autonomy flowchart. Did it get three yeses? If not, what would it take to get there? Write down what would need to change: a clearer goal, safer failure modes, or a different definition of &quot;done.&quot;&lt;/p&gt;
&lt;p class=&quot;mb-4&quot;&gt;That gap — between your task today and the three-yes target — is exactly what the rest of this track will help you close.&lt;/p&gt;

&lt;div style=&quot;border-top: 1px solid var(--border); margin: 2rem 0;&quot;&gt;&lt;/div&gt;

&lt;div style=&quot;background: var(--surface); border-radius: 12px; padding: 1.5rem; margin-top: 2rem; border-left: 3px solid var(--cta);&quot;&gt;
&lt;p style=&quot;color: var(--text-secondary); font-size: 0.9375rem;&quot;&gt;&lt;strong&gt;Next lesson:&lt;/strong&gt; &lt;a href=&quot;https://clocklobster.com/blog/tutorials/agents-working-for-you/lesson-2/&quot; style=&quot;color: var(--cta);&quot;&gt;Sandboxing: Why and How&lt;/a&gt; — because once an agent runs without you, you need to know it can&#39;t break things while you&#39;re gone.&lt;/p&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
  </entry>
</feed>
