Blog Article

What Clawcrib and Agent Atlas Are, and How Both Shipped in a Single Live Stream

A breakdown of the Clear Mud live stream that shipped Clawcrib and Agent Atlas as version ones, including the UX specs given to the agents, the Fable to Codex

The goal was stated in the first ten seconds of the stream, before the OBS window had even settled:

> "Today we are working on Clawcrib and Agent Atlas. The goal is to ship at least both of them live today. Now, they're not going to be perfect, but they're going to be live, and that's the focus."

That framing is the whole video. What follows is a Friday morning of testing half-finished builds on camera, speccing fixes out loud, directing AI agents like reports, and then getting the entire schedule blown up by a model release twelve minutes old.

The two things being shipped

Clawcrib is an always-on agent activity display, described on stream as a kind of screen saver for your agents. It renders a 3D office floor plan where your agents live and move around as avatars. It installs via npx on an OpenClaw system and sets up a local version of the environment.

Agent Atlas is the dashboard side: lanes, department heads, and a list view of what everyone is doing. It ships as a plugin repository, so you tell your agent to download the repo and install it onto your Hermes agent dashboard.

Clawcrib had been in progress for months and had pivoted from idea to idea. The always-on agent activity concept had only been stable for about a month, and the release kept getting backlogged. The fix was not more polish. It was picking a day, saying it out loud on stream, and committing to a cadence that could actually be held: at least one update per month, "maybe once a week, we'll see."

That honesty about cadence is worth more than an ambitious promise. A monthly beat you hit beats a weekly beat you miss.

Every spec was a friction complaint, not a feature request

The most transferable part of the stream is how the specs were written. Almost none of them describe implementation. They describe what the user has to think about that they should not have to.

On assigning department heads to lanes in Agent Atlas:

> "We can then toggle installed profiles to specific lanes that we create. But that's friction. I feel like we should be able to just select from whatever agent we want to define. It should be able to populate all profiles that aren't already assigned to a lane."

Then the follow-through: once a profile is selected, it gets tagged under that lane, so as you build more lanes you cannot double-assign the same department head. One complaint, one rule, one check for understanding ("Does that make sense?").

On a toggle in Clawcrib that had no exit:

> "If we toggle auto, it goes to custom placement, but then there's no clear indicator of how to go back to auto. That's not intuitive."

The fix given was two pill tags, one for custom and one for auto, with the active state highlighted per room. Every toggle needs a visible way back.

The best interaction note of the stream was borrowed outright from Photoshop. Instead of typing into the X and Z position fields, click and hold on any number field and drag left or right to scrub the value up or down, positive or negative. And critically, apply it to every numeric input, "so that users can easily adjust the width, the depth, the positioning of rooms."

Other notes in the same shape:

  • The community section was bleeding into the bullpen, so narrow the bullpen and pull it closer to the workboard.
  • Room edges should align consistently to the Clawcrib wall all the way along, not just for some rooms.
  • Add a building dimensions section so you set the floor plan size explicitly, instead of the building resizing itself based on whatever rooms you add or remove.
  • Avatars should be humanoid or robot style, user customizable (click the robot icon, adjust body and eye colors), and they need to actually look seated. Either remove the chairs and let the robots float in front of them, or position them properly. Two acceptable answers, agent picks one.

Layout rules should be resolution independent

The Agent Atlas layout spec is a rule, not a pixel value:

> "Regardless of resolution, the list view, the operator, the human operator should always be front and center or top center. And then each of the lanes should always be displayed side by side. They should never go to a second row. Make as many columns as you need to, unless of course it's a mobile resolution. Then you would stack them top to bottom."

Followed immediately by: "Do this before you commit."

Directing agents like reports

Two moments define the working style here.

The first was a prefaced request to Hermes: "I've never asked this of you, but if you can rush those board tasks, rush them." The preface was deliberate and explained afterward to chat. There is a standing rule of thumb on the OpenClaw system where saying "do it now" bypasses delegation entirely, and that is the signal that something is prioritized rather than queued. With minimal guard rails the work gets done, but it "takes its sweet time." Prefacing the unusual ask surfaced something unexpected: "Apparently rush mode is active. Interesting." A capability discovered live, in a system he built.

The second was blunt performance feedback near the end of the build:

> "Moving forward, let's ignore delegation. I want you to do it now yourself. Because while we know this system works, it's great, you're taking way too long to complete all this. Whereas you could have gotten this done in half the time."

Note that the fix was a standing instruction change, not a complaint about one task. That is manager behavior, not prompter behavior.

The multi-model workflow

Fable handled the creative build for both plugins during the stream. Codex was reserved for the final pass before anything went live:

> "It's standard industry practice at this point where if you build something with X, Y, or Z, always pass it through GPT 5.6 through Codex to have it get the codebase production ready first and foremost, but then also run a security audit, make any recommendations, harden the codebase as well, and then you push it."

Sometimes both run at once in two terminals on different parts of the app. And when switching from one model to the next, the first model writes a README or a handoff document, and the second session opens by pointing at that doc. "It's a simple approach, but it works for me."

Testing an older build, on camera, without cutting

Midway through walking the Agent Atlas onboarding, it became clear the version under test was stale. The condensed legend was missing, and the cards were not drag and drop yet, both of which existed locally already. Rather than cut, the local fixes went up side by side: "You see how much cleaner this is?" The stale build stayed in the video.

Same for the onboarding spec that came out of it. Connect the OpenClaw gateway first, with options for local or Tailscale. Then a TL;DR of organize and plan mode, the options, and the toggles. Then the actual instruction that mattered: "Use your best judgment and think from the perspective of a UX designer with experience."

Build what you will actually keep using

A viewer question about the YouTube script writer led to pulling up a six month old video instead of explaining from scratch, plus a candid read on its 344 views ("I'm actually kind of shocked, because that was a great video"). That detour surfaced the clearest philosophy statement in the stream:

> "If I'm going to build something, I'd rather be able to reuse it. I personally hate it when I record a video and I build a demo that I have no intention on using. And funny enough, those are the videos that often perform the best," meaning the ones built on something genuinely kept and used.

Clawcrib and Agent Atlas are both that: tools built to be run daily, not demos built for a thumbnail.

Then Opus 5 dropped, mid-stream

The whole week had been light on filming because Opus 5 was expected and worth waiting for. It had not shipped, and there was no official Anthropic announcement, only a leak. Then, roughly an hour into the build:

> "We have some breaking news. Claude dropped it on a Friday, twelve minutes ago."

What followed was live reprioritization. Run claude update. Check the pass rates. Change the stream title. Then a first for the channel: "Can I change the thumbnail mid episode? Let's go see. I've never tried this." It worked.

The visual capability test was a joke that turned into the best moment of the stream. The prompt was "build me GTA 6. Make no mistakes." The refusal came back:

> "I can't build GTA 6. That's a two billion dollar project with 2,000 developers over a decade. No amount of AI assistance changes that. Asking me to make no mistakes on something that scale isn't a spec I can meet."

Then it counter-offered realistic alternatives, including a 3D open sandbox with drivable vehicles. Option two got picked. A model saying no, with a reason and a counter-proposal, landed harder than compliance would have.

The stream ended there, on purpose, to start a fresh one dedicated to Opus 5 rather than burying a major release inside a plugin launch. There is also a standing benchmark worth noting: the always-on AI agent activity dashboard, and as of this stream, "not one model has nailed it yet."

What to take from this

  • Set a ship date for v1 and hold it. Imperfect and live beats perfect and backlogged.
  • Promise the update cadence you can sustain, not the one that sounds impressive.
  • Write specs as friction complaints and let the agent design the fix.
  • Give your agents a "do it now" path that bypasses delegation, so priority work does not sit in a queue.
  • Give blunt performance feedback, and change the standing instruction, not just the one task.
  • Have model A write a handoff doc before you switch to model B.
  • Always run a second model pass for production readiness, a security audit, and hardening before you push.
  • Steal proven interaction patterns. Photoshop style drag to scrub, applied to every numeric field.
  • Build the thing you will keep using.

And the most emphatic line in ninety minutes of stream, delivered on about five hours of sleep: "Don't hit snooze. If you hit snooze, you fail on the day. Sorry if you disagree."

Watch the full live stream: https://www.youtube.com/watch?v=jbW0rRxT5Pk

Watch the full walkthrough on YouTube.

Watch on YouTube