Video: "Hermes Agent OS Just Changed AI Agents Forever!" by Julian Goldie on YouTube.
What Agent OS is, briefly
Agent OS is Julian Goldie's term for the full stack he runs around Hermes Agent — a dashboard that acts as mission control for multiple AI models working in parallel. Rather than opening Claude for writing, a separate tool for research, and another for coding, Agent OS routes each task to the model best suited for it, manages a shared memory layer, and reports back when the work is done.
The previous iteration ran Claude alongside Hermes v0.19 Quicksilver. This video covers the updated stack, now incorporating Hermes v0.20 Herald — which added real-time voice, A2A agent-to-agent communication, and grounded citations when it shipped on 3 August 2026. If you want context on the v0.20 release itself, we covered it here: Hermes v0.20: say "Hey Hermes", it wakes up — then the rest of the release is equally interesting.
The stack: which model does which job
The Agent OS Julian shows in this video runs four specialist agents alongside the core Hermes and Claude combination.
Hermes Oracle handles web research and SEO work — trend spotting, keyword analysis, competitor monitoring, and drafting content from live data. It is the part of the stack that watches the web so the human does not have to.
Studio is the visual output layer, built on the Higgsfield MCP. When any other agent in the stack needs an image or video, Studio generates it — 4K images and cinematic clips — using Higgsfield's roster of over 30 models (Kling, Seedream, Veo, and others). The rest of the stack does not need to leave the workflow to produce visual content.
GLM 5.2 handles coding tasks. When the OS needs to write a script, fix a configuration, or automate something technical, it routes to GLM rather than burning a more expensive model on work that does not require it.
Open Montage handles longer-form video production — structured, edited AI film rather than short clips. It sits alongside Studio for projects that need a narrative arc rather than a single generated asset.
A shared Obsidian memory layer sits underneath all of this. Every agent reads from and writes to the same memory store, so context built during a research task is available when the content agent picks up the job, without you having to copy anything across.
What self-improving means here
The new element in this video is that the stack can now modify itself. When the system identifies that a prompt is producing weak results, that a routing rule is sending the wrong type of task to the wrong model, or that a new MCP would cover a gap in its current capabilities, it can draft the fix, test it in a sandboxed environment, and deploy it — without the human writing a single configuration line.
Julian frames this with a specific instruction: "Stop collecting AI tools. Build your own AI machine." The distinction he is drawing is between accumulating individual subscriptions to individual tools and building a single system that improves as it runs. Self-improvement is the mechanism that makes the second option viable at scale: the system's configuration gets better the more it is used, rather than staying static until the human updates it manually.
In practice, this means the Agent OS you have running at the end of a month of heavy use will make better routing decisions and produce better output than the one you set up at the start, without requiring manual tuning sessions.
What an instruction actually looks like
Julian's framing throughout is that the system should only surface itself to the human when it genuinely cannot proceed alone. For everything else — picking the right model, writing the prompt, running the task, filing the output, updating memory — it works without interruption.
A single plain-English instruction like "find the top three trending keywords in AI SEO this week, write a 1,200-word article on the strongest one, and generate a header image" routes automatically: Oracle runs the research, Hermes drafts the article with citations via the v0.20 Herald grounded-citations skill, Studio generates the image through Higgsfield. The human sees the completed output. The routing decisions and model calls happen without prompts for confirmation.
The Hermes v0.20 voice layer adds another layer to this: if you would rather speak the instruction than type it, saying "Hey Hermes" opens a session and the stack takes it from there. The output still lands in the same place.
Where this connects to NordSys
Setting up and maintaining an Agent OS stack — getting the model routing right, connecting MCPs like Higgsfield, configuring memory, and keeping the whole thing running reliably — is the kind of technical setup that looks straightforward in a video and takes considerably longer when you are doing it against your own workflows for the first time. Helping businesses get there without spending a week on configuration is what our AI Agents service covers. If you want an Agent OS running in your business, or an existing Claude Code or Hermes setup that is not working as smoothly as it should be, our AI Agents service is the right place to start.
See our AI Agents →