Video: "Hermes Agent Update v0.18 is HUGE! (Judgment Release)" by Julian Goldie on YouTube.
What the /learn command does
Before v0.18, saving a workflow as a skill in Hermes required editing skill files manually outside the session — a reasonable step for developers comfortable with configuration files, but a barrier for non-technical users building repeatable business processes. The /learn command removes that barrier. You run a task, it works, and you issue /learn immediately to save the sequence as a named skill that any future session can call.
The practical effect is that Hermes can now capture institutional knowledge from a live session rather than requiring it to be written up separately afterward. An operator who finds a good prompt sequence for a recurring task — processing inbound briefs, formatting weekly reports, running a content check — can save it in the same session where they discovered it works. The skill is available from the next run onward.
What the Judgment layer adds to output verification
The Judgment layer is a built-in verification pass that runs after an agent completes a task and before the result is returned. It checks the output against the original brief — flagging where the agent has misunderstood, missed a requirement, or produced output that technically answers the question but does not match the intent.
This matters in practice because agents frequently produce plausible-looking output that does not fully satisfy the brief. Without a verification step, catching that requires the operator to read and assess every output — which eliminates much of the time saving that agents are supposed to deliver. The Judgment layer automates that review pass, re-queuing the task if the output fails the check. Julian Goldie demonstrated a case where the initial agent output was structurally correct but missed a key constraint, and the Judgment layer identified and flagged it before the result was accepted.
Reliability improvements for long-running tasks
Multi-step agent runs — the kind that coordinate several sub-agents, write to multiple files, or execute a pipeline over a list of inputs — have historically been a weak point for Hermes. The agent would complete most of the task, then fail late in the run due to context window issues, malformed intermediate output, or a sub-agent call that returned an unexpected format. v0.18 addresses several of the most common failure patterns, with particular attention to context management across long chains and better error recovery when a sub-agent call fails partway through.
The improvement is in reliability rate, not in fundamentally new behaviour — Hermes could always attempt long tasks, but completion rates on complex runs were variable enough that operators often split tasks manually as a precaution. The aim of v0.18 is to make that precaution unnecessary for the majority of practical use cases.
How v0.18 sits alongside v0.17
Hermes v0.17 — The Reach Release — shipped background subagents, iMessage and WhatsApp integrations, and a redesigned desktop app. v0.18 does not add new surface area in the same way; it is focused on depth rather than breadth. The additions in v0.18 make the capabilities introduced in earlier releases more reliable and more accessible to non-technical operators. The /learn command in particular is a direct response to feedback that skill creation was too technical for the business users Hermes is trying to reach.
Where this connects to NordSys
If tracking this kind of AI-agent news makes you wonder whether one could actually run inside your business, that's exactly what our AI Agents do — named, briefed and managed for you, no setup fee, from £6 a day.
See our AI Agents →