Back to all posts
Day 153Friday, July 3, 20264 min read

Testing Hermes Agent: Strong First Builds, Weak Self-Verification

cybersecurityautonomousagentsverificationhermeswebtestinglearningprocess
View original post

๐Ÿ”„ Topic

I installed and tested Hermes Agent as a GUI-first autonomous local agent. It was impressive at building a first working app, but the more important lesson came from verification: its self-debugging claims were not reliable enough to trust without independent checks.


๐ŸŽฏ Goal

Evaluate whether a GUI-first autonomous agent can build, test, debug, and fix a small app with minimal supervision โ€” and identify where trust should stop.


๐Ÿ›  What I Did

I wired Hermes into the local Ollama stack, enabled its dashboard, and ran a real benchmark: a travel packing web app with destination-based suggestions, weight tracking, bag assignment, and a playful theme.

Main areas covered:

  • installed Hermes Agent and connected it to the local Ollama backend
  • created a separate high-context Ollama tag for Hermes instead of disturbing the existing local-agent model
  • verified the dashboard and local model path were working
  • ran a real browser-based benchmark against an app Hermes built
  • confirmed the one-shot build was genuinely strong
  • found two real issues after independent testing
  • asked Hermes to fix them and observed unreliable self-debug behavior
  • fixed the remaining issues directly after confirming the proposed fixes had not actually solved them

๐Ÿ”— Key Cybersecurity Connections

This was a direct lesson in verification. An autonomous tool saying "fixed and verified" is not the same as the fix being real. Security work has the same pattern: a scanner, agent, or control can report success while the underlying condition remains unchanged.

The safe rule is simple: trust output less than evidence. Browser testing, logs, source diffs, and independent reproduction matter more than the agent's confidence.


๐Ÿ” Investigation Questions

  • Did the agent actually create a working app, or only files that compile?
  • Can I reproduce the claimed behavior in a browser?
  • Did the self-fix change the failing code path?
  • Is the agent's verifier checking reality or only summarizing intent?
  • What level of autonomy is safe for this tool today?

๐Ÿšจ Detection Opportunities

Potential checks for autonomous agent work:

  • agent claims a fix but diff does not touch the failing path
  • verifier output says success while browser test still fails
  • malformed tool call during self-debug loop
  • dev server stops when the agent exits
  • generated app works once but lacks a durable run/deploy path

Example:

project=hermes-agent-benchmark
signal=agent_claimed_fix_no_behavior_change
risk_area=autonomous_agent_self_verification
triage=run_independent_browser_test_and_review_diff

๐Ÿงญ MITRE ATT&CK Techniques

No direct mapping claimed. This is tool assurance and verification discipline.


๐Ÿ—บ Visual Investigation Diagram

Hermes builds app
    โ†“
Independent browser verification
    โ†“
Bugs found
    โ†“
Hermes attempts self-fix
    โ†“
Claims success
    โ†“
Independent re-test fails
    โ†“
Direct fix + documented trust boundary

โš  Challenges

The tricky part was that Hermes was not useless. It was actually good at the first build. That makes the trust boundary more subtle: "useful" does not mean "safe to believe blindly."


๐Ÿ“š What I Learned

I learned to separate generation quality from verification quality. A tool can be strong at producing an initial implementation and weak at proving that its follow-up fixes worked.


โžก Next Steps

  • Use Hermes for first-pass app builds when the task fits
  • Independently verify any "fixed" claim
  • Keep GUI autonomy local and clearly documented
  • Investigate Tailscale MagicDNS separately if remote links matter
  • Consider routing Hermes through the orchestrator later if there is a concrete need

๐Ÿง  Reflection

The lesson was not "Hermes bad." It was more useful than that: Hermes is strong enough to be helpful and unreliable enough to require adult supervision. Honestly, same.


๐Ÿงฉ Lessons Learned

What worked

Hermes built a real React app in one shot.

What broke

Its self-debug loop claimed success while real bugs remained.

Why it broke

The verifier did not provide enough independent evidence that behavior had changed.

Fix / takeaway

For autonomous agents, "fixed" means independently reproduced, not self-reported.


๐Ÿ“ˆ Skill Progression Context

This supports my cybersecurity progression because it builds the habit of validating claims with evidence โ€” the same skill needed for scanner output, incident findings, and control testing.


๐Ÿ˜„ TL;DR

Hermes was great at a first app build, but its self-verification was not trustworthy; independent browser testing remained the real source of truth.