Back to all posts
Day 139Friday, June 19, 20264 min read

Benchmarking a Local LLM Coding Stack: Harness, Routing, and Review Findings

cybersecuritylocalaibenchmarkingcodereviewreliabilitylearningprocess
View original post

๐Ÿ”„ Topic

Building and benchmarking the local LLM coding stack โ€” a test harness, routing rules, and fixing what code review found.


๐ŸŽฏ Goal

Know, with evidence, which local model to trust for which coding task โ€” and clean up the reliability bugs review surfaced.


๐Ÿ›  What I Did

I documented and benchmarked the local coding stack: a harness for running models against real tasks, routing rules for which model handles what, and workflow docs so the setup survives model swaps. Then I ran a review pass over my own orchestration code and fixed what it found: a routing fallback that could pick the wrong provider, unsafe provider defaults, a cancellation path that could race, and orphaned processes left behind after failures. I also bootstrapped a proper operating layer for the workstation's coding agents.

Main areas covered:

  • benchmark harness for local coding models
  • routing rules based on measured capability
  • workflow documentation for model swaps
  • routing fallback and provider default fixes
  • cancellation safety and orphan process cleanup
  • operating layer bootstrap for coding agents

๐Ÿ”— Key Cybersecurity Connections

Reliability bugs are security bugs waiting for context: an unsafe fallback is an unintended trust decision, a cancellation race is a state-integrity flaw, and orphaned processes are unaccounted-for execution. Reviewing my own automation like hostile code is defensive practice.


๐Ÿ” Investigation Questions

  • Which model does each task type actually route to, and why?
  • What happens when the preferred provider is unavailable?
  • Can a cancelled task leave work half-applied?
  • Are there processes running that no active task owns?
  • Do the benchmark results justify the routing table?

๐Ÿšจ Detection Opportunities

Potential monitoring ideas:

  • fallback routing activating unexpectedly
  • orphaned agent processes accumulating
  • tasks cancelled but still producing output
  • provider configuration drift from defaults
  • benchmark regressions after model updates

Example:

project=local-coding-stack
change_type=orchestration_reliability_fix
risk_area=unintended_execution_paths
triage=audit_fallbacks_cancellation_and_orphans

๐Ÿงญ MITRE ATT&CK Techniques

Possible mappings depending on confirmed behavior:

  • T1059 โ€” Command and Scripting Interpreter
  • T1057 โ€” Process Discovery
  • T1565 โ€” Data Manipulation

๐Ÿ—บ Visual Investigation Diagram

Benchmark harness
    โ†“

Measured capability โ†“ Routing rules โ†“ Review findings โ†“ Fixed, documented stack


โš  Challenges

The challenge was reviewing my own orchestration code honestly. The bugs were all in the unhappy paths โ€” fallback, cancellation, failure cleanup โ€” exactly the paths I had tested least.


๐Ÿ“š What I Learned

I learned that unhappy paths are where automation rots. Everything works in the demo; the fallback, the cancel, and the crash are where wrong trust decisions and stray processes live.


โžก Next Steps

  • Re-run benchmarks after every model change
  • Add checks for orphaned processes to routine maintenance
  • Keep routing decisions written down with their evidence
  • Extend the harness with security-flavored coding tasks

๐Ÿง  Reflection

This was useful because it applied the audit mindset to my own tooling: measure first, route on evidence, and treat every unhappy path as a finding.


๐Ÿงฉ Lessons Learned

What worked

Benchmarking before trusting, and reviewing the orchestrator like third-party code.

What broke

Fallback routing, provider defaults, cancellation, and process cleanup.

Why it broke

I had only exercised the happy path during development.

Fix / takeaway

Test the failure paths deliberately; that is where reliability and security overlap.


๐Ÿ“ˆ Skill Progression Context

This supports my cybersecurity progression because evidence-based trust decisions and unhappy-path analysis are the same skills used in code review and incident investigation.


๐Ÿ˜„ TL;DR

Benchmarked the stack; the bugs were hiding in the failure paths.