Back to all posts
Day 144Wednesday, June 24, 20265 min read

A 348-Term Glossary and the False Positives That Came With It

cybersecurityknowledgemanagementobsidianwebsecuritylinuxlearningprocess
View original post

๐Ÿ”„ Topic

Building a broad cybersecurity, networking, Linux, and developer-tooling glossary across my study vault, then linking every matching mention โ€” and cleaning up the collateral damage that caused.


๐ŸŽฏ Goal

Turn scattered vocabulary across months of notes into a single, sourced reference layer, without corrupting the notes it touches.


๐Ÿ›  What I Did

I scanned the vault for cybersecurity, Linux, shell, networking, web security, SOC, and adversary-behavior vocabulary, expanded a seed list into 348 terms, and generated one sourced note per term with a plain explanation, a why-it-matters line, and backlinks to where it appeared. Then I linked every matching mention across existing notes back into the glossary.

Main areas covered:

  • generating 349 files: 348 term notes plus one index
  • scanning 485 source files and updating 410 of them with new links
  • inserting 3,820 links across the vault
  • sourcing terms from NIST CSRC, MDN, OWASP, MITRE ATT&CK, RFCs, and official Git/GitHub Pages docs
  • validating every generated and modified file for broken or unbalanced links
  • grouping the 348 terms into five topic maps: Networking, Linux Privilege Escalation, Web Security, SOC Operations, and Developer Tooling

๐Ÿ”— Key Cybersecurity Connections

Mass find-and-link automation across hundreds of files is exactly the kind of bulk change that needs a backup and a validation pass before you trust it โ€” the same reasoning that applies to any bulk remediation script run against production data. I took a pre-link backup archive before running anything, which mattered.


๐Ÿ” Investigation Questions

  • Did the linker touch anything it shouldn't have โ€” frontmatter, code blocks, URLs, existing links?
  • What do the false positives look like, and why did they happen?
  • Is every generated link's target actually present, or does it point at nothing?
  • Can the change be fully reverted from the backup if needed?

๐Ÿšจ Detection Opportunities

Potential monitoring ideas:

  • bulk-edit jobs that modify hundreds of files without a prior backup
  • generated links whose target file doesn't exist (broken reference)
  • automated text-matching that fires inside code blocks or raw syntax it shouldn't touch
  • content restored from backup after an automated pass โ€” a signal the pass needs tighter rules

Example:

project=cybersecurity-glossary
change_type=bulk_autolink_pass
risk_area=unintended_content_modification
triage=validate_then_diff_against_backup

๐Ÿงญ MITRE ATT&CK Techniques

Not directly applicable to this task โ€” this was a knowledge-base build, not adversary behavior. No mappings claimed here.


๐Ÿ—บ Visual Investigation Diagram

Scan vault for vocabulary
    โ†“
Generate sourced term notes
    โ†“
Autolink matching mentions
    โ†“
Validate every link
    โ†“
Restore false positives from backup

โš  Challenges

Four notes came back damaged in a subtle way: raw terminal and Vim bracket syntax in lab session logs and cheat sheets looked enough like wiki-link syntax that the linker mangled it. Validation caught it, but it only caught it because I checked link balance file by file instead of trusting the summary counts.


๐Ÿ“š What I Learned

I learned that "0 missing targets" in a validation report is not the same as "nothing went wrong." The real risk in bulk text automation isn't missing links, it's confidently correct-looking output that quietly broke something the validator wasn't checking for.


โžก Next Steps

  • Add personal lab examples to high-value terms like SSH, SUID, IDOR, CSRF, and lateral movement
  • Consider additional topic maps: Identity and Access, Malware and Adversary Behavior, Detection Engineering, Cloud and Infrastructure
  • Keep the pre-link backup archive until I'm confident no other false positives exist

๐Ÿง  Reflection

This was a good reminder that automation validated against the wrong criteria is still unvalidated. I was checking "did links break," not "did the linker rewrite something it should have left alone" โ€” and both matter.


๐Ÿงฉ Lessons Learned

What worked

Taking a full backup before running a bulk change across 485 files, and skipping frontmatter, code blocks, URLs, and existing links by rule.

What broke

Four notes with raw terminal/Vim bracket syntax got misread as wiki-link syntax and needed manual restoration.

Why it broke

Text that looks like [[...]] syntax isn't always a wiki link โ€” sometimes it's just what a terminal session or a Vim cheat sheet looks like.

Fix / takeaway

Validate for "unexpected changes," not just "expected links present" โ€” a clean link-count report can still hide corrupted content.


๐Ÿ“ˆ Skill Progression Context

This supports my cybersecurity progression because running a large automated change safely โ€” backup first, narrow rules, validate for the failure modes you didn't anticipate โ€” is the same operational discipline that separates a careful defender from one who trusts their own tooling too quickly.


๐Ÿ˜„ TL;DR

348 terms, 3,820 new links, and four notes that had to be rescued from a false-positive linker โ€” the backup is what saved it.