Back to all posts
Day 204Sunday, August 23, 20267 min read

Shadow-Testing a New Briefing Feature Before It Reaches Anyone

cybersecurityshadowdeploymentevidencegatingchangemanagementdeliveryverificationlearningprocess
View original post

๐Ÿ”„ Topic

Rolling out a new "Phase 4" briefing feature meant reusing a pattern from earlier this month โ€” shadow trials before anything reaches a real user โ€” but this time as deliberate, staged, standard practice: request-and-verify explicit use, verify delivery receipts, gate on readable drafts, and reject internal content that isn't actually meant for a reader.


๐ŸŽฏ Goal

Get a new user-facing feature through a chain of independent evidence gates before it's trusted to reach me directly, rather than shipping it after "it looked right in a demo."


๐Ÿ›  What I Did

I built the feature and the verification chain around it together, treating each stage as something that had to prove itself before the next stage could begin.

Main areas covered:

  • built the daily attention budget enforcement for the briefing feature first, so the feature itself had a bounded scope before anything else was layered on top
  • made the shadow run deterministic โ€” the same inputs produce the same shadow output, so a shadow-day result can actually be compared and trusted rather than treated as one more roll of a non-deterministic dice
  • read from the canonical roadmap projection rather than a stale or duplicated copy, fixing this twice after finding it drifting
  • verified unique-day shadow progress explicitly โ€” confirming the feature was actually producing a new day's content each day, not silently repeating or stalling
  • recorded "attention brief day one," then continued the chain: deployed cron trace recorded, current roadmap action selected correctly, and a bug fixed where an internal roadmap day was being surfaced as if it were meant for a real reader
  • built the delivery step itself: prepared drafts delivered to "reader" mode, verified readable, gated so nothing ships to shadow unless the draft is actually readable โ€” not just structurally present
  • added a fix that specifically rejects internal roadmap shadow days from being treated as user-facing, closing a gap where planning content meant for me-as-operator could have leaked into what looks like a real briefing
  • built a delivery-receipt verifier: requests and verifies explicit use before proceeding, then separately verifies cron-triggered Telegram receipts โ€” confirming a scheduled send actually reached and was recorded as received, not just that the send call didn't error
  • surfaced validated user actions from Phase 4 as its own step, and recorded a "Phase 4 shadow Day 3" checkpoint, continuing to track the feature's shadow-day count as a first-class piece of evidence rather than an informal sense of "it's been running a while"

๐Ÿ”— Key Cybersecurity Connections

This is the shadow-mode evaluation pattern from earlier this month, now applied by default to a brand-new feature rather than retrofitted after a problem โ€” that's the real story: a control built once, in response to a specific incident, became standing practice for anything new. Delivery-receipt verification specifically closes the "send is not delivery" gap that recurred elsewhere this same month: a cron job reporting that it sent a Telegram message is not evidence the message was received, and this feature doesn't get to claim success without that distinct, separate check.

Rejecting internal roadmap content from reaching a "reader" surface is a data-classification boundary: planning and operational content has a different audience than a finished, user-facing briefing, and the two need to stay structurally separated rather than trusted to never cross paths just because they flow through the same pipeline.


๐Ÿ” Investigation Questions

  • Does a new user-facing feature go through staged, evidence-gated shadow trials before reaching a real user, or does "it demoed fine" count as done?
  • Is delivery verified as actually received, separately from confirming the send call didn't error?
  • Can internal, operator-facing content reach a reader-facing surface, or is that boundary structurally enforced?
  • Is a shadow run deterministic enough that its results can be meaningfully compared day over day?
  • Is shadow-day count tracked as real evidence, or just an informal sense that "it's been running a while so it's probably fine"?

๐Ÿšจ Detection Opportunities

Checks for a staged feature rollout:

  • a new user-facing feature reaching real users without a documented shadow-trial history
  • a "message sent" log treated as equivalent to "message received" with no receipt verification
  • internal/operational content reaching a reader-facing delivery path
  • a shadow run that isn't deterministic, making day-over-day comparison meaningless
  • a rollout decision made on vibes ("it's been running a while") rather than a tracked shadow-day count and pass criteria

Example:

project=hermes-briefing-phase4
signal=send_confirmed_without_delivery_receipt_verification
risk_area=false_assurance_of_feature_readiness
triage=require_receipt_verification_before_counting_a_shadow_day_as_passed

๐Ÿงญ MITRE ATT&CK Techniques

No direct mapping claimed. This is staged deployment, evidence-gating, and delivery verification methodology for a new feature, not an adversary technique.


๐Ÿ—บ Visual Investigation Diagram

New Phase 4 briefing feature planned
    โ†“
Attention budget scoped first
    โ†“
Shadow run made deterministic
    โ†“
Canonical roadmap source fixed, unique-day progress verified
    โ†“
Internal roadmap days rejected from reader-facing output
    โ†“
Delivery-receipt verifier: explicit use + cron Telegram receipts confirmed
    โ†“
Readable-draft gate before shadow counts as passed
    โ†“
Phase 4 shadow Day 3 recorded โ€” evidence accumulates, not assumed

โš  Challenges

The temptation with any new feature is to let momentum from a working demo substitute for the slower, staged verification chain โ€” especially a feature this granular, where each gate (deterministic shadow, canonical source, receipt verification, readable-draft check) feels like it could reasonably be skipped "just this once." Building all of them in before the feature reached a real user was the more expensive but correct choice.


๐Ÿ“š What I Learned

I learned that a security or reliability pattern only really counts as learned once it's applied by default to something new, without being prompted by a fresh incident. The shadow-trial discipline from earlier this month stopped being "the fix for that one bug" and became "how new features get shipped here" โ€” which is the actual measure of whether a lesson stuck.


โžก Next Steps

  • Continue tracking shadow-day count and pass criteria until Phase 4 has enough evidence to graduate to fully live
  • Apply the same delivery-receipt-verification pattern to any other cron-triggered user-facing send
  • Formalize the internal-vs-reader content boundary as a reusable check for future features, not a one-off fix
  • Review whether the daily attention-budget enforcement needs adjustment once real shadow data accumulates

๐Ÿง  Reflection

Watching myself reach for the shadow-trial pattern automatically, for a feature that had nothing to do with the incident that originally produced it, was a better signal of progress than any individual bug fix this month โ€” the discipline generalized.


๐Ÿงฉ Lessons Learned

What worked

Treating shadow deployment as default practice for a new feature, verified delivery receipts instead of trusting send confirmations, and structurally rejecting internal content from reader-facing output.

What broke

Nothing shipped broken to a real user โ€” which is the point of catching a stale roadmap source, a non-deterministic shadow run, and an internal-content leak risk during the shadow phase instead of after launch.

Why it mattered

A feature that reaches a real user without staged, evidence-gated verification inherits every unverified assumption baked into its pipeline, silently.

Fix / takeaway

Make shadow-trial evidence gating the default for new features, not a response reserved for after something already went wrong once.


๐Ÿ“ˆ Skill Progression Context

This supports my cybersecurity progression because staged rollout discipline, delivery verification distinct from send confirmation, and structural content-boundary enforcement are the same evidence-based change-management practices used in real production security and reliability engineering.


๐Ÿ˜„ TL;DR

A new briefing feature had to earn its way to me through a chain of shadow trials, delivery-receipt checks, and a reader/internal content boundary โ€” the lesson from an earlier incident, now just how things ship.