Auditing My Phone-Approval Gate: Wrong Identity Key, Silent Policy Hook, Exposed Webhook
๐ Topic
A real audit of the phone-approval gate from a few weeks back found several concrete weaknesses: an approval queue keyed on the wrong identity, a policy hook that could be knocked over silently, and a webhook sitting on the network where nothing needed it to be.
๐ฏ Goal
Close the specific gaps a hostile actor โ or just bad luck โ could use against the approval mechanism that is supposed to be the last word before anything external happens.
๐ What I Did
I went through the approval queue's actual failure modes one at a time.
Main areas covered:
- keyed the approval queue on the sender, not the message โ the previous keying could let approval of one message get confused with a different message from the same conversation
- minted callback tokens for approvals instead of relying on freeform replies, and auto-denied automated Italian mail that had no business reaching the approval flow at all
- moved approval delivery to Telegram polling instead of a reachable URL โ a publicly reachable webhook for approving privileged actions is attack surface that does not need to exist
- authenticated the sender-policy hook and started backing up the policy file, so a bad actor can't spoof a policy-hook call and a bad edit can't destroy the policy with no way back
- made sure a policy-hook outage never turns into a lost decision โ if the hook is unreachable, the system waits, it doesn't silently drop the approval request
- bounded and expired the approval queue so requests don't linger forever, while making sure expiry reads as "this specific request timed out," not as a blanket block
- added a status surface reporting the sender-approval backlog, and treated a dead approval poller itself as an outage worth alerting on
- wrote a dedicated test covering the sender-policy hook specifically, and logged four distinct findings from this repair as standing lessons
๐ Key Cybersecurity Connections
Keying an approval by message instead of sender is a subtle identity bug with a real consequence: two messages in flight could get their approvals crossed. That is the same category of mistake as session-fixation or token-confusion bugs โ the system asked "was a thing approved" instead of "was this specific thing, from this specific principal, approved."
Removing the reachable webhook in favor of polling is attack-surface reduction in its purest form: nothing can hit an endpoint that does not exist. And treating "the approval poller died" as an outage closes the same gap the Google-auth watchdog closed yesterday โ a control that can silently stop working is not a control you can rely on.
๐ Investigation Questions
- Does the approval queue key on the requester's identity, or something looser?
- Can a policy-hook outage cause a decision to be silently lost or silently allowed?
- Is any approval-relevant endpoint reachable from outside that doesn't need to be?
- Is the policy file backed up, and is the hook that reads it authenticated?
- Does the system alert when its own approval mechanism stops functioning?
๐จ Detection Opportunities
Checks for an approval-gated pipeline:
- an approval resolved for a request other than the one it was issued for
- a policy-hook call with no authentication credential
- an approval-relevant endpoint reachable without an active session or token
- the approval poller silent past its expected heartbeat
- a queued approval expiring silently with no distinguishable "timed out" signal
Example:
project=approval-queue-hardening
signal=approval_resolved_for_wrong_request
risk_area=identity_confusion_in_authorization_flow
triage=verify_queue_key_is_sender_and_request_scoped
๐งญ MITRE ATT&CK Techniques
Possible mappings for the risks being closed:
- T1071 โ Application Layer Protocol (an exposed webhook as a reachable control surface)
- T1078 โ Valid Accounts (identity confusion in the approval keying)
๐บ Visual Investigation Diagram
Approval requested
โ
Old: keyed loosely, reachable webhook, unauthenticated hook
โ
New: keyed on sender, delivered via Telegram polling
โ
Policy hook authenticated + backed up
โ
Hook unreachable โ wait, don't drop
โ
Poller death โ alerted as an outage
โ Challenges
The sender-vs-message keying bug was the hardest to see, because it only shows up under concurrency โ two approvals in flight at once. It took deliberately imagining an adversarial timing scenario, not just reading the code in isolation, to notice the gap.
๐ What I Learned
I learned that an approval mechanism has its own attack surface, separate from whatever it's gating. A perfect deny-first policy behind a sloppy approval queue is still a sloppy approval queue.
โก Next Steps
- Re-audit the queue under deliberately concurrent test scenarios
- Rotate and monitor the callback tokens
- Extend the "control death is an outage" pattern to every gate in the stack
- Review the four logged findings for whether they generalize to other queues
๐ง Reflection
Auditing the mechanism that is supposed to be the last line of defense felt like the right kind of paranoia โ the gate itself needs the same scrutiny as whatever it's protecting.
๐งฉ Lessons Learned
What worked
Removing a reachable webhook, authenticating the policy hook, and keying approvals on sender identity.
What broke
Message-keyed approvals under concurrency, an unauthenticated policy hook, and a poller whose death went unnoticed.
Why it broke
The queue was designed for the common case, not for concurrent or adversarial timing.
Fix / takeaway
Key authorization state on the principal, not the artifact; authenticate every hook; and alert on the death of the control itself.
๐ Skill Progression Context
This supports my cybersecurity progression because identity-scoped authorization state, attack-surface reduction, and monitoring the health of the control itself are core access-control engineering skills.
๐ TL;DR
The approval gate got its own audit โ sender-keyed, webhook removed, authenticated, and alerted if it ever goes quiet.