Back to all posts
Day 147Saturday, June 27, 20265 min read

'Build Succeeded' Isn't Proof: Writing a Real UI Test for the Auth Flow

cybersecuritytestingauthenticationiosqualityassurancelearningprocess
View original post

๐Ÿ”„ Topic

Following up on the pairing-token auth I added to my agent daemon and iOS app: instead of leaving it at "compiles successfully," I wrote a real end-to-end UI test that drives the actual app against the actual live daemon.


๐ŸŽฏ Goal

Prove the authenticated flow works the way a real user experiences it, not just the way the compiler thinks it should.


๐Ÿ›  What I Did

I built a genuine XCUITest that launches the app, navigates the real UI, enters the pairing token, and asserts the connection state changes โ€” against the live daemon over the real network path, not a mock.

Main areas covered:

  • adding a UI test target to a project that had none before
  • writing a test that launches the app, handles the tab bar's overflow menu, types the token into a secure field, and taps the real "Refresh Diagnostics" button
  • passing the test token in through an environment variable at the scheme level, never hardcoded in source
  • adding accessibility identifiers so the test targets real UI elements instead of guessing at auto-generated labels
  • fixing two real failures the test surfaced, neither of which was a bug in the app
  • keeping the test in the repo permanently as a standing regression check

๐Ÿ”— Key Cybersecurity Connections

A build succeeding tells you the code compiles. It tells you nothing about whether the authentication flow a real user depends on actually completes correctly end to end. This is the same gap that shows up in security testing generally: a control existing in code is not the same claim as a control working when exercised the way an attacker or a legitimate user actually would exercise it. I treated "auth works" as a claim that needed a real test, not an inference from a successful compile.


๐Ÿ” Investigation Questions

  • Does "build succeeded" actually verify the behavior I care about, or just that the syntax is valid?
  • Is the test exercising the real network path, or a mock that could hide a real integration failure?
  • Are test credentials injected safely (environment variable) or accidentally committed to source?
  • When the test fails, is it because the app is broken, or because the test's assumptions about UI structure were wrong?

๐Ÿšจ Detection Opportunities

Potential monitoring ideas โ€” applied to my own test suite instead of a live system:

  • security-relevant flows (auth, permissions, credential handling) with zero test coverage
  • test credentials hardcoded in source instead of injected at runtime
  • a passing test suite that never actually exercises the real backend
  • UI tests that silently rot because they were never converted from a one-off script into a standing test

Example:

project=agent-companion-ios
change_type=security_flow_test_added
risk_area=untested_authentication_path
triage=confirm_test_exercises_real_backend_not_a_mock

๐Ÿงญ MITRE ATT&CK Techniques

Not directly applicable โ€” this is verification engineering, not adversary behavior. No mappings claimed here.


๐Ÿ—บ Visual Investigation Diagram

Build succeeds (compiles)
    โ†“
Assumption: auth flow works
    โ†“
Write real XCUITest against live daemon
    โ†“
Two failures surfaced (tab overflow, label mismatch)
    โ†“
Fix test assumptions, not the app
    โ†“
Third run: real end-to-end success, kept as regression test

โš  Challenges

Both failures I hit initially looked like app bugs and weren't. The tab bar folded a sixth tab under "More," and SwiftUI's LabeledContent merged the label into the accessibility string so a strict equality check failed even though the state was correct. It would have been easy to "fix" the app to match a wrong assumption in the test instead of fixing the test.


๐Ÿ“š What I Learned

I learned to be suspicious of my own first failure diagnosis. Both times, the instinct to change the app was wrong โ€” the test's assumptions about UI structure were the actual bug. Verifying which side is actually broken matters more than fixing the first thing that looks broken.


โžก Next Steps

  • Consider capturing the simulator-driving pattern (tab-bar folding, accessibility identifiers, env-var token injection) as a reusable pattern if more UI tests get written later
  • Keep this test as a standing regression check whenever the auth flow changes
  • Extend similar real-device verification to the other "build-verified only" features from the same session

๐Ÿง  Reflection

This was a good reminder that verification has a hierarchy: compiles, then runs, then behaves correctly under real conditions a user will actually hit. Skipping straight from the first to the third is how confidently wrong software ships.


๐Ÿงฉ Lessons Learned

What worked

Writing a real UI test against the live backend instead of accepting "it builds" as sufficient proof for a security-relevant flow.

What broke

Two of my own test's assumptions about UI structure โ€” not the app itself.

Why it broke

Real UI structure (tab overflow, merged accessibility labels) doesn't always match what you'd guess from reading the view code.

Fix / takeaway

When a test fails, don't assume the app is wrong โ€” verify which side actually is before changing anything.


๐Ÿ“ˆ Skill Progression Context

This supports my cybersecurity progression because distinguishing "compiles" from "verified working" is the same discipline needed when validating that a security control actually functions under real conditions, not just in theory.


๐Ÿ˜„ TL;DR

"Build succeeded" isn't proof anything works โ€” wrote a real end-to-end test for the auth flow and found two wrong assumptions in my own test before it passed for real.