🛰️

WiFi Odds

extension strategy and specs

The plan

WiFi Odds: the extension strategy

The v2.3+ roadmap, chosen by a blind Claude-vs-Codex strategy competition. Codex won, 22 to 25 against 21 to 25.

North star

Trust times depth at the decision point. Maximize the share of supported booking sessions that produce a confident, actionable choice between two or more real flight options.

v2.3 Make the current 4 surfaces decisive

No new sites

  • The "Best WiFi choice" decision strip.
  • Confidence on every badge: sample size, freshness, confirmed-tail check.
  • Price times WiFi, and date-flex reframing.
  • Bring United, Google Flights, Navan, and Alaska to one quality bar.

v2.4 Extend across time and itineraries

Reach beyond a single result page

  • Guard-and-rescue: recheck near departure, offer the best still-available alternative.
  • Local candidate tray plus whole-itinerary ranking.
  • A shareable trip-odds card.
  • Pilot Concur.

Continuous The moat and the maintenance

Always on

  • The @martinamps calibration and delta data contract (the moat).
  • Daily canaries and resilient selectors.
  • A monthly Starlink fleet report.

Surface policy

  • A surface may show only carriers with decision-grade evidence.
  • Breadth follows data, it never leads.
  • Consumer OTAs (Kayak, Expedia) are out, because Google Flights already gives metasearch reach.

The through-line

The two deepest moat items, prediction calibration and full Guard-and-rescue, are both unlocked by the same single ask to @martinamps: post-departure outcome backfill.

The product

Live UI mockups

Faithful renders of the three surfaces at the heart of v2.3 and v2.4: the injected odds panel, the honest decision strip, and the Guard-and-rescue notification. Odds run green (higher) to amber (lower). The confirmed-tail mark is a green check.

1 · Injected results panel

Rendered on the results page next to each fare. A star marks the top pick, the odds pill is color-ramped, sample size rides alongside, and a green check flags a confirmed tail.

🛰️ DEN → SFO · WiFi odds ESTIMATES
UA1596 8:30am departure
🛰️ 68% 51 flights
UA2402 2:15pm departure
🛰️ 16% 32 flights
UA2265 6:00am departure
🛰️ 52% 44 flights

Demonstrates: per-flight Starlink odds injected inline, ramped by likelihood, with sample size and a confirmed-tail check for trust.

2 · "Best WiFi choice" decision strip, all three honest states

The strip only claims a winner when the evidence supports one. When the top options are close, or when only one flight is scored, it says so and shows no action button.

Best WiFi choice · winner
🏆 Best WiFi: UA1596
Leads by 38 pts, 50 departures observed, high confidence.

A clear winner: one flight leads by a wide margin on a solid sample, so the strip offers to reorder the results.

Best WiFi choice · too close
🤝 Top options are close, no clear WiFi winner
The top two are within 5 pts. Both listed below.

Too close to call: the lead is inside the margin, so no button and no false certainty.

Best WiFi choice · one scored
🔎 Only one scored flight
UA1596 at 68%, nothing else here to compare it against yet.

Nothing to compare: a single scored flight cannot be a "best choice," so the strip stays honest and offers no action.

3 · Guard-and-rescue notification, all three honest states

Styled as a browser notification. Every card names the flight and date. The confirmed-non-Starlink and unknown cards add the best still-available alternative and a way back to the booking. The extension only informs, it never changes a booking.

🛰️
WiFi Odds · Guard & rescue
UA1812 · Jul 25 STARLINK CONFIRMED
Assigned tail N127UA has Starlink. You are set.
↩ Back to your booking

State A, confirmed Starlink: the assigned aircraft has it, so no alternative is offered.

WiFi Odds · Guard & rescue
UA1812 · Jul 25 NO STARLINK
Assigned tail N840UA (Viasat). We cannot count on Starlink for this aircraft.
Best still-available alternative: UA1234, now about 72%.
↩ Back to your booking

State B, confirmed non-Starlink or unknown aircraft: names the tail, then surfaces the best still-available option the traveller already saw.

WiFi Odds · Guard & rescue
UA1812 · Jul 25 NO ASSIGNMENT YET
Aircraft not assigned yet (tails publish about 48h out). We will keep watching.
Best still-available alternative: UA1234, about 72% while you wait.
↩ Back to your booking

State C, assignment unavailable: explicit that it does not yet know, rather than implying a negative.

Launch data story

How often does United actually fly Starlink? We measured.

Data as of 2026-07-30. All figures are historical estimates from a community tracker and United's own filings; equipment can change until departure. Sources are listed at the bottom so anyone can check the math.


Reddit version (r/flying, r/awardtravel)

Short answer: less often than the headlines suggest, and it depends heavily on what you're flying.

Per United's own Q2 2026 report (as of June 30, 2026), Starlink was installed on 450 of United's 1,552 mainline plus United Express aircraft, about 29%. A community tracker that pulls the fleet daily counts a slightly broader aircraft population and shows 485 equipped as of July 30, 2026. So call it "roughly a quarter to a third, and climbing."

But the fleet-wide average hides the real story. Three honest data points:

  1. Regional beats mainline right now. On United Express regional jets, about 51% already have Starlink (341 of 669 tracked aircraft). On mainline it is about 13% (144 of 1,141). If you are on a small two-cabin regional jet out of Chicago, your odds are genuinely better than on a mainline 737.
  2. It is really an aircraft-type story. The CRJ-550 is basically done (93 of 94, about 99%) and the E175 is close behind (248 of 256, about 97%). Meanwhile the 737 MAX 8/9, A319/A320, and 757-300 have no equipped aircraft recorded in the tracker yet, and the widebody 777 is at 1 of 96. Two planes flying the same route can give you completely different odds.
  3. It moves week to week. The tracker logged 53 installs in the last 30 days, roughly 9.7 mainline aircraft a week. United's stated target is 1,000 aircraft by end of 2026, so this number is a moving target, not a fixed stat.

Why your specific flight is a coin toss: the tail number (the actual airframe) usually is not assigned to a flight until roughly 24 to 48 hours out, and equipment swaps happen right up to departure. Route matters too: equipped departures cluster on Chicago and hub flying (today's most-Starlink routes are LGA-ORD and ORD-LGA at 15 departures each). So "does flight UAxxxx have Starlink" is better asked the day before than a month before.

Disclosure, kept low-key: I help with a free, open-source WiFi-odds tracker and extension that turns this into a per-flight estimate. The numbers above come from it and from United's filings; method and caveats are at wifiodds.com/methodology, and it is built on the community trackers (unitedstarlinktracker.com and the tracker work by @martinamps). Not selling anything, only thought the data was worth sharing.


Blog version

How often does United actually fly Starlink? We measured.

United has spent the past year telling a fast-moving Starlink story, and the headline numbers are real. But "how often will my flight have it" is a different question than "how many planes are equipped," and the honest answer is: it depends on your aircraft, your route, and the calendar. Here is what the data actually says as of July 30, 2026.

The fleet-wide number

Start with the publisher of record. In its Second-Quarter 2026 results (released July 15, 2026, figures as of June 30, 2026), United reported Starlink installed on 450 of its 1,552 mainline and United Express aircraft, about 29%.

A community tracker that pulls the fleet every day counts a slightly broader aircraft population and shows 485 equipped as of July 30, 2026. Those two numbers use different denominators (United's 1,552 is United's own end-of-period fleet; the tracker's population is larger), so we do not divide one by the other. The takeaway is the same either way: roughly a quarter to a third of the aircraft you might be assigned already have it, and the count rises most days.

Where the average lies to you

The fleet average blurs two very different fleets.

  • United Express (regional): about 51% equipped, 341 of 669 tracked aircraft.
  • United mainline: about 13% equipped, 144 of 1,141.

That is the counterintuitive part: right now the small regional jets are ahead of the big mainline jets. It comes straight from which aircraft types got done first.

It is an aircraft-type story

Tracked United aircraft with Starlink, by type (July 30, 2026):

AircraftEquipped / totalShare
CRJ-550 (regional)93 / 94about 99%
E175 (regional)248 / 256about 97%
737-80064 / 141about 45%
A321neo33 / 72about 46%
737-90046 / 148about 31%
777 (widebody)1 / 96about 1%
737 MAX 8/90 / 285none recorded yet
A319 / A3200 / 143none recorded yet
757-3000 / 60none recorded yet

Two flights on the same route can hand you very different odds, because they can be different airframes. A regional E175 is nearly a sure thing; a 777 is nearly a no; a mainline 737 is a genuine coin flip.

Why your specific flight is hard to promise

  • Tail assignment is late. The specific airframe on a flight typically is not locked until roughly 24 to 48 hours before departure, and equipment swaps happen up to boarding. A month-out guess is weaker than a day-before one.
  • Routes cluster. Equipped departures concentrate on hub and Chicago flying: today the busiest Starlink routes are LGA-ORD and ORD-LGA (15 equipped departures each) and ORD-CVG (13).
  • The number moves. 53 installs in the last 30 days (about 9.7 mainline aircraft a week), against United's stated targets of 1,000 aircraft by end of 2026 and the full fleet by end of 2027.

How United compares

United is not the leader on percentage, and it is not close to last. A few honest reference points from the same tracker (Starlink installs, latest pull):

  • JSX: 56 of 56 (100%), an all-Starlink boutique carrier.
  • Hawaiian: 42 of 66 (about 64%).
  • WestJet: 100 of 193 (about 52%).
  • Qatar Airways: 120 of 241 (about 50%).
  • Alaska: 99 of 350 (about 28%), the closest peer to United's share.
  • Emirates: 36 of 232 (about 16%).
  • Air Canada: 12 of 216 (about 6%).
  • British Airways: 5 of 261 (about 2%).
  • Southwest: 1 of 803 (about 0.1%), very early.

Small all-Starlink operators sit at 100% because their fleets are tiny. Among large legacy fleets, United's roughly quarter-to-third share is middle of the pack and moving up fast.

Methodology and caveats

  • These are historical estimates, not a live seatmap. They come from a community tracker (unitedstarlinktracker.com) pulled daily, anchored to United's own filings where United publishes a number.
  • Two denominators, kept separate. United's 450/1,552 is its own reported fleet as of June 30, 2026. The tracker's live 485 counts a broader aircraft population, so we never divide the tracker's equipped count by United's fleet total.
  • Tail assignments publish about 24 to 48h out and equipment can change until departure, so a per-flight estimate is a probability, not a guarantee.
  • Unknown is not zero. Some carriers (for example Air France and SAS) have announced Starlink but no published installed count; the tracker reports those as unpublished, not as 0%. We do not turn "we do not know" into "none."

The tool that produces the per-flight version of this is free and open source. Method, sources, and limitations live at wifiodds.com/methodology, and the whole thing stands on community tracker work (unitedstarlinktracker.com and the trackers by @martinamps).


Appendix: numbers used (for fact-checking)

Every figure in this piece and its exact source. Two files were treated as authoritative: united/data.json (the daily United tracker pull) and assets/airlines.js (the 18-airline ledger), both in the wifiodds repo.

  • 450 of 1,552 (about 29%), as of June 30, 2026: United's Second-Quarter 2026 results (Exhibit 99.1, published July 15, 2026), recorded in united/data.json under fleet.published (equipped: 450, total: 1552).
  • 485 equipped, tracker population 1,810, as of July 30, 2026: united/data.json fleet.equipped / fleet.total and updated.
  • Mainline 144 of 1,141 (about 13%); regional/Express 341 of 669 (about 51%): united/data.json fleet.mainline and fleet.express.
  • By type (CRJ-550 93/94, E175 248/256, 737-800 64/141, A321neo 33/72, 737-900 46/148, 777 1/96, 737 MAX 8/9 0/285, A319/A320 0/143, 757-300 0/60): united/data.json fleet.types.
  • 53 installs in last 30 days; about 9.7 mainline/week: united/data.json fleet.last30 and fleet.mainlinePacePerWeek.
  • Targets: 1,000 by end of 2026, full fleet by end of 2027: united/data.json fleet.targets.
  • Busiest Starlink routes today: LGA-ORD 15, ORD-LGA 15, ORD-CVG 13: united/data.json leaderboard (top entries, 2026-07-30).
  • Comparison airlines (JSX 56/56, Hawaiian 42/66, WestJet 100/193, Qatar 120/241, Alaska 99/350, Emirates 36/232, Air Canada 12/216, British Airways 5/261, Southwest 1/803): assets/airlines.js WIFI_AIRLINES, each airline's equipped / fleet.
  • Air France and SAS shown as unpublished, not 0%: assets/airlines.js; pctEquipped() returns null (not 0) for both.

All percentages are computed from the counts above and rounded; the underlying counts, not the rounded percentages, are the source of record.

The moat

Data partnership with @martinamps

The Starlink trackers own the ground truth. This is the proposed data contract that would make that data a source of record downstream tools can build on, plus the outreach email that opens the conversation.

Data-contract proposal for the Starlink trackers

From: Jeremy (wifiodds.com, the "WiFi Odds" Chrome extension)
To: @martinamps, unitedstarlinktracker.com, alaskastarlinktracker.com
Status: Draft for discussion. Nothing here is a demand; it is a menu, ordered by leverage.

Relationship to wifiodds-martinamps-proxy-note.md: That note asks one operational question, may I put an edge cache in front of your API so you see one origin instead of many browsers. This document is the layer under it: if we ever formalize how the extension consumes your data, here is the shape that would make both sides' lives easier. Read the proxy note first; this is the sequel, not a replacement.

The one-paragraph version

The highest-leverage thing you could give a downstream consumer is not more airlines. It is calibration and deltas: (1) explicit machine states so a consumer can never mistake "no data" for "your server erred" for "not published yet," (2) a changed-since feed so we poll you a fraction as often, and (3) post-departure outcome backfill so the probabilities you publish can be scored against what actually flew. Everything else is a nice-to-have. Those three turn a good dashboard into a source of record that other tools can build on without corrupting it.

Why this comes from a place of respect

The extension already consumes four of your surfaces per route, per direction, on page load:

  • GET /api/plan-route?origin=&destination=, itineraries with joint_probability, at_least_one_probability, coverage, total_flight_hours, and per-leg probability / n_observations.
  • POST /mcppredict_route_starlink, the per-flight odds table.
  • POST /mcpsearch_starlink_flights, confirmed departures with tail.
  • POST /mcpcheck_flight, per-flight, per-date assignment status.
  • GET /api/predict-flight?flight_number=, single-flight probability (0 to 1), n_observations, confidence.

We consume exactly four fields off the per-flight path (prob, obs, conf, and a derived confirmed-tail marker) and we treat everything you emit as facts to display, never to re-derive. The extension has already had to invent client-side workarounds for the gaps below. This proposal is really those workarounds, moved upstream where they belong and where you would control them.

Ask 1: a versioned, decision-grade per-flight record with explicit states

Today a single predict-flight response has to be interpreted three different ways, and the extension guesses at the difference in code:

  • A real number → show the odds.
  • A recognized "no history for this flight number" shape → show n/a and it is safe to cache.
  • A 429/500/garbled 200 → must not be cached, or we would show a false n/a for hours.

We currently distinguish these with a client-side sentinel (a magic "error" string, because Chrome's message channel silently drops undefined) and a directOk flag we synthesize ourselves. That logic is fragile precisely because the server never told us which case it was, we inferred it from response shape. One schema tweak on your side and our inference breaks.

The fix is to make the state explicit and versioned. A record shaped roughly like:

{
  "schema": "starlink-flight/1",
  "flight": { "number": "UA1812", "carrier": "UA", "origin": "SFO", "destination": "SEA" },
  "state": "PUBLISHED",              // PUBLISHED | NO_HISTORY | PENDING | UPSTREAM_ERROR
  "probability": 0.62,              // present iff state == PUBLISHED, else null
  "confidence": { "basis": "history", "n_observations": 51 },  // basis: history | aircraft_type
  "sample_window": { "start": "2026-06-01", "end": "2026-07-29" },
  "last_refreshed": "2026-07-30T04:10:00Z",
  "as_of": "2026-07-30T04:10:00Z"
}

The four state values are the whole point, and they map exactly to bugs we have already hit:

stateMeaningWhat a false version of this caused
PUBLISHEDReal odds, trust them(none)
NO_HISTORYGenuinely no data for this flight number; safe to cache as n/a(none)
PENDINGTail assignment has not published yet (about 48h out)Showing a stale-as-fresh n/a
UPSTREAM_ERRORYour server could not answer; do not cache, retry laterA false n/a pinned for 6h

Two subtler fields that already matter to us:

  • confidence.basis: you already signal confidence: "type" when odds are derived from the aircraft subfleet rather than that flight number's own history. Making that a structured field (history vs aircraft_type) instead of an overloaded enum lets us label it honestly to users without string-matching.
  • last_refreshed vs as_of: one says when your pipeline last touched this record, the other says how fresh the underlying observation is. Right now we cannot tell a value that is stale because nothing changed from one that is stale because your refresh stalled. Two timestamps settles it.

schema up top means you can evolve the record and downstream consumers fail loud on a version bump instead of silently misreading a renamed field.

Ask 2: a "changed-since" delta feed

The extension caches route data for 6 hours and caps itself at 100 MCP calls per local day, per user, with a 250 to 400ms spacing between calls, all self-imposed politeness because we have no way to know when your data actually moved. Most of those calls re-fetch data that did not change.

A single delta endpoint fixes the waste on both sides:

GET /api/changes?since=2026-07-30T04:00:00Z[&carrier=UA]
→ {
    "as_of": "2026-07-30T04:32:00Z",
    "cursor": "2026-07-30T04:32:00Z",
    "changed": [
      { "flight": "UA1812", "date": "2026-08-02", "kind": "tail_published", "tail": "N127UA" },
      { "flight": "UA544",  "date": "2026-08-02", "kind": "tail_swap", "from": "N127UA", "to": "N38473" }
    ]
  }

What this buys you: instead of N browsers each polling four surfaces every few hours, one consumer polls one cheap endpoint and only fetches full records for flights that actually changed. Your request volume from us drops by roughly the fraction of flights that are static between refreshes, which is most of them. kind matters because we already track tail swaps as first-class events (a booking that had a Starlink tail and lost it is our single most valuable alert), and a delta feed is exactly the shape that event stream should arrive in.

This pairs directly with the edge-cache question in the proxy note: a delta feed is what makes the cache cheap to keep warm without hammering you to revalidate.

Ask 3: post-departure outcome backfill (the calibration payload)

This is the one that turns published probabilities into scored probabilities. Right now you publish, for example, "UA1812 gets a Starlink tail about 62% of the time." Nobody, including you, can readily ask "and over the last 200 departures of that flight, how often did it actually?" because once a flight departs, the assignment that flew is not retained in a queryable, per-flight-per-date form.

The ask: after departure, publish the actual tail that operated (which you already know at check-flight time; it is not retained addressably):

GET /api/outcomes?carrier=UA&since=2026-07-01
→ [
    { "flight": "UA1812", "date": "2026-07-28", "actual_tail": "N127UA",
      "had_starlink": true, "predicted_probability": 0.62,
      "prediction_as_of": "2026-07-26T04:00:00Z" }
  ]

With that, published probabilities become backtestable by route, by booking horizon, and by fleet-rollout era, which matters a lot right now specifically because the fleet is moving fast (United is at 485 equipped and adding about 9.7 mainline aircraft/week toward a 1,000-by-end-2026 target). A 62% that was well-calibrated in May is stale by August as the equipped share climbs, and only outcome data can catch that drift. This is the difference between a forecast and a scoreboard, and it is the single most defensible thing a data source can offer: "here is my prediction, and here is my track record against it."

We would do the calibration math and hand it back to you, publicly, as a reliability curve for your own numbers. That is a feature for your site, not only ours.

What WiFi Odds offers in return

This has to be additive or it is not worth your time. Concretely:

  • Code and PRs. The ConnectScore model, the fleet parsing, and the whole extension are already open on GitHub. If a delta feed or an outcome table is useful to you, I will build the endpoints against your stack as PRs, not only ask for them. The calibration tooling I would write anyway; you would get it too.
  • QA fixtures. I hit your surfaces from a lot of real booking sessions and have already catalogued your response shapes (including the aircraft-type-derived confidence case and the per-carrier wording differences between the United and Alaska trackers). I can hand you a fixture suite that pins those shapes, so a future change on your side that would break a consumer fails your tests first.
  • Visible attribution and methodology links. Every surface that uses your data already credits it by name ("data: unitedstarlinktracker.com") and links back, with a fuller acknowledgement on the methodology page. That stays and gets more prominent, not less.
  • Reciprocal install and links. Happy to link your trackers from the site and the extension listing, and to carry whatever canonical URL/credit line you prefer.
  • A bright line: the trackers establish facts; the extension turns facts into a workflow. You own the ground truth, which tail has Starlink, mainline vs regional. The extension's job is narrow and downstream: put your fact next to a fare while someone is booking, and alert them if the tail swaps between booking and boarding. It is a booking/rebooking layer on top of your data, not a mirror of it and not a competitor to it. I do not want to stand up a copy of your dataset; I want your dataset to be the thing people cite.

Minimum viable version (the two asks worth starting with)

If none of the above is worth a big lift, here are the two that carry most of the value and are the cheapest to ship:

  1. Explicit state on the per-flight record (Ask 1's four-value enum). This is small, it is mostly surfacing a distinction your server already makes internally, and it immediately kills an entire class of false-n/a and stale-as-fresh bugs for every consumer, not only us.
  2. Outcome backfill (Ask 3). One append-only endpoint of "what actually flew," published after departure. It is the calibration foundation, it is the most defensible thing you could offer, and it costs you nothing you do not already compute.

The delta feed (Ask 2) is the natural third step once those two exist, because a delta feed of state changes and settled outcomes is trivial to derive from them.

Everything here is a proposal, not a request for a commitment. If the honest answer is "the current API is fine, cache it and credit me," that is a completely fine answer and the proxy note covers it. This document exists so that if you do want to make the data more useful downstream, the highest-leverage moves are already written down.

Jeremy

Outreach email to @martinamps

How this relates to wifiodds-martinamps-proxy-note.md: This complements that draft, it does not replace it. The proxy note leads with one operational question, may I put an edge cache in front of your API. This email leads with the warm introduction plus the two highest-value collaboration asks (a changed-since delta feed and post-departure outcome backfill for calibration). The technical detail behind those asks lives in the companion martinamps-data-contract-proposal.md.

Which to send: If Jeremy wants to open with the lightest-touch operational ask, send the proxy note as-is. If he wants to open with the collaboration framing, send this. They can also be merged: this email's intro plus the proxy note's caching paragraph plus a one-line pointer to the data-contract doc, but do not send both separately or the first ask will look decorative.

Send trigger (from the proxy note, now satisfiable): the note said "send after v2.0 is live in the Chrome Web Store and The Forecast is merged." v2.2.0 is now live in the store, so the store condition is met; confirm The Forecast is merged before sending.

Subject line options

  • Fan of the tracker, and a small collaboration idea
  • unitedstarlinktracker.com built my extension's best feature (a thank-you and an ask)
  • Introducing WiFi Odds, and asking before I lean on your data harder

Body

Hi Martin,

I fly around 80 segments a year and I have been using unitedstarlinktracker.com to pick flights for a while now. Knowing whether a departure is likely to draw an equipped tail genuinely changed how I book, and your rollout counts are one of the few dashboards I actually check. Thank you for building it and for keeping it current.

I have built something adjacent, and I would rather introduce myself than have you find it in your logs.

WiFi Odds is a free, open-source, unofficial Chrome extension that shows next-gen WiFi odds right on the page while you are booking: on united.com and Google Flights it puts the Starlink odds next to each fare. Your trackers are the per-tail source underneath it for United and Alaska, so those airlines rest entirely on your work. Every badge credits you by name and links back, with a fuller acknowledgement on the methodology page. It launched as v2.2 in the Chrome Web Store. It is not monetized (no ads, no accounts, no analytics, no affiliate links) and it is all on GitHub. If any of it is useful to you, the ConnectScore model, the fleet parsing, any of it, take it.

Happy to link your trackers prominently from both the site and the extension listing, and to carry whatever credit line you would prefer, reciprocal links, your call.

The reason I am actually writing: there are two things you could someday expose that would make the data far more useful downstream, and I would do the work for them rather than only ask.

  1. A "changed-since" feed: a cheap endpoint that returns only the flights whose tail assignment changed since a timestamp. Today the extension politely re-polls on a timer because it cannot tell when your data moved; a delta feed would cut the requests reaching you to a small fraction, and it is the natural shape for the tail-swap events I already track.
  2. Post-departure outcomes: the actual tail that operated a flight, published after it departs. That is the missing piece for calibration: it lets your published probabilities be scored against what really flew, by route and booking horizon, which matters a lot while the fleet is rolling out this fast. I would run that math and hand it back to you publicly as a reliability curve for your own numbers.

No pressure on any of this: if the honest answer is "the API's fine as-is, cache it and credit me," that is a completely fine answer and I will do exactly that. I mostly wanted you to know the extension exists, that it credits you, and that I would rather build with you than around you. I wrote the technical detail up separately so this email could stay short; happy to send it if you are curious, or to open a GitHub issue so there is a public record.

Either way, thank you for the tracker. It is a genuinely useful thing to have made.

Jeremy
[email protected] · github.com/jeremyinthebay

Notes for Jeremy (not for Martin)

  • Keep three claims true against the code before sending: not monetized, nothing mirrored/stored, and the credit-by-name is actually rendered on every badge (it is, as of v2.2).
  • Reaching him: his tracker sites have no contact route. His personal site ma.rtin.so (linked from his GitHub) points to his resume PDF for an email; a short GitHub issue on github.com/martinamps/ua-starlink-tracker is the good public-record channel. Email for the considered reply, issue for the paper trail.
  • He is a working engineer (@anthropics on GitHub); the register is pitched for that. Do not oversell.

Trust backstop · design

Accuracy Harness: design (routes times dates)

What this is. A build-ready spec for the trust backstop of the WiFi Odds Chrome extension: a harness that verifies the odds the extension shows are actually accurate. Spec only, nothing here is built yet. It extends the existing test/ harnesses; it does not replace them.

Status: design. Read-only pass over extension/ and test/ on 2026-07-30. No repo was modified.

One sentence. Two layers: (1) does the extension display what the tracker returned, correctly, across a route times date matrix, and (2) do the tracker's published probabilities match what actually flew, with the second layer honestly gated on data the tracker does not yet expose.


0. Where this fits (and what already exists)

The extension's predictions do not come from wifiodds.com. They come from @martinamps' Starlink trackers (unitedstarlinktracker.com, alaskastarlinktracker.com), keyed on flight number and route. wifiodds united/data.json is only the equipped-tail roster. So the system under test is the tracker, cross-checked against our roster where the two can be joined.

Three harnesses, one of which is this proposal:

HarnessFileAnswersDeterminismGate?
Phase 1: API matrixtest/phase1-api-matrix.mjsDo the tracker's endpoints agree with each other and with our roster? What would the panel render?Live trackerReport only
Phase 2: browser E2Etest/phase2-e2e.mjsDoes the content script render what a fixed tracker response says?Fully mockedRelease gate (exit 1)
Accuracy harness (this doc)test/phase3-accuracy.mjs (new)Layer A: display truth across routes times dates. Layer B: are the published probabilities calibrated against realized outcomes?Layer A deterministic; Layer B replay of recorded snapshotsLayer A gates; Layer B reports

The gap this fills: Phase 2 proves the display is correct for a response we invented. Phase 1 proves the tracker's surfaces are internally consistent right now. Neither one ever asks the only question a user actually cares about: when the extension said "68%," did a Starlink plane show up 68% of the time? That is a calibration/backtest question, and it needs realized outcomes over time, which is exactly what the data contract's Ask 3 (post-departure outcome backfill) would provide. Until then, the harness records predictions now so the backtest becomes computable the day outcomes exist, and measures everything that is checkable today.

1. Goal and the two layers

Layer A: DISPLAY accuracy (buildable today)

The extension shows what the tracker returned, correctly, across routes times dates.

This is a superset of Phase 2, widened from a handful of hand-authored cases to a route times date matrix driven by real (recorded, then replayed) tracker responses. It asserts, per cell:

  • the badge's %, obs, and confidence match what the tracker's predict-flight / predict_route_starlink actually returned (via the shared parsers in lib/tracker.mjs);
  • the four empty-state strings are chosen correctly (the Codex-approved state model: "No direct-flight Starlink history yet. Connection estimate below." vs "…for this route yet." vs "Direct-flight history unavailable right now." vs a rendered list);
  • a confirmed-tail mark appears only when the searched date is within about 3 days and a firm tail was named, never on a far date;
  • the two surfaces agree: the united.com panel number and the Google Flights / Navan chip number for the same flight are equal (endpoint parity, already proven at a point in time in Phase 1, here asserted per matrix cell).

Layer B: PREDICTION accuracy (calibration / backtest)

The tracker's published probabilities match realized outcomes.

For a prediction of p made at horizon h for (flight, date), we later learn the realized outcome: did the aircraft that actually operated it have Starlink. Aggregated over many predictions, p should be calibrated: of all flights predicted about 60%, close to 60% should have flown Starlink. This layer is blocked on outcome data (see section 6) and therefore reports, never gates. What is computable today is a partial, near-horizon version using the tracker's own check_flight firm-tail assignment as a proxy outcome (see section 5.3).

2. Sampling matrix (routes times horizons)

Reuses and widens test/routes.mjs. The design axes are metal class (retrofit maturity differs sharply) and horizon (how many days until departure, the tracker's answer changes as the assignment window opens about 48h out).

Routes (17 today; grouped so slices are meaningful):

  • hub-to-hub narrowbody: SFO-DEN, DEN-SFO, DEN-ORD, ORD-EWR, IAH-DEN, LGA-ORD, DEN-LAS, IAD-ORD. Expect direct history.
  • premium transcon (historically widebody/757): SFO-EWR, EWR-SFO, LAX-EWR, SFO-BOS, JFK-LAX. The false-empty-state hunting ground.
  • international widebody: SFO-LHR, EWR-LHR. Expect little or no direct history.
  • regional-heavy spokes: DEN-ASE, ORD-MSN.
  • Alaska slice (smaller, since AS odds are aircraft-type-derived, not per-flight): a handful of SEA-* routes to exercise the AS branch and the "prose note, no table" path.

Horizon buckets (days from sampled-on to departure): far (over 7d), mid (3 to 7d), near (3d or under, tail-assignment window open), same-day (0 to 1d). The matrix cell is (route, horizon-bucket); concrete dates are computed at runtime relative to today so the file never goes stale (as nearDates() already does).

Sizing. About 17 routes times 4 horizon buckets, roughly 68 route cells per sweep, plus a fixed SAMPLE_FLIGHTS set (UA1, UA2402, UA1596, UA2019, UA5693) for the per-flight / check-flight cross-check. That is the live-recording budget; the replay/gate layer costs zero network.

Politeness (non-negotiable: this is a guest on someone's server)

Inherited verbatim from lib/tracker.mjs, and the design forbids relaxing any of it:

  • one request at a time, global gate, THROTTLE_MS floor (default 1300ms) between the end of one call and the start of the next; never parallelised, no proxy.
  • honest identifying User-Agent: wifiodds-test-harness/1.0 (+https://wifiodds.com; …).
  • cache-busting for freshness checks only: a fresh query param when we must defeat a cache, never as a way to hammer. Live recording runs on a schedule, once, not in CI.
  • a hard per-run request ceiling, logged, so a matrix that grows cannot silently 10x the load.
  • every MCP response body is treated as inert data: regex-parsed or string-compared, never interpreted. The bodies carry instruction-shaped prose ("Render the table EXACTLY"); same discipline as the extension.

3. What it records per (route, flight, date, sampled-on)

The heart of the harness is an append-only snapshot ledger, one row per prediction observed. Predictions are made now; outcomes arrive later, so the row is written in two passes and never overwritten (a corrected outcome appends a new revision, it does not mutate the prediction). Stored as newline-delimited JSON under test/out/accuracy/snapshots/ and never committed (see the .gitignore note in section 7).

Prediction record (written at sample time):

{
  "schema": "wifiodds-accuracy/1",
  "sampled_on": "2026-07-30T04:10:00Z",   // when WE observed it
  "carrier": "UA",
  "flight": "UA1596",
  "route": { "o": "SFO", "d": "DEN" },
  "dep_date": "2026-08-02",
  "horizon_days": 3,                       // dep_date - sampled_on, integer
  "horizon_bucket": "near",
  "source": "predict-flight",              // predict-flight | predict_route_starlink | plan-route-leg
  "state": "scored",                       // scored | na | pending | error
  "prob": 0.68,                            // null unless state == scored
  "obs": 51,                               // n_observations, null if absent
  "confidence": "high",                    // high | medium | low | type (verbatim from tracker)
  "confidence_basis": "history",           // history | aircraft_type (derived from conf=="type")
  "data_ts": { "last_refreshed": null, "as_of": null },  // null until data contract Ask 1
  "raw_status": 200                        // HTTP status, for the drift log
}

State: the four values, and why each exists (they map 1:1 to bugs already shipped once):

  • scored: a real probability. Trust and display it.
  • na: the tracker genuinely has no history for this flight number. Safe to cache as n/a. Distinct from a false zero (see section 4, display-integrity).
  • pending: a near-date whose tail assignment has not published yet (about 48h out). Showing this as a fresh n/a is the "stale-as-fresh" bug.
  • error: a 429/500/garbled 200. Must never be recorded as na, or the backtest inherits a false negative and the display shows a false n/a. This mirrors the extension's PREDICT_ERR sentinel and directOk flag: the harness records the state the extension had to infer, so a later data-contract state field can replace the inference.

Outcome record (appended later, once realized truth is known):

{
  "schema": "wifiodds-accuracy-outcome/1",
  "flight": "UA1596",
  "dep_date": "2026-08-02",
  "resolved_on": "2026-08-04T00:00:00Z",
  "outcome_source": "check_flight_firm_tail",  // proxy today; "operated_backfill" once Ask 3 lands
  "actual_tail": "N73275",
  "actual_type": "B739",
  "wifi_state": "starlink",                     // starlink | non_starlink | unknown
  "roster_agreement": true,                     // did our equipped roster agree? null if tail unknown
  "outcome_confidence": "firm"                  // firm (published tail) | operated (post-departure truth)
}

The join key is (flight, dep_date). A prediction row and its outcome row are matched on that pair; a prediction with no matching outcome is pending in the backtest and excluded from the denominator: unknown is not zero, per the project doctrine. An outcome with wifi_state: "unknown" is likewise excluded, never assumed non-Starlink.

3.1 The proxy outcome available today

check_flight returns a firm tail for a near date (published about 48h out). That tail, joined to united/data.json, yields a yes/no Starlink verdict, a realized-enough outcome for near-horizon flights. This gives a limited backtest today: "for flights the tracker rated p a few days out, did the firm-assigned tail turn out to be Starlink?" It is honest but narrow: it scores the near-horizon assignment, not the far-horizon prediction, and it inherits any staleness in our roster. The full backtest (any horizon, post-departure truth) needs Ask 3.

4. Metrics

Two families: prediction-accuracy metrics (Layer B, mostly outcome-gated) and display-integrity metrics (Layer A, computable today).

4.1 Prediction-accuracy metrics (Layer B)

  • Reliability diagram / calibration curve. Bin predictions into deciles of p (0 to 10%, 10 to 20%, and so on). For each bin plot mean predicted p against observed Starlink frequency. Perfect calibration is the diagonal. Rendered as a table plus an SVG the audit report can embed.
  • Brier score. Mean squared error of p vs the 0/1 outcome, over all resolved scored predictions. One headline number, lower is better; also reported per slice.
  • Coverage of "confident" predictions. Of predictions the tracker labelled confidence: high (and separately basis: history vs aircraft_type), what fraction resolved correctly at a decision threshold (for example p ≥ 50% → predict Starlink). Surfaces whether "high confidence" is earned. type-derived odds are reported as their own cohort: they answer a different question (aircraft-type base rate, not this flight's history) and must never be pooled with history-based odds.
  • Error rate, sliced. Misclassification at the threshold, sliced by: route, horizon bucket (near vs far, the key question is whether far-horizon odds are worse), and fleet-rollout era (bucket by month, so improving retrofit coverage shows as drift over time, not noise).

Every metric carries its denominator and its exclusions in the report: how many predictions resolved, how many are still pending, how many unknown. A calibration curve computed on 40 resolved points is labelled as such and never presented as if it were 4,000. This is the "publish the floor, not the blend" rule applied to the harness's own output: a metric nobody can source yet is written as "not enough resolved outcomes," never as a confident zero.

4.2 Display-integrity metrics (Layer A, today)

These check the rendering, and each maps to a real incident:

  • No false zeros. A 0% badge is legitimate only when the tracker returned a real probability: 0 with observations (a genuine "this flight never gets Starlink, 51 obs"). A 0 or n/a synthesized from an error state, or from an unpublished count, is a false zero: the harness flags any badge whose 0/n/a is not backed by a scored or na state.
  • No stale-as-fresh. Once the data contract exposes last_refreshed/as_of, a value older than the display TTL rendered without a staleness cue is flagged. Today (no timestamps) the proxy is the 6h cache TTL: a badge served from a cache entry older than TTL that is not force-refreshed is flagged.
  • No absence-from-failure. The error → "unavailable" copy must appear on a failed direct-history fetch; the harness fails the cell if an error state renders as "No direct-flight Starlink history…" (the exact state-model bug from the 29 Jul findings).
  • Surface parity. united.com panel % equals GF/Navan chip % for the same flight, per cell.

5. Drift / canary behavior: fail closed

The harness must fail closed: when it can no longer trust that it is reading the tracker correctly, it stops and shouts, rather than emitting a clean-looking metric computed on garbage. Three drift classes, each with its own canary:

  1. Schema drift (tracker). The shared parsers in lib/tracker.mjs mirror extension/bg.js. If predict-flight renames probability, or predict_route_starlink's table format changes, the parser yields null/zero rows. Canary: a fixed known-good probe (SAMPLE_FLIGHTS on a route with reliable history) that must parse to a scored state with a plausible shape. If the probe comes back unparseable or all-na, the run aborts with SCHEMA-DRIFT and records no snapshots: recording a sweep of false nas would poison the ledger permanently.
  2. Row-recognition / route-extraction drift (extension). Layer A drives the real content script. If united.com/Navan/GF markup assumptions break (rows not found, route context mis-extracted) the panel renders nothing or the wrong route. Canary: the deterministic positive-control cell (a SFO-DEN-style fixture with a known badge) must still render its known badge, exactly as Phase 2's negative-control (E2E_NEG) already proves a seeded regression makes the gate exit 1. If the positive control fails, the gate fails.
  3. Parser drift between harness and extension. lib/tracker.mjs is a hand-copy of bg.js's regexes; they can silently diverge. Canary: a small golden-fixture test that feeds a frozen set of real tracker response bodies through both the harness parser and (in the browser layer) the extension, and asserts identical decoded output. If bg.js changes a regex and the mirror does not, this fails, the same "a rule that lives where the runtime does not read it is not a rule" doctrine, applied to the parser copy.

How it signals. Two channels, matching the project's existing shape:

  • Report: test/out/accuracy/report.md (human) and accuracy-findings.json (machine), severity-ranked like Phase 1 (HIGH/MEDIUM/LOW/INFO). A drift is HIGH.
  • Exit code: Layer A (display plus canaries) participates in the release gate and exits 1 on any failed check, drift included. Layer B (calibration) reports and never gates, because a calibration metric moving is a finding to investigate, not a reason to block a ship, and because it depends on data we do not yet control. A red drift canary, though, blocks: better no metric than a false one.

Fail-closed is literal: on any drift canary trip, the live-recording pass writes zero snapshots and the ledger is left exactly as it was. Same doctrine as ship.sh and reconcileUnited(): a process that corrupts the record unattended is worse than one that does nothing and says so.

6. Dependencies on the @martinamps data contract

Grounded in martinamps-data-contract-proposal.md (the three asks: explicit states, a changed-since delta feed, post-departure outcome backfill). What each unblocks:

Harness capabilityNeedsBuildable today?
Layer A display checks, all four statesnothing (state is inferred, as the extension does)Yes
Snapshot ledger of predictionsnothingYes
Endpoint / surface parity per cellnothingYes
Schema/row/parser drift canariesnothingYes
Near-horizon proxy backtest (firm tail vs roster)check_flight + united/data.json (both exist)Yes, but narrow
Trustworthy state (no client inference)Ask 1: versioned record with explicit PUBLISHED / NO_HISTORY / PENDING / UPSTREAM_ERRORNo
Stale-as-fresh metric (real timestamps)Ask 1: last_refreshed vs as_ofNo (proxied by cache TTL today)
confidence_basis without string-matching "type"Ask 1: structured basis: history | aircraft_typeNo (derived today)
Poll the tracker a fraction as oftenAsk 2: changed-since / delta feedNo
Full calibration / Brier / reliability at any horizonAsk 3: post-departure outcome backfill (actual tail + type + WiFi state after departure) plus a queryable sample windowNo: the core of Layer B is blocked here

The honest headline: the reliability diagram, Brier score, and any-horizon error rate, the metrics that make this a trust backstop rather than a display test, are blocked on Ask 3. What ships today is the machinery around them: the sampling, the append-only ledger, the display integrity, the drift canaries, and a narrow near-horizon proxy backtest. The day outcome backfill exists, Layer B lights up against a ledger that has been accumulating real predictions the whole time, which is the entire reason to start recording snapshots now rather than waiting.

Note also the edge-cache / proxy question (the sibling proxy note): if an edge cache is ever placed in front of the tracker, the harness must record whether a snapshot came from origin or edge, because a cached edge response can be stale in a way curl cannot see, the same "verify by response body, and the edge can rewrite what the repo serves" lesson.

7. Shape of the build

A node script plus fixtures, living under test/, consistent with the two existing phases.

test/
  phase3-accuracy.mjs        # entry: --record (live, scheduled) | --replay | --backtest | --gate
  routes.mjs                 # EXTEND: add horizon buckets + Alaska slice (exists today)
  lib/
    tracker.mjs              # reuse: polite client + mirrored parsers (exists)
    roster.mjs               # reuse: roster join (exists)
    ledger.mjs               # NEW: append-only snapshot read/write (NDJSON), two-pass join
    calibration.mjs          # NEW: binning, Brier, reliability-diagram SVG (pure, deterministic)
    canaries.mjs             # NEW: schema / parser golden-fixture checks
  fixtures/
    accuracy/                # NEW: frozen real response bodies for replay + golden parser tests
  out/
    accuracy/
      snapshots/*.ndjson     # the ledger (gitignored; may hold many days of predictions)
      report.md              # human
      accuracy-findings.json # machine, severity-ranked
      reliability.svg        # calibration curve for the audit report

Modes (one script, four verbs):

  • --record: the only networked mode. Sweeps the matrix once against the live tracker, politely, and appends prediction snapshots. Runs the drift canaries first and writes nothing if they trip. Scheduled, not in CI (for example a daily wave), because it costs real requests to someone else's server and because a backtest needs snapshots spread across days/horizons.
  • --replay: Layer A display checks against recorded fixtures in a real browser (the Phase 2 engine). Deterministic, offline, no tracker contact.
  • --backtest: Layer B. Joins the accumulated ledger to outcomes (proxy today, contract feed later), computes calibration/Brier/coverage/error, writes the report plus SVG. Pure and offline.
  • --gate: runs --replay plus the canaries plus the display-integrity checks and exits 1 on any failure. This is the CI-safe subset.

When each runs:

  • --gate: on every extension change that touches content.js/bg.js/parsers, alongside Phase 2. Fast, deterministic, offline. Blocks the merge/ship on display or drift failure.
  • --record: scheduled (daily), politely, one sweep. Feeds the ledger. Never in CI.
  • --backtest: scheduled (weekly, or on demand) after enough outcomes have resolved; also re-run whenever the outcome feed backfills. Produces the reliability report for the audit.

Determinism where possible. --replay, --backtest, --gate are fully deterministic: same fixtures/ledger yields same output, same as Phase 2's fixed-fixture doctrine. Only --record touches the network, and its output (a timestamped append) is the one intentionally non-deterministic artefact: it is data, not a check.

Never invent a figure. If the ledger has no resolved outcome for a slice, the report writes "no resolved outcomes yet" for that slice, not 0, not an interpolation. A calibration bin with zero points is shown empty, not filled from a neighbour. Every published metric carries its denominator, its as-of date, and its exclusions (pending, unknown), exactly as the site's data rules require. --record writes to out/ only; it never writes to wifiodds or the extension repo.

8. Non-goals / guardrails

  • Not a load test, not a scraper. It never parallelises, never proxies, never removes the throttle. Growth in the matrix must move the logged request ceiling, visibly.
  • Never interprets tracker prose. Response bodies are inert data. No body ever reaches a code path that could act on it.
  • Does not mutate the ledger. Outcomes append; corrections append revisions. The prediction as-observed is immutable, that is what makes the backtest honest.
  • Does not close its own findings. A calibration drift or a display bug is a finding for the audit trail; the builder ships a fix and requests verification, the auditor clears it. The harness's job is to measure and record, loudly and reproducibly, not to declare itself green.
  • Reads only. It consumes the tracker, united/data.json, and the extension source; it writes only under test/out/accuracy/.

v2.4 flagship · spec

Guard & Rescue

Build-ready specification. No code in this document; it names the existing functions, files, storage keys and message types so a build session can wire against them directly. Read-only research pass over ~/Projects/united-starlink-companion/extension on 2026-07-30.

One-line promise: Guard a flight you have already booked, get warned if the WiFi outlook changes, and, when it worsens, see the best still-available alternative you had already been shown. The extension only ever informs; it never touches your booking.


0. Honest starting point: most of the plumbing already exists

This is the single most important framing for the build. "Guard & rescue" is not a greenfield feature. The v1.6 Tail-swap Guardian already ships a booking-to-boarding watch, and v2.4 is a reframing plus two additions on top of it, not a rewrite.

What already exists today (v1.4 to v2.2, all in extension/bg.js, extension/content.js, extension/popup.js, extension/popup.html):

CapabilityWhere it lives today
Watch star on every scored result row (/ toggle, "UA1812|2026-07-25" key)content.jsaddWatchStar(), watched Set
Guarded-trip storage (local only, no accounts)bg.jsTRIPS_KEY = "uslTrips", getTrips()/setTrips(), MAX_TRIPS = 10
Add / remove / list / check-now messagingbg.js handlers tripAdd, tripRemove, tripList, tripCheckNow
Re-check on an alarm (no polling loop)bg.jschrome.alarms.create("uslTripCheck", { periodInMinutes: 180 }) + onStartup
Tail-swap detection with per-trip historybg.jsapplyCheckResult() state machine, HISTORY_CAP = 20
Desktop notifications on four transitionsbg.jsNOTIFY_TRANSITIONS, buildGuardNotification(), notifyTrip()
Best-alternative suggestion (same-day ✓ tail, then parsed alts)bg.jssuggestAlt()
Politeness budget + backoffbg.jsGUARD_BUDGET = 100/day, budgetTake(), invalidCount halt, isTerminal()
Popup "Guarded trips" view with per-trip timelinepopup.jsrenderTrips(), tripLine(), renderHistory()
The three data endpoints the whole feature rides onbg.jscheck_flight (MCP), predict-flight (REST), plan-route/search_starlink_flights

So what is genuinely NEW in v2.4:

  1. A three-state honest notification model that replaces the current four-plus-transition scheme with exactly three user-facing states, each of which is defensible against the auditor's "unknown is not zero" rule.
  2. A saved alternatives shortlist: the alternatives the user actually saw at booking time, stored with the trip, so a "rescue" suggestion is grounded in options the user already considered rather than a fresh, possibly-empty query.
  3. A route back to the booking surface carried on every notification.
  4. A decision-strip save affordance (optional; the star already covers the minimal path) so a guard can be created with its shortlist in one gesture.

The build should treat sections 3 to 8 as deltas against the existing code, not as new subsystems. Where a behaviour already exists, the spec says so and points at the function to extend rather than replace.

1. User story & promise

As a traveller who has already booked a United or Alaska flight, I want to "guard" it so that if the aircraft assignment changes and my WiFi outlook gets worse, I am told, clearly, honestly, and with the best still-available alternative I had already been shown, without the extension ever changing my booking for me.

Three sub-stories, in priority order:

  • Guard (minimal). One click on the next to a flight's odds badge, or one entry in the popup, registers a flight+date to watch. Already shipped.
  • Re-check (the "guard" half). The extension re-queries the tail assignment as departure approaches and again whenever it detects a change, and tells me which of three honest states my flight is in. Mostly shipped; the honesty reframing is new.
  • Rescue (the "rescue" half). When the outlook worsens, the notification and the popup show me the best still-available scored alternative from the shortlist I saw when I booked, a suggestion, never an action. New.

Non-goals, stated up front (auditor-facing):

  • The extension never rebooks, never holds, never selects a fare, never opens a checkout. "Rebooking-assist" means surfacing a better option and a link back to the booking surface, full stop. This mirrors the hard rule in SURFACES.md ("Never inject into checkout, payment, or auth pages") and the content-script GF_DENY_PATH gate.
  • No accounts, no server, no per-user tracking, no analytics beacon. All state is chrome.storage.local. Flight-number + date is still the only registration input, exactly as the v1.6 comment in bg.js promises.

2. Save flow

2.1 What is saved today

newTrip(fn, date, route) in bg.js stores:

{ fn, date, route, added, history:[], asOf, lastError, lastNotifKey,
  invalidCount, departs, lastStatus, tail, prob, typeDerived, equip, alts, routeSeen }

route is the only booking-context field, and it is a bare "SFO-SEA" string. There is no record of the alternatives the user was looking at when they guarded the flight.

2.2 What v2.4 adds: an optional shortlist

Extend newTrip() with one new optional field:

shortlist: [ { fn, date, route, prob, obs, conf, source, savedAt }, … ]   // ≤ 5 entries, may be []
  • prob/obs/conf are copied at save time from probMap (content.js) or the route/predict-flight data (popup), a frozen snapshot of what the user saw, not a live value. This matters: the rescue suggestion must be honest about being "what you saw when you booked," and re-scored live only at rescue time (section 4.4).
  • source records where it came from ("united", "navan", "gflights", "popup") for the route-back link (section 5).
  • The shortlist is capped at 5 and is allowed to be empty. An empty shortlist degrades gracefully to the v1.6 suggestAlt() live-query path (same-day ✓ departure → parsed alts → generic advice). This is the phasing seam: v2.4.0 ships with shortlist always []; the full rescue populates it.

2.3 Where the save hooks in

Three entry points, in ascending effort:

  1. Result-row star (shipped, minimal). addWatchStar() already sends tripAdd { fn, date, route }. For v2.4.0 this is unchanged and shortlist stays empty. For the full rescue, the star's click handler additionally gathers the other scored on-page flights for the same route/date from probMap / registry (top 5 by prob, excluding fn itself) and passes them as shortlist on the tripAdd message. This is the cheapest place to capture a real shortlist because content.js already has every on-page flight's odds in probMap.
  2. Popup "Guarded trips" add form (shipped). watchForm submit sends tripAdd { fn, date } with no route and no shortlist. For the full rescue, if the active tab is a booking surface for the same route, the popup can attach pageFlights / lastData.flights as the shortlist. Optional; the popup path may always ship an empty shortlist and rely on the live-query fallback.
  3. New "decision strip" (full rescue only, optional). A one-line strip rendered under the injected panel on a result page: "Guard this trip, we'll watch the tail and flag a better option if it slips." with a single Guard button that captures { guarded flight, date, top-5 shortlist } in one gesture. This is a nicety, not a requirement; the star already covers the minimal path. If built, it reuses the existing .usl-panel styling and the same tripAdd message, no new message type.

2.4 Validation (unchanged)

tripAdd already validates ^(?:UA|AS)\d{1,4}$ + ^\d{4}-\d{2}-\d{2}$, rejects past dates, and enforces MAX_TRIPS. Shortlist entries are validated with the same flight-number regex and silently dropped (not rejected) if malformed, a bad shortlist must never block guarding the primary flight.

3. Re-check logic

3.1 What exists: reuse it, do not add polling

runTripChecks() / runTripChecksInner() in bg.js already implement the whole re-check engine, driven by one alarm (uslTripCheck, 180-minute period) plus onStartup and the popup's manual tripCheckNow. There is no setInterval polling in the service worker and the spec must not add one, the alarm is the single scheduler.

The existing cadence inside runTripChecksInner():

  • Trips ≤ 4 days out are checked every run (every about 3h).
  • Trips farther out are checked at most daily (d > 4 && now - t.lastChecked < 24h → skip).
  • Expired (daysUntil < -1), terminally-invalid (invalidCount ≥ 2), and published-and-departed (isTerminal()) trips are skipped.

3.2 v2.4 deltas

  1. T-48h emphasis (mostly already there). Tail assignments publish about 48h out (documented throughout the codebase). The existing ≤ 4 days → every run band already guarantees multiple checks across the T-48h window. v2.4 formalises this as the contract ("at least one re-check between T-48h and departure") and nothing in the cadence needs to change to honour it. Do not add a dedicated T-48h alarm, it would be a second scheduler for a window the 3h alarm already covers.
  2. Re-check after a detected change (already there). applyCheckResult() already appends a history entry and re-notifies whenever status or tail differs from the newest entry. Because the trip stays in the active set until isTerminal(), the next alarm run naturally re-checks after any change. No change-triggered extra fetch is needed; the swap is caught on the following scheduled pass, which is the correct politeness posture.
  3. Bounded / backoff: reuse the v2.2 sentinel discipline. The re-check path must inherit the same failure discipline already proven elsewhere:
    • GUARD_BUDGET = 100 MCP calls/local-day via budgetTake(): when exhausted, checks are skipped, trips go stale (asOf shown in popup), no state loss.
    • invalidCount ≥ 2 halts checks on a bad flight number.
    • A transient failure (check_flight throws / unparseable) returns { status: "unknown" }, which applyCheckResult() treats as transient: it is never stored as lastStatus, never a transition, only sets lastError and leaves asOf stale. This is the exact analogue of the content-script PREDICT_ERR sentinel and the route-fetch backoff (ROUTE_BACKOFFS): an attempted-but-failed check must be distinguishable from a genuine negative, and must never be cached as a fact. The v2.4 build must preserve this and must not let a network blip flip a trip to any of the three honest states.

4. Notification states: exactly three, honest

This is the heart of v2.4 and the part most exposed to the auditor. The current code emits four notify transitions (publish-yes, publish-no, swap-lost, swap-gained) plus finer internal transitions. v2.4 collapses the user-facing surface to exactly three states so that every notification maps to one defensible claim.

4.1 The three states

#StateFires whenHonest claim
AStarlink confirmedAn authoritative tail is assigned and that tail is a known Starlink aircraft (status: "yes")"Your assigned aircraft has Starlink."
BNot Starlink, or unknownA tail is assigned and it is not Starlink (status: "no"), or the aircraft is known but its WiFi cannot be confirmed either way (type-derived / ambiguous)"Your assigned aircraft does not have Starlink" or "…we cannot confirm Starlink for your aircraft." Never a bare "0%."
CAssignment unavailableNo tail is published yet (status: "early"), the flight number cannot be resolved, or the check has not succeeded (outage / budget exhausted / transient)"No aircraft assigned yet, we'll keep watching." The extension is explicit that it does not know, rather than implying a negative.

Why three, and why this grouping: the auditor's standing findings include "Unknown is not zero" and "verify by response body, not status code." State B deliberately fuses confirmed-non-Starlink with known-aircraft-but-unconfirmed because both are honestly summarised as "do not count on Starlink," and neither should ever be dressed up as a confident "0%." State C is the pressure-relief valve that keeps a transient failure or an unpublished assignment from being mis-shown as B. A trip must never present state A or B off a status: "unknown" result, that is exactly the applyCheckResult() "unknown is transient" guarantee, and it is load-bearing here.

4.2 Mapping the existing state machine onto the three

The existing applyCheckResult() transitions map cleanly; do not rip it out, add a notifyState(transition, res) classifier that folds transitions into A/B/C:

Existing transitionv2.4 stateNotify?
publish-yes, swap-gainedAYes
publish-no, swap-lostBYes
swap-yes-yes (tail changed, still ✓)A (informational)Yes, tail changed, reassure
swap-no-no (tail changed, still ✗)BYes, still worth re-offering a rescue
withdrawn (was assigned, back to early)CYes, outlook regressed to unknown
first-early, noneC / no-opNo notification (timeline only)
unknown (transient)no state changeNever

The "worsened" predicate that gates the rescue payload (section 4.3) is: A→B, A→C, or B re-confirmed with a different, still-non-Starlink tail (swap-no-no). B→A and C→A are improvements and carry no rescue.

De-dup stays as-is: lastNotifKey = transition + "|" + tail already prevents an identical re-fire; the classifier keys off the same value.

4.3 What every notification carries

Each of the three notifications includes:

  1. The trip identity: fn + date (already in buildGuardNotification()).
  2. A route back to the booking surface (section 5), new.
  3. The best still-available scored alternative, only when the outlook worsened (state A→B, A→C, or worse-tail B). New wiring around the existing suggestAlt().

State A (improved/confirmed) and state C-from-unpublished carry no rescue line, surfacing an alternative when nothing got worse would be noise and would undercut the honesty of the feature.

4.4 Sourcing the "best still-available alternative"

Preference order, extending the existing suggestAlt():

  1. From the saved shortlist first (full rescue). Re-score each shortlist entry live at rescue time via the cached predict-flight / check_flight path (cache-first, so usually free against the budget). Pick the highest live prob that (a) beats the guarded flight's current outlook and (b) is still a plausible same-day option. Present it as "Better option you saw: UA1234, now about 72%." The shortlist is what makes this a rescue rather than a guess: it is grounded in flights the user already considered.
  2. Fall back to suggestAlt() live query when the shortlist is empty (v2.4.0 always) or yields nothing: same-day confirmed ✓ departure on the same route (getRouteDatadeps), then the parsed alts table from check_flight, then generic advice. This path already exists and is budget-aware.
  3. If nothing is found, say so: "No better same-day option found." Never invent one. (Auditor rule: never invent a figure.)

The extension never rebooks. The alternative is a labelled suggestion plus the route-back link; acting on it is entirely the user's, on the airline's own site.

5. Route back to the booking surface

New, small, and shared by all three notifications and the popup.

  • Store on each trip (and each shortlist entry) a source and, where available, the origin URL the guard was created from (for example the united.com results URL, captured in content.js at star-click time; the popup's activeTab.url).
  • On notification click (chrome.notifications.onClicked), open (or focus) that URL in a tab. If no URL was captured, fall back to the carrier's search page for the stored route + date (deep-link when the URL shape is known, else the carrier home). Never fabricate a booking-confirmation deep link.
  • In the popup "Guarded trips" row, the route-back is a small link/affordance on the trip's main line.

Add one message/handler (chrome.notifications.onClicked listener in bg.js), there is none today. Keep it fail-silent (wrapped in try/catch like every other bg.js entry point).

6. Data needed, and what is blocked on the @martinamps data contract

6.1 The four fields v2.4 wants per re-check

FieldNeeded forAvailable today?Source today
Assignment timestamp (when the tail was assigned)Honest "as of" + detecting fresh changes vs stale dataNO (inferred only): the extension times changes by when its own poll first saw them (t.asOf, history[].ts), not by an authoritative assign timederived locally
Tail / aircraft typeStates A/B, swap detectionYEScheck_flight MCP → parseCheck() tail, equip
WiFi state (Starlink yes/no/type-derived)The three statesYES (UA/AS), partialcheck_flight status + predict-flight confidence
Change historySwap timeline, "worsened" predicateYES, but locally builtapplyCheckResult() appends to history[] from successive polls

6.2 What today's endpoints do and do not give

Give (sufficient for v2.4.0 on UA + AS):

  • check_flight (MCP, per airline host): status in {yes, no, early, invalid, unknown}, tail, equip, route, Departs …Z (a departure time), and a parsed alts table. This is the tail/outcome feed the guard rides on.
  • predict-flight (REST): probability, n_observations, confidence (including "type" = aircraft-type-derived, not per-flight history).
  • plan-route / search_starlink_flights: itineraries and confirmed ✓ departures used by suggestAlt().

Do not give (the gaps):

  • Authoritative assignment timestamp. The tracker returns the departure time, not when the tail was assigned. So "the assignment changed" is really "our poll noticed a change," bounded by the 3h-or-under cadence. Honest, but coarser than a push feed.
  • Push / event feed. Everything is pull. A change is caught on the next scheduled poll, not the instant it happens.
  • Post-flight actual outcome. No "did it actually fly with Starlink" signal; the guard is predictive-up-to-departure only.
  • Per-flight coverage beyond UA/AS. Fourteen of eighteen carriers have no per-flight API; Hawaiian is tracked but publishes no per-flight probability (probe transcript in bg.js). Guarding is UA/AS-only, by design.

6.3 The @martinamps data contract: what it would unblock

SURFACES.md records that Martin Amps "owns real data and real distribution" (a UA Google-Flights Starlink indicator, currently dormant) and that collaboration is worth pursuing. A formal outcome/tail feed from that source is the dependency that upgrades Guard & Rescue from "honest but pull-based" to "authoritative":

  • Unblocks: a real assignment timestamp (so state changes are timed by the source, not the poll), potentially a push/webhook change feed (so swaps are caught immediately instead of within 3h), and possibly post-flight outcome (closing the loop on whether a guarded flight actually delivered Starlink).
  • Blocked until it exists: anything that claims an assignment time rather than a last-checked time; any "instant" swap alert; any post-flight reconciliation. v2.4 must not print an assignment timestamp it does not have, it shows asOf (last successful check) and labels it as such, per the existing popup "as of …" convention.

Build guidance: design the re-check result shape with an optional assignedAt field that is null today and populated only if/when the @martinamps feed lands. Every template must read it as "unknown → show asOf instead," never substitute the departure time or Date.now() for it. (This is a direct application of the "unknown is not zero / never invent a figure" doctrine.)

7. UX

7.1 Popup "Guarded trips" view (extend, don't replace)

The section exists: popup.html "Guardian · booking-to-boarding tail watch", rendered by renderTrips() / tripLine() / renderHistory(). v2.4 changes:

  • Retitle to "Guarded trips" (the user-facing name for v2.4), keeping the same DOM ids so popup.js bindings survive.
  • State chip per trip using the three-state model: A = green "Starlink ✓", B = coral "No Starlink ✗ / Unconfirmed", C = dim "Awaiting assignment" or "Unavailable, as of …". Reuse the existing .usl-t-yes / .usl-t-no / .usl-t-early / .usl-asof classes; no new palette.
  • Rescue line under a worsened trip: "Better option you saw: UA1234 (about 72%)" with the route-back link. Reuses tripLine()'s existing better: rendering, now sourced from the shortlist when present.
  • The per-trip history timeline (renderHistory()) is unchanged.

7.2 Alert copy (the three states)

Extend buildGuardNotification() to emit exactly these shapes. Title carries the state; message carries the honest claim + (when worsened) the rescue + route-back.

  • A, confirmed:
    Title: 🛰️ UA1812 · Jul 25 · Starlink confirmed
    Body: Tail N127UA has Starlink. You're set. (no rescue line)
  • A, reassuring swap (still ✓):
    Title: 🛰️ UA1812 · Jul 25 · tail changed, still Starlink
    Body: New tail N201UA also has Starlink. No action needed.
  • B, not Starlink / unconfirmed:
    Title: ✗ UA1812 · Jul 25 · no Starlink
    Body: Assigned tail N840UA (Viasat). Better option you saw: UA1234 (~72%). Open booking ↗
  • C, unavailable / awaiting:
    Title: ⏳ UA1812 · Jul 25 · no assignment yet
    Body: Aircraft not assigned yet (tails publish ~48h out). We'll keep watching. Open booking ↗

Copy must pass build/slop-gate.js (the prose ratchet), keep it plain, no em dashes where a period will do, and load the no-slop skill before finalising strings that ship.

7.3 Watch-star affordance (reuse verbatim)

The page-row / star (addWatchStar()) is the guard entry point and does not change shape. Its two titles already read as a guard promise ("Guard UA1812 … alerts from booking to boarding if its Starlink tail changes"). The only v2.4 change is the optional shortlist capture on click (section 2.3). The read-only panel marker (.usl-guarded gold ★, GUARD_MARK) stays as the "this row is guarded" indicator.

8. Edge cases & safety

  • No per-user tracking / no analytics. Nothing about guards, shortlists, or notifications leaves the browser. No beacon, no server call except the same tracker MCP/REST endpoints the extension already uses for odds. (Directly responsive to the auditor's edge-injected-beacon P0.)
  • Local storage only. All guard state is chrome.storage.local (uslTrips, uslGuardBudget, selector cache). No sync storage, guards do not follow the user across devices, and that is intentional (no account model).
  • Data unavailable. State C absorbs every failure: MCP outage, budget exhausted, unparseable body, flight-not-found. The trip shows "as of <last good check>" and keeps its last real state in the timeline; it never silently flips to A or B on a failed check (applyCheckResult() "unknown is transient").
  • Assignment genuinely 0% / no coverage. Never render a bare "0%." A known-non-Starlink aircraft is state B with the aircraft named; an unknown aircraft is state C. This mirrors the site-side equippedPublished / pctEquipped()==null doctrine in CLAUDE.md.
  • Shortlist goes stale. Shortlist odds are a save-time snapshot; the rescue path re-scores live before presenting, and drops any entry that no longer beats the guarded flight or cannot be re-scored. A stale shortlist can only ever fail to offer a rescue, never offer a wrong one.
  • Unguard flow. Unchanged and available in three places: the page star (toggle off → tripRemove), the popup trip row × button, and implicitly on expiry (daysUntil < -1 → filtered out in runTripChecksInner). Removing a guard also discards its shortlist. Optimistic UI in addWatchStar() stays: flip the star first, tell bg.js after; both handlers are idempotent.
  • Budget starvation is safe, not silent to the user. When GUARD_BUDGET is exhausted, the popup's "as of …" line is the visible signal; no notification is suppressed silently into a wrong state, a skipped check leaves the prior state and its timestamp.
  • Checkout/auth pages. The decision strip (section 2.3) and star obey the same GF_DENY_PATH / never-inject-on-checkout rule as the rest of the content script; guarding is only ever offered on search/results surfaces.

9. Phasing

v2.4.0: Guard (minimal, shippable alone)

  • Re-check on the existing uslTripCheck alarm (no new scheduler, no polling).
  • The three-state notification model (A / B / C) replacing the four-transition copy, with the "unknown is transient" guarantee preserved.
  • Route-back link on every notification (onClicked handler + captured source URL, falling back to carrier search).
  • No shortlist. shortlist is always []; the rescue line uses the existing suggestAlt() live fallback only, and is shown only on a worsened state.
  • Popup "Guarded trips" retitle + three-state chips.

This is a small, self-contained delta over v2.2 and carries no new data dependency, it works entirely on today's check_flight / predict-flight endpoints.

v2.4.x: Rescue (full)

  • Saved shortlist: capture top-5 scored on-page alternatives at guard time (star + optional decision strip + popup-when-on-route).
  • Shortlist-first rescue sourcing (section 4.4), live-re-scored, with the suggestAlt() path as fallback.
  • Decision-strip save affordance (optional UI).

Blocked on the @martinamps data contract (post-v2.4)

  • Authoritative assignment timestamp (assignedAt) → honest "assigned at" instead of "as of last check."
  • Push / event change feed → near-instant swap alerts instead of 3h-or-under poll.
  • Post-flight outcome reconciliation → closing the loop on whether a guarded flight actually delivered Starlink.

Until that feed exists, v2.4 ships fully on the pull-based tracker endpoints and labels every time it shows as a last-checked time, never an assignment time.

10. Build checklist (pointers, not code)

  • bg.js newTrip(): add optional shortlist: [], source, sourceUrl, assignedAt: null; extend migrateTrips() to default them on old trips.
  • bg.js: add notifyState(transition, res) classifier folding existing transitions into A/B/C; keep applyCheckResult() and NOTIFY_TRANSITIONS.
  • bg.js buildGuardNotification(): rewrite copy to the three states + route-back; gate the rescue line on the "worsened" predicate.
  • bg.js suggestAlt(): add shortlist-first branch (live re-score, cache-first, budget-aware); keep the existing fallback chain.
  • bg.js: add a fail-silent chrome.notifications.onClicked handler for route-back. Do not add a new alarm or any setInterval.
  • content.js addWatchStar(): optional shortlist capture from probMap on the tripAdd message (full rescue only).
  • popup.js renderTrips() / tripLine(): three-state chips + rescue line + route-back link; popup.html retitle to "Guarded trips" (keep DOM ids).
  • Copy through build/slop-gate.js; load the no-slop skill first.
  • Ship via bash build/ship.sh "…" only; leave any auditor finding OPEN until the auditor clears it on the shipped artefact (per CLAUDE.md).