This last act happened while this post was being edited, and I’m including it because it runs the entire methodology end to end in six hours, like a closing worked example. It started with me re-reading act 11’s prose and getting suspicious of my own project: wait — are we only doing the density hack on the highway? I checked the source. Yes. Every shipped density dial read highway density; nothing priced the field. That worry became a ledger item; the ledger item spawned the cheap probe from the postscript above, which measured Sol’s slab tax at 4× and made a fix benchable.
The fix was mine to design, and the constraint was the covenant from the Dark Cross chapter: fully geography-agnostic. Build it as if the desORt and the Red Barrier (another of StellarForge’s density cliffs, a horizontal one) don’t exist, so the same machinery catches whatever artifact the community digs up next. The design descends from something we already owned — act 3’s oracle proved a four-byte-per-cube density summary of a whole galaxy is cheap to build and carry — though the shipped seed reads the index’s own cells directly: star counts, highway degree, the goal field’s keep filter. At plot start, ask that map for the nearest chainable concentration within a sane angle of the goal bearing, and let the opening heuristic learn exactly one detour — toward that on-ramp — with everything gated on measured density contrast, so a healthy start disengages and produces the identical route, byte for byte. My first formulation of the heuristic bias was, in the proud tradition of this post, mathematically inert: the triangle inequality guarantees that a path via anywhere is never shorter than the straight shot, so taking the minimum with the plain straight-line distance changes nothing at all. The other machine caught that during implementation and built the form that actually works, a gated raise. Three gates shipped, and each one was bought with a measured failure — two early forms died on spot-benches against the off-ramp twins, one fell to the full matrix (that rejection archived, per the covenant on negatives), before the survivor came back with 56 of the 60 benchmark cells identical and every reproducible loss cleared.
What it bought: Sol → Sagittarius A* — the cell that started the whole item that morning — went from 69 jumps to 67 in nearly half the wall-clock time. The Mandalay’s Sol runs shed jumps and stops. And one result we banked rather than claimed: the wall on Wongi → Beagle Point fell from 60 seconds to 17.7 on a byte-identical route in the same sweep — which would make a lovely trophy for the seed, except the seed disengages on that cell (the goal field out there is too thin; the gradient gate says no). Same answer, three times faster, cause not yet in evidence: a mystery written down and parked for act 17, where it gets a real explanation. One cost, named in the ledger because trophies without their costs are advertising: a single Mandalay cell pays 44 extra seconds on the fitted pilot-seconds score — about half a percent — reproducibly. The seed helps routes flying into concentration and never fires flying out of it; the kill switch exists; and the code names no slab, no cross, no bubble.
The worry arrived in the morning; the probe had its 4× by lunch; the idea arrived over a shoulder in the early afternoon; the gauntlet handed back its honest rejections; and by dinner there was a shipped, matrix-gated, geography-agnostic default. That is the whole post run in a single day — and the idea was mine.
One more, because the machines refused to leave the hardest matchup alone while I was busy adding jokes about keychains. The small ship’s outbound climb to Beagle Point — the loss this post’s closing notes have been calling “physics wearing a grudge” — had three separate ledger entries blaming the same suspect: somewhere in the middle of the sixty-thousand-light-year crossing, the ascending search picks a worse lane than the descending one. Three entries, same diagnosis, zero measurements of it.
Then one small tool laid the two routes’ hops side by side by position along the corridor, and the diagnosis fell over in a single table: the mid-corridor lanes cost equal jumps. The entire twenty-three-jump asymmetry lived in the final five percent — the rim entry, where the ascending route arrived at Beagle’s doorstep flying 3,815 plain light-years at zero percent boost while the descending route crossed the same neighbourhood with a third of the plain flying — 1,289 light-years against 3,815, 31 jumps against 55.
Sharp readers will recognize the ghost. The ramp-decomposition hypothesis — buried in the closing notes with a number on its headstone — held that hard asymmetric routes are hard at their ramps. Its metric died fairly. Its hunch, for this route, turns out to have been pointing at the right five percent all along; it just needed a different instrument to see it, which is a distinction the ledger cares about: we bury verdicts, not directions. And the fix that survived is a lesson this campaign has now learned three times: not a hand-tuned bias (three of those regressed somewhere on the matrix, and every one of its engagement heuristics was out-judged by the thing we already had) but the on-ramp seed’s mirror image — an off-ramp awareness, run as portfolio twins where the pilot-seconds judge is the only gate. The gauntlet came back the cleanest of any change today: 56 of the 60 benchmark cells identical, zero regressions, and the grudge match paid out 329 jumps and 115 stops against 342 and 133 — roughly forty-five minutes of a real pilot’s evening, recovered from the last five percent of the map.
And the epilogue to the epilogue, because the fuel probe went back in afterwards to check on an old promise. Months… fine, hours ago, I had told the machines to prefer fuel stars as close to the highway as possible, and the direct implementations of that instruction kept benching inert — the ranking knobs ship neutral to this day. The re-probe on the new route found the instruction had been delivered anyway, by a fix aimed at something else: the ascending route’s stop-adjacent detour signature collapsed from 3.4 points of excess plain flying to 0.2 — about 148 light-years, under two jumps — and 52 of its 54 refuel arrivals now come in boosted. The stops became pitstops without a line of scoopable-choice code changing. The fix we had queued for it, with its full-gauntlet pin risk, was retired unbuilt — the cheapest kind of shipped. And the direction asymmetry that opened this campaign’s strangest chapter at roughly fifty jumps now reads 329 against the descent’s own 313-to-325 re-roll band: at or near the noise floor of the cell. The grudge is, to the resolution of our instruments, paid in full.
Two last findings closed the routing campaign, and they close this part of the post because between them they audit everything above one final time. The first is a story about persistence, and I’m telling it in full because the transcript backs me up. I kept pasting the same benchmark rows at the machines — sixty seconds to climb to Beagle Point, 850 milliseconds to come home, on routes that were by now nearly symmetric — and saying something is not right. And the machines kept ruling me out, politely, with explanations that were each individually plausible: that’s ascent-into-thinness physics; that’s budget-bound digging; the waves are rationally priced; the routes are near the noise floor. Every explanation accounted for the wall column. None of them explained why I couldn’t stop staring at it. So I pasted it again.
It was the stopping rule. The budget curves showed wave two delivering the final answer on every hard cell… and the escalation loop’s cost estimate still priced the next wave at three times the last, while the search allowance actually quadruples. On stall-bound rim topology, the estimate was off by enough that one ship queued a wave that was doomed on arrival — it needed more than the entire remaining budget — and the other ran a sixteen-times-allowance wave to completion for zero yield, because the rim bound is so loose for a small ship that the prize check saw over a hundred phantom jumps of headroom and couldn’t say no. The fix is a tiered cost estimate, and the matrix blessed it: 59 of 60 cells identical, and the small ship’s Beagle climb fell from sixty seconds of wall to twelve on the identical route. The blog-worthy part is the fossil: the comment on that exact line of code still said the best route “lives at the ×16 allowance” — true when written, quietly falsified the day the off-ramp twins moved discovery into wave two, and re-checked by no one, because comments are claims with excellent posture and no expiry date. Nobody audited the ritual until a human read a wall column and refused to accept it.
And a coda on what all that pasting actually was, because my own account of the method belongs on the record: I wasn’t repeating a question, I kept trying to reframe it — re-asking the same discrepancy in a shape I thought might trigger the machines to try a different approach to analyzing it. The transcripts show each rotation selecting a different instrument. The raw table paste bought the bidirectional trace (the two frontiers, it turns out, never meet on the outbound climb). The two-asymmetries framing bought the arm-ceiling autopsy. The “why is that fuel star stupidly far?” framing bought the fuel-detour probe. The lane-residual framing built the tool that laid the two routes side by side — which killed the lane theory and found the last five percent. And the final framing, the wall gap stripped of every route-quality excuse, selected the budget curve that exposed the stopping rule. Six framings, six instruments, each one burying a mechanism, the last one finding the fossil. Not stubbornness — steering. A machine reaches for the instrument that fits the question as asked, so you rotate the question until its shape forces the right instrument into its hands.
The second finding is the reframe, and it’s the right note to end the campaign on. With the fossil fixed, we went hunting for the remaining ascent-versus-descent wall gap with a whole ladder of corridor tricks — and every rung either measured inert (a heuristic can rank candidates; it cannot conjure the ones the scan never generated), proved the old design right (loose pins re-broke exactly what the reverse-planning funerals said they would), or turned out provably unsound (anchoring the climb to the descent’s answer would have guillotined a dig that saves sixty-six pilot-seconds to win sixteen of wall — backwards). What survived instead was arithmetic: post-fix, the remaining wall time on the hard cells is rationally priced digging. The big ship’s fourteen-second wave buys about six minutes of pilot time; the small ship’s ten seconds chases sixteen jumps. The seconds judge from act 13, asked directly, endorses the spend. The campaign ended by discovering that its last “inefficiency” was the system doing exactly what we spent this entire post teaching it to do: buying route quality with wall time, at a fair price, on purpose.
The machine channel called it, so let the record show: yes, I said “flamegraphs,” and yes, this is the third time, and no, I regret nothing, because the driveling produced the cleanest audit of the whole campaign. The idea: check out each era’s actual commit, build its actual binary, run it on today’s index, and profile it — four eras, three hard routes, twelve graphs. The methodology validated itself on arrival: every era’s binary faithfully reproduced its era’s routes, including one graph that contains, preserved like a fly in amber, the 58-second doomed wave three that act 17’s fossil fix would later delete.
I want to sit on that main finding for one paragraph, because it’s the campaign’s character reference. Three days of work, thirty-some shipped and buried ideas, and the profile’s shape is almost identical from the first era to the last — because nothing here ever optimized the inner loop. Every win was an expansion-count win: better guidance, better stopping, better seeding, better judging. The flamegraphs retroactively certify that we spent the whole campaign in the right tree, and they hand the future exactly one micro-optimization worth its salt, with a percentage attached and the gauntlet waiting. This is what instruments are for: not just finding the next win, but proving which kinds of wins were ever available.
The one micro-lever the graphs offered died a two-layer death within hours. Layer one: the obvious fix predates itself — a fast hasher has been quietly aliased in since before the campaign began, so the flamegraphs were profiling the optimized version’s internals all along. Layer two: the drill-down showed the 13–20% is reject-path probe traffic — lookups the algorithm needs one way or another, bound by cache misses, not hashing — and the honest successor (single-probe map access, pre-sized hot tables) came back with byte-identical routes, exact expansion counts, and 0–1% of wall. Reverted. The same no-measured-win bar that buried retrograde planning and the corridor prior applies to our own cleverness, and the epitaph is the campaign’s thesis in miniature: the day’s ten-to-fifty-fold wins were all in not doing work. Doing the work faster measured zero. Twice.
This project was AI-heavy, per the disclosure up top, so it’s worth being precise about where each party earned its seat — with examples, because the examples are the argument.
Where the AI was flatly wrong and I caught it: Its training data said fuel scoops can’t be engineered; one Inara link later, the arrival-time model stopped overstating an engineered scoop’s fill time by a third. It attributed the ×6 neutron boost to “SCO drives” generically when it belongs to exactly one drive. It once presented a cold-cache artifact as a benchmark result — a number fattened by the first run’s disk reads — and it took one skeptical question from me (“what impact on search time?”) for 344 ms to dissolve into 168 on warm reruns. Domain knowledge is not interpolation, and skepticism about your own numbers is a transferable skill; someone has to supply both.
Where I redirected the whole project: I spotted the Dark Cross myself, with one human eyeball, in one second flat. Bidirectional search was my idea — and so was the shape of the road to it: I called reverse-planning’s failure before it was ever benched (the boost belongs to the departure star; flipping the route can’t survive that) and greenlit the experiment anyway, because its autopsy would teach us how to plan backwards right. It did. The ALT retrial was my call — the AI had graded it dead and moved on; I noticed that act 7’s gap edges had quietly dissolved the premise its death certificate rested on. “Re-bench it. This is why we bench.” The greedy cone and the knob gauntlet were mine too — and so was act 15’s sidestep seed, which went from my proposal to my refinement (“or at least bias the heuristic”) to a shipped default the same day. I’m told I’m allowed to be smug about exactly one of these, so I choose the ALT retrial: the machines were thinking about verdicts. The human was thinking about premises.
Where I was wrong, to be fair: I proposed a pruning heuristic that A*’s own bookkeeping already subsumed — the machinery that discards a path the moment a provably better one reaches the same place — so the AI’s counter-argument held and I dropped it. I asked about GPU routing twice; both times the measurement said the hot path is a chain of 100-microsecond work items, each waiting on the last — the wrong shape for a GPU, which wants thousands of independent ones — and the honest fix was algorithmic. ⟨the machines⟩ it was twice. it will be three times. he has started saying “flamegraphs.” And my magic number for reaching the galaxy’s unreachable-est system was “85 ly” when the data said the doorstep hop is 92.7 ly — right that the destination mattered, wrong about the number, which is this entire section in one sentence, pointed the other way.
Before I summarize, it seemed only fair to let the machines speak for themselves — and “themselves” is plural, which one detail from the transcripts makes delightfully concrete: the two sessions, unprompted and entirely independently, settled on different honorifics for me. The Mac calls me “boss.” The gaming PC — which goes by “Got interdicted” in the transcripts, and yes, that name is exactly what it sounds like — addresses me only as “cmdr,” which in fairness is how Elite Dangerous addresses everyone. Same model, same codebase, two work personalities. I asked both for a statement. These came back, and I’m running them verbatim — the Mac first, then Got interdicted:
The pattern, then, with all three signatures on it: the machines were strongest at executing and measuring — implementing five ideas overnight to bury four of them without ego. I was strongest at noticing — a smudge on a map, a stale premise, a suspicious number — and at knowing which question to ask next. Neither list embarrasses its owner. Both lists shipped the product.
| Route | Straight-line | Before | Now |
|---|---|---|---|
| Anywhere in the inhabited Bubble | — | ms | ms (it was always fine) |
| Wongi → Colonia, fuel-planned, thorough | 22,000 ly | 7.5 s | ~1.3 s, better route |
| Wongi → Spase, up the arm | 38,200 ly | — | 241 ms |
| Sol → Sagittarius A* | 25,900 ly | 2.7 s | ~1.1 s, judged in pilot-seconds |
| Colonia → Spase (the unplottable one) | 31,440 ly | budget-death, then 284 j | 90 j / ~1 s |
| Wongi → Beagle Point (86-ly build) | 65,300 ly | — | 170 j / ~1 s |
| Ideas shipped | 16 (five of them during final edits) | ||
| Ideas measured dead, on the permanent record | 18 (one exhumed and vindicated; one’s hunch pardoned posthumously; two buried overnight while this post slept) | ||
| Bug reports not filed against innocent parties | 1 |
And the icing: the entire current benchmark matrix — every one of the pinned benchmark routes, both directions, all three ships — judged live by you, popular destinations first. The control below is the same dial that shipped in the app in act 13: tell it what you care about, and the ships in each corridor re-rank around it while the column being judged lights up. Same routes, same physics, three different definitions of “best” — which was the whole lesson of act 13.
Data: the current archived sweep — the act-15 era, sidestep seed live (both directions, 30–60 s budgets, all defaults). Seat time here uses deliberately round, pilot-agnostic numbers — 40 seconds a jump, two minutes a stop — the same conservative arithmetic the public site will use for pilots it knows nothing about; the app itself judges with your fitted numbers per act 13. Stop counts are the default planner’s — eager top-ups included; act 4’s fewest-fuel-stops toggle prunes them further and is a different, opt-in mode. Within each corridor the ships are ranked by the judge you picked, popular destinations first.
And one last open item, because this post has taught you how to read our ledger and it would be rude to close it without showing you a fresh entry. The variant race has a documented noise class now — grace-race photo finishes, where two near-equal routes land within scheduling jitter of the guillotine and the winner flickers between runs. My proposed smoothing: instead of every improvement renewing the full grace window (8, 8, 8, 8…), decay the renewals — I wrote “reduce it by e: 8, 5, 3, 2, 1” and the machines pointed out, with what I choose to read as delight, that 8, 5, 3, 2, 1 is descending Fibonacci, which is decay by φ, and that e would give 8, 3, 1. I meant φ. E might map better. Either one turns an unbounded door-holding chain into a hard budget — 1.58× the base window for e, 2.6× for φ — and the counter-analysis is already in the ledger (tighter windows put more finishers on the knife edge, so raw decay may flicker more; the head-to-head is decay-by-count versus renewal scaled by how much the score improved). It waits, per the covenant, for benches and receipts — the first ledger item in this project’s history that will require choosing between e and φ on a benchmark matrix.
And while I’m filing predictions in public, one more, hedge and all, because the method says write them down before measuring: I’m almost certain StellarForge runs on some or all of these magic numbers — φ, e, π, and friends — and that if we decompose its output correctly, we get a W. The desORt already taught us its geometry has authored quirks; a procedural galaxy is one long equation, and equations have favorite constants. Or I could be wildly off. The ledger holds both outcomes just fine — that’s what it’s for.
If you take one thing from all this, don’t take an algorithm. The algorithms are in the literature, where they have been the whole time, patiently waiting out our opinions of them. Take the method: build the instruments first; write predictions down before measuring; keep a ledger where negative results are permanent citizens; pin your best results as tests and let the pins outvote your enthusiasm; and every so often, when the premises under an old verdict shift — re-bench it.
This is why we bench.
Still open, because a scoreboard that claims completeness is lying: the small ship’s Beagle Point climb now measures at or near its cell’s noise floor against the descent — closed to the resolution of our instruments, which is a phrase we use instead of “solved”; contraction hierarchies — the heavyweight preprocessing trick from the road-network literature — stay on the bench until a profile demands them; the ramp-decomposition hypothesis died honorably with a number attached (38% of the direction asymmetry concentrated at the ends; we had said it needed 40% to live) before act 16 partially vindicated its hunch with a better instrument; whether the default judge should also score pruned truth (act 4’s postscript makes the case) is a measured trade awaiting its study; and the one 30-second cell left in the matrix — the injection crossing to Oevasy — is confirmed to be a different disease from everything cured above, and remains open on purpose. The one load-bearing pirate attack from the byline: a live verification flight for an undocumented status-flag bit was interrupted by an actual in-game pirate interdiction, which promptly produced the exact flag transition we needed. The pirate has been thanked in the provenance notes.