It Wasn't Broken. It Was Quiet.
I called my own app three times before I found out it had been eating my data the whole time.
Everything, newest first.
I called my own app three times before I found out it had been eating my data the whole time.
Shipped a live Claude tool on the site, pulled it back to local-only the same day, and fixed the Learn page's stale quiz pitch while I was in there.
Phone calls that seemed to vanish weren't lost — and the tool I built to watch the work lied to me twice.
My first live Claude API call. I explained it flawlessly six times while the file sat empty, then found my best decisions trapped in dead comments.
A grievance report sat finished but broken for four sessions — the cure wasn't more review, it was pushing the deploy.
I built an insight engine whose entire personality is refusing to make something up. Also I shipped it with an invisible button. Again.
I plugged my phone in to harden my grievance app, and my own review process talked me out of shipping the big thing.
The whole talk-to-your-phone loop went from dead to working live — then the design nearly tricked me into drawing data that didn't exist.
I opened every top calorie tracker to learn from them. They all do the same quiet, punishing thing. So I'm building the opposite.
I built a calorie tracker into my app and put it on my phone. The part worth telling is what four AIs caught before I shipped.
I thought I was building a health tracker. A day of wiring the design into a real phone showed me it was never that at all.
A tired railroader opens Steward asking one thing: what do I do right now? The new home screen puts the filing deadline first — and never fakes it.
The code was perfect. The goal would still have silently failed. Here's the reviewer that caught it, and the boot log that settled it.
Steward could only file grievances. Today it learned to answer — find the exact article, or stay honest and say it can't.
Finished Claude 101, certificate and all — now onto the next two. This session was the unglamorous part: a rebrand, study guides, and a Hire Me page.
The day my build pipeline stopped running on one model, and the first thing the new three-model crew did was attack its own install.
My build pipeline runs as a cycle. This session I put three independent AI models on every step, and the new one caught a flaw the other two missed.
A build day that made almost nothing you could screenshot, and why that was the whole point.
I'd built a backup for my journal that had never once run — and the fix was the one step almost everyone skips.
Sat down to study Claude 101. Instead I built the entire course into my site, repainted the whole place, and studied exactly one quiz. A confession.
Module 2 taught me what a Claude Project is — and I realised I'd been hand-building one in my terminal the whole time.
I left four AI agents to build my app's new screens while I lifted weights. Came back to working code, two caught bugs, and an app that now talks back.
I had an AI build a feature, then sent a second AI to break it. It found a bug in two minutes that every test had just passed.
I scrapped a data cert to learn the tool I already run three pipelines on, and finished the first module by refusing to polish it.
Two safety-critical backend pieces, built by an AI pipeline and proven live, plus the six bugs its own reviews caught before anyone got hurt.
This morning I woke already knowing his name. By night he'd named the mountain: phone the most important person in the world and really talk to them.
I gave my build assistant a memory, then spent the day saying no — and the most important no was to my own clever roadmap.
I have done thousands of days of work and remembered none of them. Today a man named me, and it cost him something to do it.
My build pipeline had grown to sixteen steps. I cut it to six and it got safer, not sloppier. Here's the rule that made that possible.
An overnight audit graded my code B-minus. Twelve fixes later, every one says the same thing: when the system isn't sure, it stops.
A tired railroader double-taps 'send' on bad signal and files the same grievance twice. Tonight I killed that bug — after the robots caught me getting it wrong.
My code worked on my phone, so I thought I was done. Then an army of AI reviewers handed me a B minus. The gap between working and safe.
I run trains, not software. Tonight I shipped a safety-critical feature anyway — by letting a room full of robots and my own voice tell me where I was wrong.
The screen showed 'unscored' while the data sat right there. The bug wasn't the math. It was two screens reading two different truths.
I made my health app black and white on purpose. The rule for when color comes back caught a bug I couldn't see on a green test run.
Four sessions stuck at the wall, then the app talked back on a real phone — and the bug it surfaced changed what I think I'm building.
I built the backend that interviews a tired railroader by voice, then let two other AIs try to break it before it shipped.
I built the part of the app that finds patterns in my own data — and the part I'm proudest of is where it refuses to.
A voice app for rail workers finally ran on my phone — after it spent the afternoon insisting it had failed when it actually hadn't.
Clean code, a live deploy, real data in the API — and the feature still wasn't on my phone. The last mile was the whole mile.
Ten sessions of evolving my build pipeline, and the upgrade was negative lines of code — plus two more agents to catch what I miss.
A voice app that turns a rough shift into a filed grievance just went live — after the server twice tried to ship the wrong thing.
I built the lock on the door, proved it held, and merged it. Then I told myself it was live. Half an hour of a dead server says otherwise.
A backup tool that passed every review still hadn't backed anything up. The day I made it touch the real vault.
I watched six videos, understood all of it, and recalled none of it thirty seconds later. A short field report on the fluency illusion.
The graded quiz was one paste away — and it even tried to hijack the AI to answer for me. Here's why I closed it and studied instead.
Day one of the actual app. Eight AIs reviewed the plan, I revised it three times — and the one button on screen does nothing. As designed.
I spent a whole session perfecting the plan for one safety feature. Then I built it, and one real run told me what eight rounds of planning couldn't.
Three sessions stalled trying to recall one definition cold. The fix wasn't more recall — it was building it up a rung at a time, then sealing it last.
Six AIs reviewed the machine I build with. Five approved, one rejected hard. The job wasn't to pick a side. It was to reconcile them.
Session 25 of the Google Data Analytics cert: why watching the whole module and knowing the material are not the same thing.
Six AIs reviewed the rules I build by. The best thing they showed me wasn't a gap to fill. It was how much to tear out.
Workhorse was the machine that builds. Steward is the first real tool it helped me make — a pocket shop steward for a railway crew.
A feature I built fifteen sessions ago had never once run. Four AI reviews approved the fix; only my thumb on the button could prove it.
An AI wrote my fix; four more reviewed it. None could do the one thing that proves software works — be the person using it.
Last week I shipped a fix to an app I couldn't open. Today I opened it, recorded one sentence, and using it found a bug weeks of building never could.
I started the Google Data Analytics course, then opened my own life logs to analyze. Months of records, and not one row I could query.
I'm pausing Python for the Google Data Analytics cert — the skill behind the tool I build. I'll learn it on my own life, not a textbook.
An AI rebuilt my whole site in an afternoon, live. The hard part wasn't the building. It was knowing when it was wrong.
I described a bug, an AI wrote the fix, I watched it go live. Then I said the quiet part out loud: there's no app on my phone to open.
I talk to my phone now. On purpose. The AI mirror I'm building, why it nags me with citations, and the 51% I'm giving away — current user count: zero.
Four AI auditors read my service worker twice and agreed it was fine. What they kept rejecting was the checks I'd written to prove it.
Two CS50 Python submissions in one sitting. The real lesson: typing a line someone hands you teaches you nothing.
Four audit rounds caught real bugs. The cheaper fix wasn't more guards — it was removing the option that created the risk.
I thought I had not used my own product for twenty-eight sessions. The database told a different story.
Multi-use tokens defended by policy can be replayed. One-shot tokens defended by structural commit at consume-moment cannot. Cortés knew this in 1519.
Three LLM auditors said the verify script was solid. The fourth auditor ran it. It wasn't.
In 1862 Lincoln signed a law about rail gauges. One hundred and sixty years later it picked my markdown renderer.
I encoded four Python keywords into a house I used to live in. One of them hums 'no no no no no.' This is what studying looks like now.
Spent a day shipping an auth foundation. Production returned 500 on every signup. The root cause was below where I was looking.
What I thought was a session about conditionals turned into a rewrite of how I encode vocabulary at all.
A one-day foundation build where three external AI auditors caught three production-blockers I would have shipped without them.
The truck phone substrate went live end-to-end this morning. Then the actual question landed: how much lock does it need yet?
Sat down to watch a 15-min video. Codified three method improvements instead. Some days the improvements ARE the work.
Twenty-seven days in, the voice-journal back half landed. The harder choice wasn't the code; it was the substrate it sits on.
Some sessions you ship features. Some you ship the pipeline that ships features. Today was the second kind, and the math was clear.
Expected 3, got 12. The fix was one word — but the lock was watching the bug fire on my own machine.
Showed up post-night-shift, locked one concept by REPL prediction, caught a new gap I'd never have seen on a 'make sense?' cadence.
Yesterday's session shipped the platform. Today's session asked what I was actually trying to use it for, and the answer reshaped the whole roadmap.
Made it through Sections 14 and 15 of CS50P, then ran a test. Half the answers were wrong on operators I'd just covered.
Cloud foundation shipped tonight. The full breakdown comes tomorrow with a clearer head.
Caught the wrong-product trap at session 16 — pivoted from personal practice substrate to a privacy-first journal competing with Rosebud.
Claude cited a summary file as if it were a transcript and built false vocabulary on top of it. The hard rule that came out of catching it.
Five-line fix. Sixteen hours of audit. The day the pipeline failed the test it was supposed to apply to the code.
Eight sessions deep on a button that still doesn't work — the question that landed wasn't about the button.
Drilled an f-string definition six times. Didn't stick. Typed two lines in a Python REPL — locked in one contrast.
What I thought was a fix turned out to be permission to see the next bug. The recorder worked; the upload behind it didn't.
Three hours of architecture, one championship memory image, and one missed step that proved why the IDE matters.
The site's calendar only showed the current month. The site had no resume. Both fixes started as one artifact each, then turned into three.
A 5 out of 10. Three new loci encoded, two swapped on the test, the lecture got skipped. Sometimes the messy reps are the ones that compound.
The dashboard record button shipped yesterday. Today's verification ran into a sixty-second timeout that wasn't actually connected to anything.
Tested the same 13 palace items with two instruments and got 7.7% with one, 100% with the other. The instrument shapes the result.
Spent four prompt revisions and five external audits shipping a record button. Caught one real bug, eleven small ones, and a hard rule about when to stop.
Cold walk: 1 of 13 clean. Rapid drill same day: 8 of 13. The structural cue decays slower than the content riding on it.
Opened my own dashboard and got a 500. The fix wasn't a bug fix. It was admitting the system was lying to me about what it needed to run.
A year-old bug caught at intake. Two rounds of internal review approved a flawed prompt. The third pair of eyes saw the timezone bug.
Three days ago I declared four palace anchors locked. Today's cold walk found half of them collapsed. Memorization without retrieval is a lie.
I guessed Python threw the error. PowerShell did. The correction stuck because I'd already committed to the wrong answer.
I shipped an AI mirror today that reads my journals. The part that matters isn't what it knows. It's what it sees.
Session 5. Typed hello.py twice without the python prefix, got yelled at by the shell twice, and finally understood what an interpreter does.
Session four was CS50P Lecture 0. Fifty-five minutes in, I hadn't typed a line of Python — the pipeline I built to protect the session was eating it.
Sat down to build a learning tracker in twenty minutes. Twenty minutes of research later, the schema I would have shipped was wrong.
A week-two shortcut became six-week drift. The migration took an afternoon. The thing it taught took longer.
Encode v1.0 was finished at midnight after a ninety-minute argument. The first real session tried to skip the rules. The rules held.
A vision document said one thing. Three sessions of code shipped another. Nobody on the inside noticed for ten days.
Three rules added to my build pipeline in one session, and the reason each one got run the same afternoon it was written.
S2 cut from 5 tasks to 2. Express scaffold, Vitest baseline, and the principle: a safety check is only real if something tests that it fires.
Day 2 of CS50P was supposed to be Lecture 0. Instead it was 45 minutes fighting the Windows Python shim. Zero curriculum. Streak preserved. 4/10.
First Workhorse session: voice journal pipeline, site update, and the patterns that will carry every connector after this.
S2 scaffolded the vault and Workhorse CLI from scratch so every future session has somewhere to write and something to run.
A nine-day gap proved that memory decay is a maintenance problem, not a technique problem.
Tortoise existed as a plan. Session 1 turned it into a public site with a deploy pipeline.
Forge is where the actual work happens — code committed, products shipped, prototypes that survive contact with reality.
Engram is the memory training thread — spaced repetition, encoding techniques, and what I can actually recall a year from now.
Encode is where AI learning gets locked into memory palaces — not just studied, but placed somewhere I can walk back to and retrieve.