Skip to content
← Back to the build log

Part 9 · 16 September 2026

Fourteen hours behind a green health check

A database migration accidentally replaced the deployed app's password with a local one. Sign-in failed for fourteen hours while the health check stayed green, exposing a serious gap in my deployment checks.

Postgres Authorisation Playwright AI agents Incidents

series · Rebuilding a school-ops dashboard

Version one is waiting on the calendar. The last exit criterion is fourteen consecutive nights of backups in the bucket, and the fourteenth night is the twenty-sixth. I did not want to sit for two weeks, so I asked the agent for an inventory of everything the four milestones had deferred and picked the smallest piece that changes what the app is.

Every milestone since the second had put off the same two things. When a teacher files an expense, the row lands with approval_status = pending. The domain refuses to bill it, warns about it, and refuses to finalise the month while any exist. Nothing could decide it. The only way to close a month with a pending expense in it was to delete the row. Four columns for a manager’s adjustments had existed in the schema since the second milestone, the domain read all four, and nothing wrote them. Until this week the manager could read money. Now she can change it.

Eight tasks were built on Sunday and deployed on Monday evening. Reviews changed three of the planned tests, and the browser suite found a bug the smaller checks had missed. The deployment is the important part of this entry. Sign-in on the box failed for fourteen hours while the health check remained green. The cause was recorded in a line I had filtered out of the command output because it contained the word PASSWORD.

Writes with rules

The design decision was to put nothing new on screen. A pending block and a material-fee dialog on the draft page the manager already has. An adjust panel on the view-as month page she already opens. Four row-level routes: decide one pending row, patch one row’s rate override or billing month or one-off amount, put a person’s material fee for a month, read a person’s rows for the panel. Manager and head can do all of it, which widens the original portal’s manager-only rule, and I made that call on purpose because the head is the person who covers for the manager.

Every write has the same shape.

request
  |
  v
AuthGuard ........ who is this, really
  |
RolesGuard ....... manager or head, or 403
  |
month open? ...... invoiced_through and no final invoice, or MONTH_LOCKED
  |
UPDATE ... WHERE approval_status = 'pending'
  |
  +-- 0 rows? ...... somebody else got there first: re-read, report theirs
  |
audit row ........ actor, target, what changed

The domain package did not change. It already billed approved rows, ignored rejected ones, and honoured the three row columns and the fee table. The migration is one GRANT, because the third milestone’s column grants had never given the manager role write access to the fee table, and nothing had needed it. The oracle baseline from the first milestone stayed untouched, which was the constraint I cared about most.

Three tests the plan got wrong

Each task was built by one agent and reviewed by another, and the reviewers argued with my plan three times. They were right each time.

The decision route has a race guard. Two managers open the same pending expense, both click approve, the second update matches zero rows and the code re-reads the row to report what the first manager did. My plan’s test for that case made the row not pending before the call, so it passed through the early status check and never reached the fallback it was named after. The reviewer flagged that the concurrency contract in the spec had no test exercising it. The fix was a hook on the fake repository that lets the update return nothing while the read still says pending.

The draft’s pending list should show only rows the domain warned about in that month. A row moved into another month by its billing-month override belongs to that other month. My plan’s test had one pending row, so the half of the filter that handles moved rows ran against nothing. Two more rows in the same case, one moved out and two sharing a date so the sort tie-break runs too.

The third was mine in a way I found embarrassing. The proof that a rejected expense adds no money asserted that the teacher’s October work total was zero after the rejection. The implementer stopped with fifty-eight of fifty-nine passing and said the seed pays that teacher a twenty-thousand-yen monthly retainer, which bills whether or not she teaches. I had written an assertion from the fixture in my head instead of from what rejection means. The test now reads the total before the decision, rejects, reads it again, and asserts the two are equal and that no line in the month names the project. That is a stronger proof than the one I asked for, and it was not what I asked for.

The manager who could not read her own panel

The last task runs the real browser suite, and its fifth journey is the manager opening a teacher’s month via view-as and adjusting a lesson’s rate. The panel’s own read, the person’s rows for the month, answered 403 with the code FORBIDDEN_ROLE, under both time zones the suite runs in.

View-as was designed in the third milestone as “the app as that person”. The web client attaches a header naming the viewed person to every GET while a person is viewed, and the API swaps the request’s identity for that person’s before any guard runs. So the manager’s read of an admin-only route arrived at the role guard wearing the teacher’s role, and the guard refused the teacher. The route and the guard were both correct. The header was on the wrong request.

There were two fixes and one was wrong. The guard could learn about the real role behind the substitution and accept it. That would let a manager pass every manager-only guard while impersonating a teacher, which turns view-as into a costume rather than a viewpoint. The right fix was one line in the client: the header rides only on the /api/me routes, because those are the routes that mean “as that person”. Every admin read and every write reaches the API as the real user.

Nothing before the browser suite could see this. The panel’s unit test mocks the fetch layer, so the middleware never ran. The API’s end-to-end test sends the view-as header with writes, where the refusal is the tested and correct outcome, and never with that read. Two green suites, each honest about a world without the other in it.

Two sessions in one tree

Halfway through the last task the implementer stopped: a rate limit ended its session with nothing committed. That part was fine. The ledger and the files on disk let a fresh agent resume without redoing anything. The fresh agent then reported something else. Files it had not touched were changing under it. A second session of mine, in another window, was adding Sentry tracing on the same branch in the same working tree, and it had edited eighteen files, including all four state files the milestone’s own agent was trying to update.

I told the milestone’s session that the tree was its own. It moved the other session’s files into two named stashes and committed its own by explicit path. The whole-branch review still found a paragraph from the tracing work inside the operations document, describing a feature the branch did not contain, and cut it. One session per branch, or a worktree each. I knew that and did it anyway because both jobs were small.

The tracing work is now merged, three days later, on its own branch. One trace id runs from the browser’s navigation to the month page through the API’s guards and its ten queries, and the health check is never traced, which matters for the next section.

Fourteen hours behind a green health check

Monday evening I deployed. The milestone carries a migration, so the sequence was the one from the third milestone: open a port forward to the box’s database over ssh, run the migration tool from my laptop against it, then pull the images. The migration applied, the grant showed in the catalogue, the four new routes appeared in the API log, the health check reported the new release. I brought the tunnel up for the hand checks and left them for the morning.

At ten the next morning the sign-in button did nothing.

The API log said password authentication failed for user "minim_app", from 19:34 the previous evening, on the first database write of the sign-in flow. The migration tool loads the repository’s .env with dotenv before it connects. On my laptop that file holds my laptop’s app password. The tool has an optional step that sets the app role’s password from that variable, so I could rotate it on a fresh box. I had not passed the variable. I did not need to. It was already in the environment, from the wrong machine, and the tool ran ALTER ROLE minim_app PASSWORD on the production database with my laptop’s value.

The tool printed a line saying so. I had piped its output through grep -v PASSWORD to keep the record clean of secrets, and that line was the one it removed.

The health check answered {"status":"ok","release":"sha-7a50d78"} all night because it touches no database. The deploy script’s proof, the one I was proud of in the last entry because it refuses to declare success until the release matches, was true and useless. The box was serving the release I asked for, and nobody could sign in to it.

The fix was an ALTER ROLE back to the box’s own value, proven by connecting as the app user from a second container, then a restart. The rule now sits in the runbook and the operations document: a migration run from a laptop passes APP_DB_PASSWORD= empty, in front, every time, and the tool prints left unchanged when it gets it. And nothing filters the migration’s output. If a line contains a secret, elide it in the record by hand afterwards.

There was a second, smaller find the same evening. To approve an expense by hand I needed a pending one, and the teacher test account could not file into September, because September was already finalised. I had finalised it myself on the eighth, during the previous milestone’s hand check, as a test. A finalised month is the thing the app is built to protect, so there is no route to reopen one. I did it in one transaction as the database owner, with an audit row saying so, and left October alone. The month is open again and the hand checks passed: two expenses approved, a one-yen rate override set and cleared, a fee set and cleared, every one of them audited with me as the actor.

What I am keeping

A health check that touches nothing proves nothing about the thing you care about. A green light that cannot go red in the failure you had tells you nothing. A health check that reads the database is now on the backlog.

Never filter the output of a tool that can change a secret. Read every line. Elide by hand afterwards.

A value from one machine is not a value on another. This is the fifth time in this project a tool has carried my laptop’s world onto the box. The carriage returns, the key mode, the clipboard, and now dotenv. Each time the fix was the same: prove the value on the consumer’s machine, as the consumer.

Write the assertion from what the word means, not from the fixture you remember. “Rejected adds no money” is a before-and-after comparison. “Work is zero” was a guess about the seed, and the seed disagreed.

View-as is a viewpoint. The header goes on the routes that mean “as that person” and nowhere else. Widening the guard would have made the test pass and the model wrong.

One session per branch. Two small jobs in one tree cost a stash, a leaked paragraph, and an afternoon of flaky browser runs.

The manager can approve and adjust now, and every change she makes leaves a row with her name on it. Ten nights of backups remain. Then it is version one, and the next piece is the one the teachers will notice: absences, cover, and who is actually in the room.