Nhat Linh standing on a path through a bamboo grove

Nhat Linh

Get to know me

I'm Nhat Linh, a junior DevOps/MLOps engineer who likes turning experiments into useful tools, from Kubernetes labs to AI workflows. Away from code, you'll find me behind a camera or at the piano.

I’m Nhat Linh, a DevOps/MLOps Junior. I love exploring and experimenting with new ideas. I believe that my curiosity in tech will bring me to new horizons.

French-Vietnamese dictionary

French-Vietnamese dictionary

  • React native
  • Tauri + Astro
Dictionary detail view for “être”
Dictionary entry for “prénom”
Dictionary search results

Nail salon SaaS

Booking, management, and client application ecosystem for nail salons

  • React native
  • Tanstack + Better Auth
Nail salon dashboard

Discord agentic orchestrator

Agentic workflow from Discord threads to Claude Code tasks

  • TypeScript
Discord bot reporting a validated Claude Code task

KubeLearning

Learn Kubernetes with a self-hosted terminal in one click

  • Astro + Go
KubeLearning landing page

Lecture Notes

Local-first desktop app that turns lecture recordings into structured Markdown notes with Whisper and Claude Code

  • Tauri 2 + React
  • Whisper + Claude Code
Project

Lecture Notes

Case Study: Lecture Notes

A macOS desktop app that turns a lecture recording into structured study notes. Transcription runs locally with whisper.cpp. An AI model of the user’s choice then turns the transcript into organized Markdown notes.

Role Solo: product, design, engineering
Timeline 15–18 September 2026 (a working MVP on day one, then two iteration rounds)
Platform macOS (Apple Silicon) desktop app
Stack Tauri 2 · Rust · React 19 · TypeScript · Tailwind CSS v4 · whisper.cpp · ffmpeg
AI providers Claude Code, Codex CLI, OpenCode, Pi (CLI-based) · OpenRouter (API)
Size ~8,300 lines of Rust + TypeScript · 31 commits · ~110 Rust tests

1. The problem

A recorded lecture is hard to study from. A one-hour recording is only useful when you scrub back and forth to find the part you need. A raw transcript isn’t much better: it’s a wall of text full of filler words, false starts and misheard technical terms.

What a student actually wants from a lecture is:

  • the structure: which topics were covered, in what order, and where in the recording each one starts
  • the definitions and key points
  • the parts the lecturer stressed, especially the ones that sound like they’ll be on the exam
  • a way to test themselves afterwards

The goal: drop in an audio file and get back a clean, well-organized Markdown note in a folder you control. It should need no account and no upload service, and nothing should leave the machine except the transcript text sent to the AI model the user picked.


2. What the app does

 lecture.m4a
     │
     ▼
 ┌──────────────┐   ┌────────────────────┐   ┌──────────────────────┐   ┌────────────────┐
 │  Preprocess  │──▶│     Transcribe     │──▶│       Generate       │──▶│     Write      │
 │  ffmpeg →    │   │  whisper.cpp       │   │  AI provider →       │   │  Markdown +    │
 │  16 kHz mono │   │  (local, Metal GPU)│   │  schema-checked JSON │   │  frontmatter   │
 └──────────────┘   └────────────────────┘   └──────────────────────┘   └────────────────┘
                                                         │
                                                         ▼
                                           classify into a lecture folder
  1. Preprocess: ffmpeg converts any audio format to 16 kHz mono WAV, and ffprobe measures the duration.
  2. Transcribe: a bundled whisper-cli sidecar (whisper.cpp built with Metal) transcribes the audio locally with timestamps and detects the spoken language.
  3. Generate: the selected AI provider receives the transcript and a carefully written prompt. It returns a LectureNote JSON object that must match a JSON Schema: title, summary, time-stamped sections with key points, definitions, examples, exam-relevant points and review questions.
  4. Write: the note is rendered to Markdown with YAML frontmatter and saved to the user’s output folder.
  5. Organize: a second, small AI call assigns the note to an existing lecture (course) folder or creates a new one.

Around this pipeline:

  • Home: drag-and-drop or pick a file. A stage stepper shows live progress, and jobs can be cancelled.
  • Result: the rendered note, the full transcript with search, the detected language, and “Open .md” / “Reveal in Finder”.
  • Lectures: the notes grouped into course folders.
  • Calendar: import .ics files (Google Calendar exports work) so your course schedule lives in the app.
  • Settings: AI provider and model, note-style toggles (keep examples, keep code, timestamps, exam points, review questions), Whisper model download with progress, language, output folder and binary paths.

3. Stack decisions

Tauri 2 + Rust instead of Electron

The core of the app is orchestrating external processes: ffmpeg, whisper-cli and AI CLIs. That means streaming their output, reporting progress, and cancelling them cleanly. Rust with tokio::process is well suited to this, and Tauri gives a small native bundle that uses the system WebView instead of shipping Chromium.

Tauri’s Channel IPC fit the job model well. Each long-running command streams typed progress events to the UI and is cancelled through a CancellationToken held in app state.

whisper.cpp as a sidecar binary instead of a cloud speech API

  • Privacy and cost. Lecture audio never leaves the machine, and there are no per-minute transcription fees.
  • Speed on Apple Silicon. whisper.cpp with Metal is fast enough for hour-long recordings on a laptop.
  • Sidecar instead of FFI bindings. Running whisper-cli as a separate process keeps the Rust build simple. It also isolates crashes and makes cancellation a simple process kill. A build script (scripts/build-whisper.sh) compiles a pinned whisper.cpp tag and checks that the binary only links system libraries.
  • The user chooses the model (base → large-v3-turbo) in Settings, and the app downloads it with streamed progress, so the app bundle stays small.

AI CLIs instead of an API SDK (at first)

The first provider was the Claude Code CLI in headless mode (claude -p --output-format json --json-schema …). This reuses the user’s existing subscription and authentication, so the app never stores an API key. --json-schema gives structured output that is checked against a schema.

This decision later made multi-provider support cheap. Codex, OpenCode and Pi are all “send a prompt on stdin, get text back”. The differences are argument construction and output parsing, so both were kept as pure, unit-testable functions.

React 19 + TypeScript + Tailwind v4, no router

The UI has four or five screens, so navigation is plain React state. ts-rs generates TypeScript types from the Rust structs. The frontend and backend share one source of truth, and cargo test regenerates the bindings.

Deliberately small dependency surface

The plan listed the allowed Rust crates up front (tokio, serde, reqwest, ts-rs, uuid, chrono, …). Adding anything else needed an explicit decision. There is no database yet: jobs are directories under the app-data folder (meta.json, transcript.json, note.json), and the files on disk are the source of truth.


4. How I built it

I built this with an AI-assisted, spec-driven workflow, with Claude Code as a pair programmer. The process mattered as much as the stack:

  1. Design spec first. I wrote down the product, the pipeline and the constraints before any code.
  2. Implementation plan. The first commit in the repo is a ~420-line plan. It contains the exact data model, the module layout, the global constraints (for example: “modules that talk to external processes take explicit binary paths so they can be unit-tested”), the verified CLI invocation, and a UI brief.
  3. Task-by-task execution with review rounds. Each task produced a feat: commit. A review pass followed, and its findings landed as separate, well-explained fix: commits. That is why the history alternates feat → fix.
  4. Verify against the real thing. Unit tests use fixtures. An #[ignore]d end-to-end test runs real ffmpeg, whisper-cli and the Claude CLI on a generated sample recording.

My role was to write the specs, set the constraints, make the product and architecture decisions, review every change, and test on real lectures.


5. Development timeline

Day 1 (15 Sep): from plan to working MVP, 26 commits

Step Commit What happened
0 docs: implementation plan The plan defines the data model, constraints and task breakdown
1 chore: scaffold Tauri 2 + React + Tailwind v4 project
2 feat: data model and markdown renderer Shared Rust types → generated TS; LectureNote → Markdown with snapshot tests
3 feat: ffmpeg preprocessing and whisper transcription Shared streaming process runner with cancellation; whisper JSON parsing (ms → s)
4 feat: Claude Code note provider NoteProvider trait, prompt building, JSON extraction with a fallback for fenced output
5 feat: pipeline, settings, model download and Tauri commands The full job lifecycle, wired to typed commands
6 build: whisper.cpp sidecar build script Reproducible sidecar build pinned to a tag
7 feat: app shell and Home page Sidebar, drag-and-drop, live stage stepper
8 feat: Result and Settings pages Note and transcript viewer; settings with debounced autosave
9 test: end-to-end pipeline verification Real-binary integration test
10 hardening pass Lifecycle, threading and isolation fixes (see §6)

Day 2 (16 Sep): provider-agnostic

  • feat: support multiple AI CLI providers: the Claude-specific code was split into a generic CliProvider. Each provider is an OutputMode (Claude JSON, Codex output file, OpenCode JSON lines, plain text) plus an argument builder. Users can set a custom binary path, because apps launched from the macOS GUI don’t inherit the shell’s PATH.
  • Small UX fixes: hide the YAML frontmatter in the preview, and show the detected language on the Result page.

Day 3 (18 Sep): direct API, organization and a roadmap

  • docs: design OpenRouter support: a written design spec (scope, what’s excluded and why, architecture) before implementation.
  • OpenRouter provider: a direct HTTP provider behind the same NoteProvider trait, so it reuses the pipeline’s cancellation, progress, retry and writing unchanged. Model search runs in Rust, not in the browser, so the API key never touches frontend fetch code. The UI shows input/output pricing per model.
  • Lecture library: the AI classifies each new note against the existing course names (case-insensitive reuse) and files it into a folder. A new Lectures page lists them.
  • Calendar: an .ics parser with line unfolding, text unescaping, all-day events and de-duplication by UID.
  • Roadmap: ten feature plans (docs/plans/00–09), each with dependencies and a suggested order. They cover audio-synced notes, a SQLite + FTS5 library, flashcards and spaced repetition, PDF/slide import, export (PDF/HTML/Anki), privacy controls, and accessibility and multilingual support.

6. Interesting problems solved

Each of these is documented in its commit message.

Pipe deadlock with large prompts. The first version wrote the whole prompt to the CLI’s stdin before it started reading stdout and stderr. An hour-long lecture’s prompt is larger than the OS pipe buffer, so once the child started writing output, both sides blocked forever. The fix spawns the output readers first and writes stdin concurrently. A regression test pipes more than 2 MB through cat under a timeout.

Keeping the AI’s output uncontaminated. The Claude CLI loads CLAUDE.md and user settings from its working directory. Running it from the project folder, or with the user’s global config, could quietly change the notes. Each job now runs in its own empty workspace directory with --no-session-persistence and --setting-sources project, and with tools disabled. Every run is a stateless, reproducible transformation.

Schema validator rejections. The CLI’s --json-schema validator couldn’t resolve the standard $schema meta-URL and rejected every request. Removing that key fixed it. Parsing now prefers structured_output and falls back to extracting JSON from a fenced block.

Race conditions in the job lifecycle. If you pressed “New” while a job was running, events from the orphaned job could corrupt the new one’s UI state. Each run now gets a run token, and the reducer drops events from stale runs. Job IDs are checked as UUIDs, and deletion confirms that the path stays inside the jobs directory, so it can’t be tricked into deleting elsewhere. Child processes are kill_on_drop, and all in-flight jobs and downloads are cancelled when the app quits.

Keeping the UI thread free. Dependency probes (which spawn up to 4 processes) and native file dialogs had been blocking the async runtime or the main thread. They moved to Tauri’s blocking pool / spawn_blocking.

Serde and ts-rs disagreeing on a name. LargeV3TurboQ5_0 serialized as largeV3TurboQ5_0 in serde but was exported to TypeScript as largeV3TurboQ50. An explicit rename plus a test that checks every enum variant’s wire string against the generated TS prevents this whole class of bug.

A Rust lifetime puzzle. #[async_trait] expanded a method that takes an on_log: &mut dyn FnMut(&str) callback incorrectly, tying the &str lifetime to the outer borrow. The trait uses hand-written Pin<Box<dyn Future + Send + 'a>> signatures instead, which stay object-safe so providers can be boxed.

Prompting for noisy ASR input. The prompt tells the model that the transcript is speech-recognition output. It should expect misheard jargon and [BLANK_AUDIO] markers and fix them quietly, split sections on topic shifts rather than by time, and write every field in the lecture’s own language, never translating.


7. Design

The UI brief was: “a calm, focused study tool for one student; warm, quiet, typographic; not a SaaS dashboard.”

  • A warm paper-and-ink palette with one ink-blue accent, and light and dark tokens that follow the system setting.
  • System fonts only (the app works offline, so no webfonts). Headings use the macOS serif (New York) to give a notebook feel.
  • Overlay title bar, a quiet sidebar with recent jobs, and errors shown as a small dismissible line instead of modal alerts.

8. Testing and quality

  • About 110 Rust unit tests: parsers (whisper JSON, each CLI’s output format, OpenRouter responses, .ics), Markdown rendering snapshots, prompt building, backward compatibility of settings (older settings.json files still load), lecture-folder reuse, and path-safety checks.
  • Fixtures for real CLI outputs (structured, fenced and error responses).
  • An end-to-end test (cargo test --test e2e_pipeline -- --ignored) against real binaries and a real model.
  • pnpm build runs strict TypeScript over bindings generated from Rust.

9. Outcome

In three days the project went from an empty repo to:

  • a working, cancellable, local-first transcription → notes pipeline
  • five interchangeable AI backends behind one trait
  • automatic organization into course folders, plus calendar import
  • a written, prioritized roadmap for the next phase

10. What I learned

  1. Planning pays off most when the implementer is fast. A precise plan with global constraints and exact interfaces meant the AI-assisted build stayed consistent across 26 commits in a single day.
  2. Review rounds catch the real bugs. The deadlock, the lifecycle races and the serde mismatch didn’t show up in the happy path. They came out of deliberate review passes and real end-to-end runs.
  3. Design for replaceability early. Keeping argument building and output parsing as pure functions behind a trait turned “support four more providers” into an afternoon’s work.
  4. Shelling out is a legitimate architecture. Treating whisper.cpp and the AI CLIs as sidecar processes kept the app small, private and easy to swap. The cost is careful process handling.

11. What’s next

From docs/plans/, in priority order:

  • Disclose privacy behaviour and fix a contrast issue.
  • Audio-synced notes: click a section to jump to that point in the recording.
  • A SQLite + FTS5 library with search.
  • Editable notes with AI actions.
  • Flashcards, quizzes and spaced repetition.
  • PDF/slide import.
  • Export to PDF, HTML and Anki.
Project

Nail salon SaaS

NailFriends — a two-app booking system for nail salons

Case study · May – September 2026 · solo build

A multi-tenant salon back-office on the web, a customer booking app on iOS, and one availability engine shared between them through a versioned public API.

Role Sole designer, architect and engineer
Duration 2026-05-22 → 2026-09-14 (~4 months, part-time)
Surfaces Admin web app (nail-shop-admin), customer mobile app (NailFriends)
Stack TanStack Start · React 19 · Prisma 7 · PostgreSQL · better-auth · Expo SDK 57 · React Native 0.86
Scale of code ~25.7k lines TypeScript (admin) + ~6.1k (mobile), 35 test files, 508 translated strings × 3 locales
Status Admin running in Docker against Postgres; iOS app at v1.3.0 submitted to TestFlight; AR try-on engine is an open decision

1. The problem

Small nail salons run on paper diaries, Instagram DMs and phone calls. Two things break repeatedly:

  1. Double bookings. Availability isn’t one rule, it’s seven overlapping ones — is the salon open, is that technician working, is she off, is the slot blocked, does another appointment overlap, can she actually do a chrome set, and does the requested service duration even fit before closing.
  2. Bookings arrive in a channel nobody owns. A DM is not a record. There is no customer history, no no-show tracking, no reliable notice period for cancellations.

So the product is two-sided by necessity: staff need a back-office that is the source of truth, and customers need a self-serve app that can only ever offer slots the back-office would accept.

Constraint that shaped everything: the customer app must never be trusted. It picks the slot, but the server decides whether the slot exists.


2. What I built

Admin web app — the source of truth

Nine feature modules behind an owner/manager/receptionist/technician role model:

  • salon — settings, timezone, languages, deposit and cancellation policy
  • auth — staff sessions, organization membership, RBAC
  • services — catalog with fixed/from/range pricing and an onlineBookable flag
  • technicians — staff profiles and per-service skills
  • working-hours — salon opening hours, per-technician hours, days off, blocked time
  • customers — customer records, history, internal notes
  • appointments — lifecycle: pending → confirmed → checked in → in progress → completed, plus cancelled / no-show
  • availability — the slot engine
  • calendar — day/week grid with drag-to-reschedule

Plus customer-auth and a public /api/public/v1 surface added later for the mobile client.

Mobile app — NailFriends

Expo Router app with 32 screens: tabbed home, look gallery, a five-step booking flow (service → staff → date → note → confirm), booking management, a loyalty wallet, account screens, and an AR nail try-on flow.


3. Stack decisions, and why

TanStack Start over Next.js

I wanted end-to-end type safety from the database to the JSX without a separate API layer or codegen step. TanStack Start’s createServerFn gives a typed RPC boundary where I control validation explicitly, and TanStack Router gives fully typed route params and search params — including typed loader data. The tradeoff I accepted: a v1 framework with a smaller ecosystem and fewer answers on Stack Overflow. That cost showed up in real ways (Vite allowedHosts and auth trusted-origin configuration both needed dedicated commits), but the payoff was that a Prisma schema change surfaced as a type error in a component rather than as a runtime 500.

Modular monolith over microservices

A single deployable, but with enforced internal seams. Every feature lives in src/modules/<feature>/ and exposes exactly three things:

modules/<feature>/
  index.ts     # PUBLIC API — the only legal import surface
  schema.ts    # Zod schemas — the contract between server and UI
  server/      # createServerFn handlers + business logic
  components/  # React
  README.md    # responsibility, dependencies, tasks

The rule is one line long: import another module only via its index.ts. Never reach into another module’s server/, components/ or schema.ts. That single constraint is what made it possible to build modules in parallel (see §5) and what kept the refactors later in the project cheap.

Prisma 7 + PostgreSQL, with tenancy as a helper not a convention

Every tenant table carries salonId with an index, and every tenant query goes through withSalonScope(salonId, where) from shared/db/tenant. Making it a function rather than a code-review rule means tenant leakage is a visible omission in a diff, not an invisible one. The Prisma schema is sectioned by module (// MODULE: services) so parallel work on it produced predictable, non-overlapping diffs.

better-auth with the organization plugin

An organization is a salon. Membership gives tenancy and role in one object, so “which salon am I in” and “what may I do here” come from the same session rather than from a bolted-on salons table.

Compile-time i18n (Paraglide) over runtime i18n

The target market is French salons with Vietnamese-speaking staff and English tourists, so three locales were a requirement, not a nice-to-have. Paraglide compiles messages into tree-shakeable functions: a missing key is a TypeScript error, and unused strings don’t ship. 508 keys across en / fr / vi. The cost is a generated, gitignored src/paraglide/ directory that has to exist before tsc runs — solved with a pretypecheck/postinstall hook rather than by committing build output.

Expo + React Native for the client, not a mobile web app

The try-on feature needs the camera and, eventually, a native AR SDK. Expo Router gives file-based routing that matches the web app’s mental model, EAS handles iOS signing and TestFlight submission, and expo-secure-store puts the session token in the keychain instead of AsyncStorage.

Bun as package manager, Node as runtime

bun install and script running for speed; Node stays the runtime for Vite and TanStack Start, including in Docker (a node:22 image with the bun binary copied in). A mid-project commit — “Replace completely npm by bun in the system” — made this consistent rather than half-migrated, which is usually where this kind of choice goes wrong.


4. How it was built — the actual timeline

Reconstructed from git: 65 commits on the admin app, 17 on the mobile app.

Phase 0 — Mobile prototype first (May 22 – 23)

The mobile app is the older repo. Before any backend existed, I scaffolded NailFriends with tabs, a booking flow and a loyalty card, all on local mock data, then pushed it far enough to have a confirmation screen with calendar integration.

Building the consumer surface first was deliberate: it forced the data model to be designed against a real screen rather than against a guess. By the time I wrote schema.prisma, I already knew what a booking flow actually asks for.

Phase 1 — Design before data (June 2 – 3)

A written design brief (DESIGN_BRIEF.md) fixed a single aesthetic — “Coastal Spa”: Fraunces display serif, Manrope for UI, glass surfaces, soft teals and sand — and explicitly forbade business logic during the UI pass. First commit: “Run design brief and implement design for each module.” Mock data and useState only.

Then a 32KB architecture spec, then the foundation modules: auth, salon, shared tenancy helpers, Docker dev setup.

Doing the whole UI first, with no data, meant the design never got quietly compromised by plumbing decisions — and it meant the module seams were already visible before any of them had a server function.

Phase 2 — Core modules in parallel (June 3 – 4)

Services, technicians and customers have no dependencies on each other, only on salon. They were built concurrently against their schema.ts contracts. Commit “Implement salon, auth, service, technician, customer module” closes the wave.

Phase 3 — The calendar, spec-driven (June 4)

The largest single feature and the most disciplined stretch of the project. The sequence in git is exact:

docs: add option a calendar design spec        ← 9KB design
docs: add calendar real-data implementation plan ← 55KB plan
feat: share technician display helpers
feat: enrich appointment list items with technician names
feat: add calendar interaction math              ← pure functions
feat: support appointment form defaults
feat: wire calendar to real appointment data
fix: require drag threshold before calendar reschedule
test: guard calendar against mock data imports

Two details worth pointing at. Calendar interaction math was extracted into pure functions and committed before the UI was wired — pixel-to-time conversion and drag snapping are exactly the kind of logic that is miserable to debug inside a component and trivial to unit test outside one. And a test that fails if the calendar imports mock data: the migration from mock to real data is enforced by CI, not by memory.

The fix: require drag threshold before calendar reschedule commit is the giveaway that this was used, not just built — a click was being read as a one-pixel drag and silently rescheduling appointments.

Phase 4 — Localization and design polish (June 5 – 6)

Three commits localizing dashboard, calendar, customers, appointments, services, technicians and working hours. Retrofitting i18n is painful; it’s included here because it’s honest — doing it per-module after the fact was slower than writing the strings localized would have been.

Phase 5 — Opening the API to the mobile app (June 30 – July 6)

The most security-sensitive work in the project, and the most carefully sequenced. A design spec, then a 66KB implementation plan, then fifteen commits in dependency order:

feat(db):            onlineBookable, Customer.userId, User.principalType
feat(config):        getPublicSalonId resolver
feat(contracts):     client-safe public API Zod schemas
feat(services):      onlineBookable flag + public catalog handler
refactor(availability): extract computeAvailabilityForSalon(salonId, …)
refactor(appointments): extract bookAppointmentCore from createAppointment
feat(customer-auth): requireCustomer guard + signUpCustomer linking
feat(public-api):    json + error-envelope route helpers
feat(public-api):    wire services/technicians/availability/me/sign-up
fix(public-api):     stop internalNote leak, enforce onlineBookable/visibility
docs:                record customer-auth module + public API surface

Three things I’d call out as the good decisions here:

The two refactor commits come before the feature commits. Public booking had to run the same availability and booking logic as the admin app, not a parallel implementation that would drift. So the shared cores were extracted first, and the public endpoints became thin callers.

Clients cannot choose a salon. getPublicSalonId resolves the tenant from server-side config (PUBLIC_SALON_SLUG). There is no salonId parameter on any public endpoint to tamper with — a whole class of tenant-crossing bug removed by API shape rather than by validation. A later commit, “Remove salon id api,” finished the job.

Separate, client-safe response schemas. The admin Service type and the public one are different Zod schemas, which is what made fix(public-api): stop internalNote leak a one-line fix at a single seam instead of an audit of every handler. That commit is also a fair illustration of the risk: a shared model would have leaked staff-only notes to customers, and I caught it because there was one place to look.

The API is documented in a hand-maintained openapi.yaml (9 endpoints), which also served as the mobile app’s build contract.

Phase 6 — Mobile app against the real API (July 1 – 6)

feat: Implement booking context with API integration, then immediately feat: Update date formatting functions and add timezone handling. Then auth with guest browsing and post-login redirect, then v1.1.0.

Phase 7 — Hardening and shipping iOS (Aug – Sep 14)

  • feat: remake completely UI components for looks and booking management — a full visual second pass once the flows were proven
  • Expo SDK 57 upgrade, React Native 0.86
  • fix(ios): add missing Info.plist purpose strings; bump to 1.3.0 — the classic App Store rejection: camera, photo library and location all need human-readable purpose strings
  • chore(eas): set ascAppId for non-interactive iOS submission — CI-able eas submit
  • Admin side: real seeded working hours replacing the last fake data, remember-me on login, and an error-handling/user-feedback pass across components
  • Rework tab bar (Home, Gallery, AI Try-on, Booking, Profile) — the last commit, promoting try-on to a primary tab

5. Working with AI agents as a delivery method

This project was built with Claude Code, and the workflow is part of the case study because it’s why a solo build has module READMEs, 32–66KB specs and a documented module boundary rule.

The loop was spec → plan → implement → test, per module:

  1. A design spec in docs/specs/ — what and why, no code.
  2. An implementation plan in docs/plans/ — ordered, verifiable steps.
  3. Implementation against the schema.ts contract.
  4. Tests, then the next module.

CLAUDE.md at the repo root encodes the rules that make this work: module boundaries, contract-first Zod schemas, tenancy through withSalonScope, availability as pure functions. A docs/ tree and per-module READMEs (owner, dependencies, tasks-from-spec) keep the state of the project legible between sessions.

What actually made the difference: the module boundary rule and the schema.ts contracts. Parallel work across modules only stays coherent if each one has a narrow, written interface — and that’s equally true whether the parallel workers are agents or people. The architecture discipline wasn’t overhead on top of the AI workflow; it was the thing that made the AI workflow produce something maintainable.

I still own every architectural decision in §3, the security model in §5’s public API phase, and the calls documented in §7. The specs exist because I had to write down what I wanted precisely enough to hand off — which is not a bad habit to be forced into.


6. Engineering problems worth showing

The availability engine

Seven independent rules, one answer. The core is pure functions with no database access — it receives working hours, appointments and services as inputs, which makes every rule unit-testable in isolation. The test file is 23.7KB against a 15.5KB implementation.

It also does something more useful than returning a boolean: every unavailable slot carries a typed reason.

export const availabilityConflictReasonSchema = z.enum([
  "salon_closed",
  "technician_unavailable",
  "technician_day_off",
  "blocked_time",
  "appointment_overlap",
  "missing_skill",
  "duration_invalid",
]);

The schema enforces the invariant in both directions — an unavailable slot must have a reason, an available slot must not have one:

if (!slot.available && slot.conflictReason === undefined) {
  ctx.addIssue({ message: "Unavailable slots require a conflictReason", ... });
}

So the UI can say “Mai is off on Mondays” instead of “no slots available,” and it can say it in three languages, because the reason is an enum and not a string.

Timezones — the bug I wrote a document to prevent

Salons keep wall-clock time. Devices keep their own. A customer whose phone is in another timezone — traveller, expat, wrong clock — will see the wrong appointment time, and nothing will crash. I wrote a 10KB integration guide (CLIENT_APP_TIMEZONE_GUIDE.md) before writing the client’s date handling. Five rules:

  1. The wire is always UTC, ISO-8601, Z-suffixed.
  2. Display in the salon’s timezone, not the device’s.
  3. The date passed to availability is a salon-local calendar day, never derived from the device clock.
  4. To book, send the slot’s startAt back verbatim.
  5. The server owns the conversion.

This is the kind of defect that ships, survives QA, and gets reported months later as “the app is wrong sometimes.” Writing the rules down first was cheaper than debugging it once.

The staff / customer privilege wall

Customers and staff are both better-auth users. The distinction is principalType plus “no active organization” — a single seam carrying the entire privilege boundary. I documented it in CLAUDE.md as a known security-sensitive area with an explicit pre-ship audit checklist rather than declaring it done:

  • a customer session must never reach a staff endpoint, or vice versa
  • customer mutations must only touch that customer’s own records — ownership, not just authentication
  • public endpoints need rate limiting against booking spam and enumeration

That TODO is still open, and it’s in the case study on purpose. Knowing where your single point of security failure is, and saying so, is a more useful signal than a clean README.

AR nail try-on — the decision I didn’t make

The try-on flow in the app is fully built: camera feed, lens carousel, hint pills, capture ring, result screen. The AR is simulated — tapping the hand guide fakes a “tracking lost” state. The UI is done; only the engine is missing.

Rather than commit budget on instinct, I wrote a 14KB decision brief comparing two Snap Camera Kit integration paths, with the already-settled questions separated from the open ones (we author our own Lenses in Lens Studio; iOS first; v1 is live preview + lens switching + real tracking state + photo capture, no video; fallback is the existing plain expo-camera screen). It’s explicitly marked as seeking review from engineers with production React Native + native SDK experience, and it names the real blocker: Snap’s hand-and-nail segmentation is still beta.

Shipping a convincing UI behind an undecided engine, with a documented degraded fallback, is a deliberate sequencing choice — the product is demoable and the expensive decision stays open.


7. Honest assessment

What went well

  • The design-first pass produced one coherent product across nine modules and two platforms.
  • Extracting shared cores (computeAvailabilityForSalon, bookAppointmentCore) before adding public endpoints meant admin and customer booking can’t drift.
  • Pure-function boundaries around availability and calendar math made the two hardest pieces of logic testable and the tests cheap to write.
  • Separate public Zod schemas turned a data-leak class of bug into a one-line fix.
  • Tenant resolution from server config removed a whole category of multi-tenant vulnerability by API shape rather than by validation.

What I’d do differently

  • i18n from the first line. Retrofitting 508 keys across seven modules cost more than writing them localized would have.
  • Security review before the public API shipped, not after. The audit checklist exists; the audit doesn’t.
  • A framework this young needs an escape hatch. Several days went to TanStack Start + Vite + better-auth integration issues (allowedHosts, trusted origins, Docker networking) that a mature stack would have documented. I’d still choose the type safety, but I’d budget for it.
  • Four commits in the final week are still pulling out fake data. Seeding real data per module as it was built would have avoided a cleanup tail.

Known open items

  • Public API security audit and rate limiting (documented, not done)
  • AR try-on engine undecided; UI ships against a simulated engine
  • Android build not wired
  • Some content — salon profile, look gallery, loyalty wallet — is intentionally build-time static, because the API has no endpoint for it yet. The code says so explicitly: “Anything the API does serve is fetched, never mocked.”

8. Numbers

Admin app ~25.7k lines TypeScript, 9 feature modules, 35 test files
Mobile app ~6.1k lines TypeScript, 32 screens, 15 shared UI components
Public API 9 documented endpoints, hand-maintained OpenAPI 3.0.3 spec
Localization 508 message keys × 3 locales (en / fr / vi)
Specs & plans 7 documents, ~180KB of written design before implementation
Commits 65 (admin) + 17 (mobile)
Availability engine 15.5KB implementation, 23.7KB of tests, 7 typed conflict reasons
Onboarding cost bun start — Postgres, migrations, seed and app in one command
Project

French-Vietnamese dictionary

FV Dict: an offline French–Vietnamese dictionary for the desktop

Role: Solo designer and developer
Stack: Tauri v2 · Rust · SQLite (rusqlite) · Astro 7 · TypeScript · Tailwind CSS v4 · Bun
Platforms: macOS (.dmg / .app), Windows (NSIS .exe), Linux (AppImage / .deb)
Timeline: Desktop port built in one day (14 July 2026, about 11 hours from scaffold to Windows release fix)

1. The problem

Vietnamese speakers learning French have few good tools. Most French dictionaries translate into English, so a Vietnamese learner ends up translating twice (French → English → Vietnamese) and loses nuance each time. The French–Vietnamese dictionaries that do exist tend to be:

  • Online-only. They’re slow or useless on a train, a plane, or a weak campus Wi-Fi connection.
  • Lookup-only. They give you a gloss, but no conjugation tables, grammar help, or way to review words later.
  • Strict about input. Searching nang won’t find nắng, and searching a conjugated form like aimais won’t find the verb aimer.

The goal: one small desktop app that works fully offline and covers the whole learning loop. You look a word up in either direction, conjugate it, read the grammar behind it, save it to a collection, and review it with flashcards until it sticks.


2. What the app does

Area What it does
Bidirectional search Search French → Vietnamese (46,891 French entries) and Vietnamese → French (37,152 Vietnamese headwords). Search ignores accents and case, and đ matches d.
Conjugated-form lookup Typing an inflected form (aimais, allions, fûmes) finds the infinitive through an index of about 400,000 verb forms.
Word detail Numbered senses with French/Vietnamese example sentences, part of speech, gender, synonyms that link to their own entries, related words, and the source of each entry. French text-to-speech uses the Web Speech API.
Conjugation Full tables for 10,894 verbs across Indicatif, Subjonctif, Conditionnel, Impératif, Infinitif and Participe. Auxiliary choice, pronominal forms and elision (j', qu', s') are handled.
Grammar guide 48 short grammar topics explained in both French and Vietnamese, with examples and a mini practice exercise. Topics can be searched.
Collections Unlimited personal word lists, with a “mastered” flag on each word.
Study Flashcards with keyboard controls (Space flips a card, ←/→ marks it, P pronounces it), multiple-choice quizzes, and a list view.
Motivation Daily streak, weekly activity chart, monthly heat map, and ten ranks from Débutant I to Érudit based on how many words you’ve saved.

3. Stack decisions

Why a desktop app, and why Tauri instead of Electron

The main requirement was that everything runs offline, including a 187 MB SQLite dictionary. That ruled out a plain website, which would either need a server or would have to download the whole database into the browser. So the app needed to be native.

Tauri v2 Electron
Runtime shipped The OS’s own WebView A full copy of Chromium plus Node
Installer overhead A few MB plus the data About 80–120 MB plus the data
Backend language Rust: compiled, memory-safe, fast SQLite access Node.js
Security model Commands must be explicitly allow-listed Broad Node access unless you lock it down

The database was already large, so it made sense not to add another ~100 MB of runtime on top of it. Tauri also forces a clean split between UI and data access. The frontend can only call seven named Rust commands. It never opens the database itself.

Why Rust and rusqlite for the backend

Project

KubeLearning

KubeLearning — Case Study

A hands-on Kubernetes course with a real terminal and auto-graded labs, built as a static site.

Role Solo — product, curriculum, design, front end, Go daemon, CI
Timeline 13 July – 27 August 2026 (three focused build phases)
Stack Astro 7 · React 19 · MDX · Tailwind CSS v4 · TypeScript · Go · xterm.js · Vitest · GitHub Actions · k3d
Scale 87 lessons · 17 modules · ~120 hours of material · 234 quiz questions · 469 lab steps · 16 generated diagrams
Repo activity 41 commits · ~48k lines added

📸 Screenshot suggestion: homepage hero + a lesson page with the terminal panel open.


1. The problem

I wanted to learn Kubernetes properly: not just kubectl apply a tutorial manifest, but reach the point where I could run and debug a cluster. The official docs are complete but they’re reference material, not a path. Most courses have one of two problems:

  • They’re passive. You read, watch, and tick a box. Nothing checks that you did the work, so the checklist records intent, not skill.
  • They’re hard to start. Before the first lab you need a terminal, a cluster, and a working toolchain, and the course usually leaves you to sort that out.

The audience was me and a small internal group of 3–5 engineers, all coming from web/TypeScript backgrounds. That small, known audience shaped nearly every technical decision below.

2. What it is

KubeLearning is a structured Kubernetes curriculum, from Linux and container basics up to cluster administration, taught around one evolving application (KubeJourney: a React frontend, a TypeScript API, PostgreSQL, Redis and a worker) rather than dozens of unrelated demos.

Every lesson follows the same learning loop:

  1. Objectives and a mastery target: recognize → reproduce → apply → diagnose
  2. Concept + diagram: each diagram is generated from a JSON spec
  3. Predict, then verify: commands and manifests you run yourself
  4. Lab: a checklist of concrete steps, some auto-graded against your real cluster
  5. Break/fix: a deliberately broken scenario with progressive hints
  6. Quiz: with an explanation for every answer

Progress, streaks, quiz scores, notes and lab state are stored locally in the browser, so there are no accounts and no backend.

The 17 modules: Prerequisites · Environment · Containers · Architecture · kubectl · Workloads · Configuration · Networking · Gateway API · Storage · Scheduling · Autoscaling · Security · Observability · Packaging (Helm) · GitOps · Administration

📸 Screenshot suggestion: a module detail page showing lessons, mastery levels and progress.


3. Stack decisions

Astro + MDX for the site

The product is mostly long-form content with small pockets of interactivity (quizzes, checklists, a terminal). Astro fits that shape: lessons are MDX files in a typed content collection, pages are static HTML by default, and React is hydrated only where it’s needed.

The content collection schema (src/content.config.ts) became the backbone. Each lesson declares its difficulty, mastery level, Kubernetes version, objectives, lab steps, break/fix scenario and quiz in frontmatter, and a malformed lesson fails the build.

React islands, only where needed

Quiz, LabChecklist, LessonTerminal, navigation and progress components are React. The terminal mounts with client:visible, so 87 lesson pages don’t each pay for the xterm.js bundle on load.

Tailwind CSS v4, after a detour through a component library

On day one I tried adopting Astryx, a design-system library, across the whole app. Within about an hour it was clear it was the wrong trade: it added a layer of abstraction over a UI that was mostly typography and prose. I wrote a short design doc, removed it completely, and rebuilt on semantic HTML + Tailwind v4 tokens (a single commit: +664 / −1,927 lines). That commit was net-negative in code and kept every route, the progress behaviour and all the content intact.

The visual language is documented in DESIGN.md as a token system: a warm cream canvas, a tiered green palette with specific roles, full-pill actions and 12px cards, with Manrope, Lora and JetBrains Mono as the fonts. Colours are CSS variables with light and dark themes. Code blocks stay dark in both themes because a terminal should look like a terminal.

Go for the local agent

The terminal and grading need a process with real shell and cluster access. I chose Go for that daemon because it compiles to one static binary per platform (make release → darwin/arm64, darwin/amd64, linux/amd64), has a strong standard library for HTTP and processes, and has good PTY and WebSocket support.

Keep the site static, with no hosted infrastructure

This was the decision with the most impact. The obvious approach is a hosted backend with per-user cluster sandboxes, which means servers, isolation, cost and on-call. For 3–5 engineers who already run clusters on their laptops, none of that pays off. Each learner runs their own cluster and a small local daemon; the site stays a pure static build. (More in section 5.)

Testing & CI

  • Vitest + Testing Library for the agent client and the verified-step checklist
  • Go test suite for every agent package (config, guard, catalog, matchers, runner, server, update)
  • A curriculum contract (npm run check:curriculum) that validates lesson inventory and ordering, and that every check ID referenced by a lesson exists (and vice versa)
  • GitHub Actions that start a real k3d cluster and run every published check against it

4. How it was built: timeline from git history

Phase 1 — Foundation & curriculum (13–14 July)

Step What happened
Scaffold Astro starter, then a snapshot of the first working version: content schema, progress library, React pages, and Modules 0–3 (18 lessons). The whole curriculum was planned in Path.md, a ~2,000-line blueprint.
Design pivot Astryx migration designed → reversed the same morning → Tailwind v4 restoration (see above).
Modules 4–10 Written spec → plan → test → content: a design doc, an implementation plan, then a curriculum contract test committed before the lessons, and only then 38 lessons (~7,400 lines). A follow-up commit made every lab executable end to end.
Modules 11–16 The same loop: extend the contract, then publish Autoscaling, Security, Observability, Packaging, GitOps one module per commit, each followed by a “deepen” pass where a lesson fell short of the reference depth.
Diagrams Added a JSON spec per diagram, rendered as a structured placeholder until an SVG exists.

Writing the contract test first meant “done” had a checkable definition: correct lesson IDs, contiguous ordering and a valid schema, before a single lesson was written.

Phase 2 — Content audit (mid–late July)

Before adding features I audited the whole course and wrote it up as a remediation plan. The findings:

  • ✅ All 87 lessons pass the schema; ordering is contiguous in every module
  • ✅ 234 quiz questions, no wrong answer keys, every one with an explanation
  • ✅ No deprecated APIs (autoscaling/v2, policy/v1, gateway.networking.k8s.io/v1, native sidecars)
  • ✅ Clean build: 108 pages in 1.9s
  • ❌ The critical finding: 56 of the 87 lessons assume a running KubeJourney app that the course never provides

That last finding changed the roadmap. It is why the grading catalog below is deliberately small.

Phase 3 — Embedded terminal & auto-graded labs (30 July)

A design doc, then a 15-task TDD implementation plan, then one commit per task:

feat(agent): config loading and pairing token
feat(agent): origin, host, and token guard
feat(agent): check catalog loading and validation
feat(agent): expect matchers
feat(agent): cluster command runner
feat(agent): health and verify endpoints
feat(agent): update command and CLI entry point
feat: publish check catalog and allow verified lab steps
feat: validate check catalog against lesson references
feat: agent client and verified-step progress storage
feat: verified lab steps in LabChecklist
feat(agent): interactive PTY over websocket
feat: embedded terminal in lesson pages
feat: pilot check catalog for app-independent lessons
ci: run the check catalog against a real cluster; release builds

Finally, the starter README was replaced with real setup docs. I ran every command and error string in them against the real system before writing it down.

Phase 4 — Diagram generator (27 August)

The hand-made SVGs had clipped text and overlapping labels, so I replaced them with a generator (scripts/generate-diagrams.py) that:

  • measures text using the real advance widths of Manrope and JetBrains Mono (extracted by a second script), so wrapping is exact
  • computes the viewBox from the content bounds and normalises every diagram to the same type size
  • places edge labels so they never collide with nodes, strokes or other labels
  • uses only CSS-variable colours, so every diagram switches between light and dark themes

The SVGs are now build artefacts: to change a diagram you edit its spec or the generator, never the SVG.

📸 Screenshot suggestion: one diagram in light and dark theme side by side.


5. Deep dive: grading labs without a backend

browser (static Astro page)          kubelearn-agent (127.0.0.1:7788)
  ├── LessonTerminal  ──ws────────►  WS  /pty      → a real shell
  ├── LabChecklist    ──POST──────►  POST /verify  → runs a named check
  └── agent.ts        ──GET───────►  GET  /health  → version, context, kubectl

The page sends a check ID, never a command. The agent looks the ID up in its local catalog and runs the matching command itself. So /verify can’t be used to execute arbitrary commands, and the only shell exposure is /pty, which exists to be a shell.

Checks are data, not code. Each one is a list of assertions. Every assertion is an argv array (no shell, so nothing to interpolate or inject) plus exactly one matcher (equals, contains, matches, exitZero) and a hint:

{
  "run": ["kubectl", "get", "nodes", "-o", "jsonpath={.items[*].spec.unschedulable}"],
  "expect": { "equals": "" },
  "hint": "A node is still cordoned. You drained it but never uncordoned it — run `kubectl uncordon <name>`."
}

The first failing assertion supplies its hint, so a failed check tells the learner what to try next, not just that they failed.

Security model. A localhost shell is a serious attack surface, so there are four layered controls:

  1. Loopback-only binding. The agent refuses to start on any other address.
  2. Origin allowlist. Browsers set Origin and pages can’t forge it, so another tab can’t open a shell.
  3. Host validation. This is the anti-DNS-rebinding control; without it, Origin alone isn’t enough.
  4. Pairing token. This blocks non-browser processes on the same machine. It is accepted in the query string only on /pty, because browsers can’t set WebSocket headers and query strings leak into logs.

Graceful degradation. With no agent running, the terminal panel explains how to start one and every lab step falls back to a manual checkbox, exactly as before. The terminal and the Check buttons add to the course; nothing requires them.

📸 Screenshot / GIF suggestion: running a lab step, clicking Check, seeing a failed hint, fixing it, then a pass.


6. Challenges & what I learned

Running the checks found bugs that reading didn’t. The pilot catalog uncovered three problems, all fixed in the same commit:

  • lsof -i :8080 can report a closed socket; it needs -sTCP:LISTEN
  • the nodes check said “a node is not Ready” when kubectl couldn’t reach a cluster at all, so every check now starts with a connectivity assertion
  • a failing kubectl dumped ~700 characters of stderr into the page, so output is now capped at 200 characters

A check that can never pass destroys trust. That’s why CI starts a real k3d cluster and runs the entire published catalog against it. Checks that depend on state the learner creates are marked requiresLearnerState, and CI skips them instead of faking them.

If a step can’t be asserted, it can’t be graded. A step like “create a namespace with an identifiable ConfigMap” can’t be checked because there’s no fixed name to assert on. Lesson prose now has to pin resource names.

Refuse, don’t ignore. The appVersion field is reserved for a future app release ladder. Because nothing can report the app version yet, validation rejects the field; silently ignoring it would let authors believe it works.

SSR edge cases. xterm.js is CommonJS and touches the DOM when imported, which broke Astro’s server render. It’s now imported dynamically inside the effect.

Scoping honestly. The grading mechanism is complete and tested, but the catalog has only 3 checks. That was a deliberate decision: a check has to assert on something that exists, and 34 of the labs depend on the KubeJourney app, which doesn’t ship yet. So I built everything that wasn’t blocked and documented the rest as sequenced work instead of padding the numbers.

7. Process

  • Spec → plan → test → build for every major feature. The design docs and plans live in docs/superpowers/.
  • Contract tests before content, so “done” is something a script can check.
  • Small, single-purpose commits. The terminal feature is 15 commits that each map to one planned task.
  • AI-assisted development. I used Claude Code as a pair programmer for planning, TDD implementation and review. I owned the product direction, the architecture and scope decisions, and verifying everything against a real cluster.

8. Results

  • A complete 87-lesson, 17-module curriculum with a consistent learning loop and a verified-accurate quiz bank
  • A static site that builds 108 pages in under 2 seconds, with no backend and no accounts
  • Auto-graded labs against the learner’s own cluster, with a threat-modelled local daemon
  • A CI pipeline that proves every published check can actually pass on a real cluster
  • A deterministic diagram pipeline in which every figure is generated from a reviewable spec

9. What’s next

  1. Ship the KubeJourney reference app. This unblocks 56 lessons and the bulk of the grading catalog.
  2. Backfill checks across the remaining labs, now that the mechanism is in place.
  3. Close the remaining pedagogical gaps and the small defects listed in the remediation plan (including 7 broken external links).
Project

Discord agentic orchestrator

Case Study: Discord Dev Agent

From a Product Owner’s forum post to a reviewable merge request, with an AI coding agent in between that is never allowed to ship on its own.

Role Solo: design, architecture, implementation, testing
Timeline September 2026
Stack TypeScript · Node.js 22 · discord.js 14 · Claude Agent SDK · SQLite (better-sqlite3) · GitLab CLI (glab) · Zod · Pino · Vitest
Size ~2,200 lines of application code, ~720 lines of tests, 83 passing tests across 10 suites
Status Verified live against a private GitLab project (nailshoppoc); full Discord end-to-end run pending

1. The problem

On a small product team, the Product Owner spots bugs and wants small changes all the time: a spinner that never stops, a label with a typo, a button that doesn’t close its modal. Each one follows the same slow path:

  1. The PO writes it up somewhere (chat, a doc, a ticket).
  2. A developer reads it, switches context, and finds the code.
  3. The developer makes a change that often takes ten minutes, then opens a merge request.
  4. Someone reviews and merges it.

Step 3 is the part a coding agent can now do well. Steps 1 and 4 are where people need to stay involved: the PO knows what is wrong, and a developer must decide whether the fix ships.

The goal: let the PO write feedback in a place they already use, and have an agent turn it into a draft merge request. The agent works inside hard limits: it never touches main, never deploys, and never merges.

The main design constraint was the review boundary. An agent that can change production because someone typed in Discord is a liability, not a tool.


2. What it does

A PO opens a post in a Discord forum channel and tags it agent-ready. The bot then:

Discord forum post (tagged `agent-ready`)
      ↓
FeedbackJob persisted in SQLite, keyed by thread ID
      ↓
Isolated git worktree from origin/<baseBranch>
branch: discord/<threadId>-<slug>
      ↓
Headless Claude coding agent runs inside the worktree
      ↓
Project validation commands (typecheck, tests…)
      ↓
Safety check on the real diff → commit → push → DRAFT merge request
      ↓
Progress reported back in the same thread:
🔎 Investigating → 🛠 Implementing → 🧪 Validating → 📦 MR ready

The PO never leaves Discord. The developer gets a draft MR containing the original feedback, the agent’s root-cause summary, and a ✅/❌ list of validation results.

Behaviour at a glance

PO action What happens
Creates a post without agent-ready The bot acknowledges it. No job runs and nothing is spent.
Creates a post with agent-ready A job is queued. The tag moves agent-ready → in-progress → review.
Adds agent-ready later The job starts then.
Edits the post mid-run Edits are merged into one follow-up run that resumes the same agent session, so the agent does not start from scratch.
The subscription hits its usage limit The thread shows “⏳ paused” and the job retries automatically after 15 minutes. It is not marked as failed.
The bot crashes mid-run On restart the job is marked failed, the thread is notified, and the worktree is cleaned up.

3. How it was built

Phase 0: Brainstorming the architecture (docs/brainstorming.md)

Before writing any code I did a design pass and kept the result in the repo. The starting idea was:

Discord bot → handle feedback → launch Claude → create a prompt → launch agent

The brainstorm changed that plan in several ways, and most of those changes made it into the final design:

  • Forum channels, not a regular text channel. Each piece of feedback becomes its own thread, so the bot has a natural place to reply and the thread ID becomes the job’s primary key.
  • Tags as commands. Nothing runs without the agent-ready tag. This keeps brainstorming posts from starting expensive agent runs, and it makes approval a visible action in the UI instead of a hidden setting.
  • Never give the agent main. Every job gets its own worktree and branch, and the output is a merge request for a human to review.
  • One pipeline stage instead of two. I dropped the separate “prompt-writing” LLM call. The coding agent can do the analysis itself. A triage agent is noted as a possible future step, not an MVP requirement.
  • Serialize work per thread. If the PO edits the post while the agent is working, don’t start a second agent on the same branch. Queue the edit and resume the session.
  • Report state changes, not every tool call. Four phases in Discord instead of hundreds of Read/Grep events.

Phase 1: The full pipeline (commit 32cadd8, Initial commit)

The first commit contains the whole system: 39 files and about 6,000 lines including the lockfile. I built it in layers, each with its own folder and tests:

Layer Files Responsibility
Config config/env.ts, config/repos.ts Zod-validated env; one forum channel → one repo mapping
Jobs jobs/store.ts, queue.ts, worker.ts, reporter.ts SQLite persistence, per-thread serialized queue, the pipeline itself
Git git/worktree.ts, gitlab.ts, guard.ts, exec.ts Worktree lifecycle, glab MR creation, forbidden-path check
Agent agent/prompt.ts, runner.ts Prompt builder, Claude Agent SDK runner, phase detection
Discord discord/client.ts, forum.ts, report.ts Events, tag resolution, in-thread reporting
Validation validate/run.ts Runs the repo’s configured checks
Dev loop scripts/dry-run.ts The whole pipeline without Discord

The dry-run script mattered most during development. It runs worktree → agent → validation → MR from the command line, so I could test the core against a real GitLab repo before connecting a Discord bot. It also has a --preflight mode that checks git and glab credentials without a TTY. That check exists because a headless git push with no credential helper blocks forever on a password prompt.

Phase 2: First live run, first real bug (commit 0d08455)

The first live run against the nailshoppoc repo found a bug that unit tests had missed.

Symptom: Discord went straight from 🔎 Investigating to 🧪 Testing. The 🛠 Implementing update never appeared.

Root cause: Phase detection treated any Bash command containing the word “test”, “build”, or “check” as a validation run. The agent’s first investigation commands, such as grep -rn test src/, matched. Phases only move forward, to stop a late Read from moving the report backwards, so once the phase reached testing it could never go back to implementing.

Fix: I extracted a PhaseTracker class with two stricter rules:

  1. A command counts as validation only if it actually runs a validation tool (pnpm test, npx tsc, vitest, cargo test…), not if it merely contains the word. A bare test or build counts only behind a package runner.
  2. testing cannot be reached before an edit has happened. An early npx tsc during investigation is exploration, not validation.

I added 18 phase tests, including the exact regression. The same commit added a step-by-step Discord setup guide (application, privileged intent, invite permission integer 292057861120, forum tags, channel ID) and a troubleshooting table, because setup was the other thing the live run showed to be hard.

I kept this commit in the case study on purpose. The first design looked right and passed its tests, but real agent behaviour broke an assumption I didn’t know I was making.


4. Stack decisions

TypeScript on Node.js 22

Both discord.js and the Claude Agent SDK are first-class in TypeScript. The domain has many state machines (job status, agent phase, workflow tag), and strict types catch invalid transitions at compile time.

Claude Agent SDK instead of shelling out to claude -p

The brainstorm suggested claude -p as the fastest prototype route. I used the SDK because the orchestrator needs to observe the run, not just wait for its output:

  • Session IDs from the init message are stored so an edited post can resume the same session.
  • Tool-use events drive the phase reporting in Discord.
  • Structured result messages separate success, error, rate limit, and timeout.
  • Tool allow/deny lists: the agent can Read/Edit/Bash but has no WebFetch or WebSearch.

Subscription billing, not the metered API

The SDK starts the local claude binary, which uses the Pro/Max subscription. A detail that is easy to miss: if ANTHROPIC_API_KEY reaches the subprocess, the CLI silently switches to metered billing. Because the SDK replaces the child environment, subscriptionAgentEnv() copies process.env and explicitly removes the key. At runtime the bot logs a warning if the CLI reports any apiKeySource other than none.

This choice also shaped capacity. Subscription rate limits are the real ceiling, so the defaults are MAX_CONCURRENT_JOBS=1 and the agent-ready gate, and a usage-limit hit is treated as “paused, retry later” rather than as a failure.

SQLite (better-sqlite3), not Redis or Postgres

This is one process with low throughput, and the state has to survive a crash. SQLite in WAL mode gives durable state with no extra infrastructure. The job table is keyed by thread ID, and create is an idempotent upsert, so re-tagging a post or restarting the bot never creates duplicate jobs.

An in-process queue instead of BullMQ

The queue enforces two rules and only needs about 80 lines to do it:

  1. Never two agents on the same thread. Enqueuing a running thread marks it “dirty”. After the current run, exactly one follow-up runs, however many edits came in meanwhile.
  2. A global concurrency cap.

A distributed queue would add a Redis dependency to solve a scaling problem this tool doesn’t have.

Git worktrees instead of clones

Worktrees share the object store with the main clone, so creating one is almost instant and uses little disk. Each job gets its own worktree from origin/<baseBranch>, never from a local checkout. The worktree is always removed in a finally block.

GitLab CLI (glab) instead of the REST API

glab already handles auth through the OS keyring, and the target project was on GitLab. Two details make it safe to run unattended:

  • It is called with execFile and an explicit argv array, never a shell string, because MR titles come from Discord post titles that users write.
  • It uses --yes, an explicit -t, and --description-file -. Without them glab opens an editor and the worker hangs. --fill is deliberately avoided.

Zod for configuration

Invalid env or repo config fails at boot with a readable list of problems, not mid-job. A forum channel without an agent-ready tag also fails startup, because otherwise the bot would never trigger and give no sign of why.


5. Safety as code, not just prompt text

The prompt tells the agent to stay away from secrets and deployment config. A prompt is not a control, so the important rules are enforced in code:

Rail How it’s enforced
Never touch main Worktrees always branch from origin/<base>, and main is never checked out
Never modify secrets, CI, or infra git/guard.ts checks the actual diff before push and refuses changes to .env*, CI configs, Dockerfiles, k8s/Helm/Terraform/Ansible, deploy manifests, and key files
Never merge Every MR is a draft, with no auto-merge
No shell injection from Discord text execFile with argv arrays everywhere
No silent failures Failed validation does not block the MR. It shows ❌ with the output tail in the thread and the MR body, because a red MR is better than one that silently disappears
No attachment leakage Discord attachments are staged in .agent-context/, excluded through the worktree’s own info/exclude so they never enter the diff
No billing surprises The API key is stripped from the child environment, and the credential source is checked at runtime

6. Engineering details worth highlighting

  • Crash recovery. Job status is saved at every step. At boot, any job still marked as in progress is marked failed, because a crashed worktree can’t be resumed safely. The thread is told why, and jobs with unprocessed feedback go back on the queue.
  • A low-noise Discord reporter. One status message is edited in place with a debounce, so long runs don’t flood the thread. The final result is a new message so it triggers a notification.
  • Edit-aware prompts. When a run resumes after an edit, the prompt includes the previous feedback text so the agent can see what changed.
  • Branch-safe slugs. Post titles are normalized: accents removed, and every character git forbids in ref names stripped.
  • Tests at the right level. Pure functions (the glab argv builder, prompt builder, slugify, the path guard, phase detection) get unit tests. The worktree lifecycle has an integration test against a real git repo, including recovery when an earlier run left a worktree behind.

7. Results

  • ✅ Verified live against a private GitLab project: preflight, headless push with no TTY, worktree from origin/master, agent run on subscription billing, both validation commands passing, a draft MR opened, and the worktree cleaned up (confirmed with git worktree list).
  • ✅ 83 automated tests passing in about 1.6 s.
  • ⏳ Still to verify: a full end-to-end run in a private Discord guild (post → tag → MR link in thread → mid-run edit resuming the same session_id) and a crash-recovery drill.

8. What I learned

  1. The hard part of an agentic system is the orchestration around the model, not the model call. The agent call is one function. Isolation, idempotency, crash recovery, credential handling, and cost control make up the rest of the codebase.
  2. Watching real agent behaviour catches bugs that unit tests miss. The phase-tracking bug only appeared once a real agent ran real grep commands. A heuristic about what an agent will do should be tested against what it actually does.
  3. Headless means no prompts, ever. Password prompts, editors, and confirmation dialogs can each hang an unattended worker. The preflight check exists for this reason.
  4. Make the approval step visible. A Discord tag is a simple UI, but it makes “yes, spend compute on this” an explicit, visible, reversible decision by a person.

9. What’s next

  • A triage agent in front of the coding agent: turn loose PO feedback into structured acceptance criteria, or ask a clarifying question in the thread before any code is written.
  • GitHub support next to GitLab, behind the same createMergeRequest interface.
  • MR status sync. When a developer merges or closes the MR, move the Discord post to completed.
  • Run as a service (launchd or a container) using GITLAB_TOKEN instead of the desktop keyring. This is already supported through authMode.

Fujifilm XT-30 II

Fujifilm X-T30 II with a Viltrox 35mm f1.8 mounted

Viltrox 35mm f1.8

Fujifilm XT-30 II + Viltrox 35mm f1.8

SS
1/250
f
4.5
ISO
160

Contact

Have an idea, a role, or a question? Let’s talk.