Productivity Programming Language

I Replaced My Productivity Apps With One I Built With AI. Here Is What Worked and What Broke.

Phone: the Today screen

A personal, non-commercial project. Concept to working phone app, with the mistakes left in.

TL;DR

  • I built one app for notes, databases, tasks, planning, habits, journal, food log and mind maps with two AIs in separate roles: Gemini designed the architecture and wrote the specification, Claude Code wrote and tested the code.
  • The big win was not code generation. It was writing a spec first, then making the AI audit that spec by executing it.
  • The big pain was not features. It was sync, data safety and the gap between "tests pass" and "works on my phone".
  • AI can replace the features of your apps. It cannot replace the operations: hosting, auth, backups, device quirks, and deciding what is safe.

1. The idea

I had the usual sprawl: one app for notes, one for databases, one for tasks, one for highlights, one for the calendar. Every tool added a place to look and none of them agreed with each other.

The spark was a walkthrough video in which Tom Solid (Dr. Thomas Roedl, co-creator of the myICOR method with Paco Cantero) shows Claude Code generating a bespoke productivity system from a few prompts. The method organises life as Inner World, Key Elements, Goals, Projects, then Tasks and Habits, and runs on four moves: Input, Control, Output, Refine.

My question was simple: if an AI can write the app, why keep paying attention to ten apps that almost fit?

This is a private tool for one user. It is not a product, it has no customers, and it is not affiliated with myICOR. Its colours and fonts were chosen to match that method's look for personal enjoyment only.

2. What the app does today

Job What the app has Honest status
Notes (Obsidian style) Live-preview Markdown, [[wiki links]], backlinks, tags, graph view Replaced
Databases (Airtable style) Grid, Board, Gallery, Calendar, List views, formulas, rollups, linked records, CSV Replaced
Planning Day hour-grid, weekly planner, habit matrix, weekly review, recurring tasks Replaced
Tasks - [ ] Call supplier #todo in any note becomes an Inbox task; "by friday" in a title sets the due date Replaced, with two-way sync to Todoist as a bridge
Reading highlights Readwise import into a Books database without duplicates Integrated, not replaced
Calendar Read-only ICS feed on the Today page Integrated, not replaced
Food log Photo or text to calories and macros, reviewed before saving Replaced
Mind maps Nodes become linked notes, AI expands a branch, FreeMind export Replaced
Phone Installable PWA with bottom tabs, offline, capture in one field Replaced
Agents Terminal for Claude Code inside the app, and a 15-tool MCP server over my Markdown folder New capability

Desktop: the My Life dashboard, with the Inner World, key elements with habit health, and projects in motion
Desktop installed app, dark theme: the My Life dashboard (sample data).

Desktop: the Tasks database as a board grouped by status
The same Tasks database as a Board. Grid, Gallery, Calendar and List are saved views over the same rows.

Size at the time of writing: about 30,000 lines of TypeScript, 71 test files, 623 test cases, 9 database migrations, 8 serverless functions.

3. Two AIs, two jobs

I did not use one AI for everything. I gave each the job that matches its strengths.

Gemini Claude Code
Role Architect Engineer and reviewer
Works in A chat, with the whole picture in one long conversation The repository, terminal and test runner
Produced The architecture and the specification (v1.0.0 to v2.1.0): data model, rules, wireframes, design tokens The code, the tests, the migrations, the pull requests, and written reviews of each spec version
Strength used Broad design thinking and drafting a long, structured document Reading a real codebase, running commands, fixing what the run shows
Cannot do Run the SQL it just wrote against a real database Know what you want unless the spec says so

The hand-off between them was the specification. Gemini wrote it, Claude executed it. Running Gemini's SQL in a real SQLite is what exposed the defects in Stage 2, because a design chat has no database to run against. A second model also reads a document without the author's blind spots.

Gemini also ended up inside the app: the in-app assistant asks Gemini first and falls back to OpenAI if it fails. That is a separate use from building.

Treat this split as my observation of what each tool did well for this project, not a ranking. Models change quickly, so test the split on your own task.

4. From concept to screen, in five stages

Stage 1: A prompt that states invariants, not just features

The first prompts did not say "build a task app". They said what must always be true:

  • Personal items may only be in the PKM or PPM quadrants, business items only in BKM or BPM, enforced by database constraints.
  • Project progress is computed by database triggers from its tasks, and a project whose tasks are all done moves to Review.
  • Capture must take under two seconds and land in one Inbox.
  • The colour tokens, fonts and spacing are listed explicitly.

Rules the AI can test beat adjectives it can interpret.

Stage 2: The spec, and the AI auditing it

Gemini drafted the spec, and it went through six versions, v1.0.0 to v2.1.0. After each draft I handed it to Claude to review by running the SQL in a real SQLite, not by reading it. That found real defects before any UI existed:

  • A view selected t.created_at from a table that has no such column. It would have failed on first query.
  • A full-text index pointed at a text primary key. A routine VACUUM can renumber rowids and silently corrupt such an index.
  • A title was stored in nine different tables, so a rename left eight stale copies.
  • A completed project with no tasks showed 0%.
  • Two sources of truth for a task's scheduled day.

Even v2.1.0 still had seven open defects when audited. I stopped revising the document and recorded them as deviations to fix during the build.

Lesson: a spec is code that has not run yet. Run it.

Stage 3: Build in phases, in layers

React UI (pages, components)
        |  hooks
Data layer in TypeScript (no UI)
        |  query() / run() / transaction()   <- the only door to storage
SQLite (WASM) in the browser, with triggers for rollups
        |
Folder sync: plain Markdown files, readable by Obsidian
Cloud sync:  push / pull API on Vercel -> one Postgres table of JSON rows

Two decisions paid for themselves many times:

  1. Everything above one file speaks SQL through three functions, so the database can move without touching screens. This is what made the cloud move possible later.
  2. Tests run against the real SQLite engine in memory, and the cloud tests run against a real PostgreSQL in memory (PGlite). Mocks of the database would have hidden most of the sync bugs listed below.

Stage 4: Visualization

The spec contains wireframes drawn in ASCII, then the design system as tokens:

Token Value
Ink background #0C0E12
Cream text #F6F3EC
Accent orange #FF5A2D
Fonts Bricolage Grotesque (headings), Instrument Sans (interface), Spline Sans Mono (code)

From there the AI built the pages: the day hour-grid with overlapping blocks side by side, the weekly planner with drag and drop, the board view, the knowledge graph, mind maps, and a separate phone shell with a bottom tab bar that respects safe areas and the on-screen keyboard.

Desktop: the weekly planner in the cream Paper theme, with the Inbox triage drawer and the habit matrix
The weekly planner in the cream "Paper" theme: Inbox triage on the left, tasks with days attached, habit matrix below.

Desktop: the knowledge graph linking goals, projects, notes, people and habits
The knowledge graph. This is the view that once went blank.

On the phone the app is not a squeezed desktop. It has its own shell with a bottom tab bar, big tap targets and the day as a list to tick off.

Phone: the Today screen with mood, tasks, habits and a food card
Phone: the Capture screen, one field that files into the Inbox, Today or the journal
Phone: the More screen listing every page and database

Installed phone app: Today, Capture and More (sample data).

Visual work is where AI is weakest, because the failure is not an error message. A blank graph canvas, a layout that overflows on one phone, a button hidden behind a keyboard: none of these fail a unit test.

Stage 5: The cloud, which was the real project

Local-first (the idea described in the Ink and Switch paper "Local-first software", Kleppmann et al., 2019) gave me speed and offline use. It also meant my data lived in one browser profile, and browsers may evict it. A development database was once found empty, cause unknown. That single event drove everything after it:

  1. Folder sync to plain Markdown files.
  2. One-file JSON backup including attachments.
  3. Cloud sync to a free-tier Supabase PostgreSQL through my own small push and pull API, with the browser SQLite staying as the working copy.

Decisions I wrote down before coding, each with a reason: no hybrid logical clocks (device clocks never decide anything), a three-way merge against the last synced copy, the server assigns versions, and the server holds data only, with no triggers, so rollups never fire twice.

5. The workflow that held it together

  • Docs as the AI's memory. A new session starts cold. PLAN-history.md and CLOUD-MIGRATION.md hold decisions, a backlog of small stories with sizes and "Done" marks, and a dated progress log. Every session read them first.
  • One question per decision. The AI listed choices (A to E) with a default; I approved or changed each. It never decided architecture silently.
  • Small pull requests. In one four day stretch (6 to 9 October) the repository gained cloud sync, attachments in the cloud, Readwise import, a phone layout, mind maps, a food log and MCP server connections. 14 separate Claude sessions appear in the commit trailers, each ending as a pull request.
  • Every bug fix ships with a test. The sync fixes below each added a test that reproduces the original failure first.
  • A guard against AI "tidying". A test hashes every shipped migration and fails if one is edited. The right change is always a new migration, because browsers that already ran the old one never run it again.

6. The pitfalls, with the evidence

# What happened Why it happens What I do now
1 The spec's SQL had defects that only appeared when executed Reading looks right; running does not lie Make the AI execute every spec and schema
2 Cloud pull skipped rows for good The cursor was compared as text, so '1000' sorted before '200' Test with more than 1,000 rows; never trust implicit types
3 One missing parent row rolled back a whole sync A pull was one transaction; a child arrived before its parent Rows whose parent is missing wait, then settle in a loop
4 Duplicate tasks from #todo lines A note arrived before its task; the phone journal also saved over the task id Wait while a sync cycle runs; merge edits into the current text
5 Readwise worked in tests, failed live The mock accepted a trailing slash that Vercel's rewrite does not match Call each real service once by hand before trusting a mock
6 Todoist and AI providers were never run against live APIs in tests Fakes prove my assumptions, not the vendor's behaviour Label them "not verified live" in the README
7 The access prompt never appeared on the installed phone app The browser login box needs a page request; an installed app opens from its offline cache A modal key screen and an HttpOnly session cookie
8 Claude would not start from the terminal bridge on Windows (exit code 193) npm installs an extension-less script Windows cannot run Look up .exe, .cmd, .bat, and start .cmd through cmd.exe
9 The graph view went blank Force layout with hub nodes diverged to NaN Stabilise the simulation, and add a browser test
10 The phone layout crashed on a newer Chrome A browser behaviour the unit tests cannot see Browser end-to-end check in CI
11 19 native dialogs (alert, confirm, prompt) AI reaches for them by default; installed apps handle them badly One in-app modal
12 Free tier surprises Supabase free has no automatic backups, pauses after 7 days idle, and its direct connection address is unreachable from Vercel Weekly encrypted dump restored into a throwaway database with a row count check, a weekly keep-alive call, and a clear error message
13 Spec and code disagreed on palette and fonts An old spec can make a later prompt "correct" the code back to the old values Update the spec in the same change
14 The architect's spec and the engineer's code drifted apart Two models, two documents, one truth The spec records every deviation; the engineer's review notes go back into it

Three themes sit under this table.

Mocks and green tests lie politely. My own README says it plainly: real Gemini and OpenAI calls are tested against a mock, and the Todoist client was tested against a fake. I would rather say so than imply otherwise.

Security is yours to think about. The AI will happily build a shell on a socket if you ask for a terminal in the browser. I locked the bridge to loopback, a token on every request, an origin allow-list, a host header check, allow-listed commands and folders under my home directory. Passwords pasted into a chat are rotated before use. Tokens that follow me across devices sit as plain text in my own database, and I wrote that trade-off down instead of hiding it.

Give AI inside the app a leash. The mind-map AI and the assistant only propose; I tick what to apply. The MCP server never deletes. External tools are read-only unless I allow more. Food photo estimates can be off by 20 to 30%, mostly portion size, so the result fills a form I check before saving.

7. What AI did not replace

  • Operations. Hosting, secrets, backups, restore drills, keeping a free database awake.
  • Judgement on data loss. Every rule on delete, conflict and restore was a human decision. Fred Brooks's point in "No Silver Bullet" (1986) still holds: tools remove accidental complexity, not essential complexity. Conflict handling between devices is essential complexity. See also Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017).
  • Testing on real devices and real services.
  • Knowing what you want. The method (ICOR) came from people who had used it for years. The AI built the tool, not the thinking.

So can we replace all our apps? For features, largely yes. For the guarantee that your data is safe tomorrow, you now own that job.

8. A playbook for early developers and vibe coders

Vibe coding is Andrej Karpathy's 2025 term for building by describing what you want and accepting the code the AI writes. It is a fine way to start. These steps are how to stop it from collapsing at scale.

  1. Write the spec first. Goals, data model, rules that must always hold.
  2. Split the roles. Use one AI to design and draft the spec, and a different one, with repository access, to build and to audit.
  3. Make the AI run the spec. Execute the SQL, the formulas, the edge cases, then fix the document.
  4. Pick a boring architecture with one door to storage, so you can change the database later.
  5. Test on the real engine. In-memory SQLite and in-memory Postgres are cheap and honest.
  6. Keep a decisions file and a backlog the AI reads at the start of every session.
  7. Ask for options with a default, and approve architecture yourself.
  8. Ship small. One pull request per story, one test per bug.
  9. Protect history. Never edit a shipped migration; append.
  10. Plan data safety before features: export, mirror, backup, restore drill.
  11. Look at it on your phone, every day. Dogfooding finds what tests cannot.
  12. Label what is unverified. A README line saying "not tested against the live API" is worth more than ten green checkmarks.
  13. Treat any power tool as a security task. Terminals, proxies and key storage need a threat list before code.

9. Closing

The most useful thing I learned is that AI changes the cost of building, not the cost of being wrong. Building got cheap. Finding out you are wrong, on a real phone, with real data, is still the expensive part. Design your process around finding out early.

Sources

  • Andrej Karpathy, post introducing "vibe coding", X (Twitter), February 2025.
  • Martin Kleppmann, Adam Wiggins, Peter van Hardenberg, Mark McGranaghan, "Local-first software: You own your data, in spite of the cloud", Ink and Switch, Onward! 2019.
  • Martin Kleppmann, Designing Data-Intensive Applications, O'Reilly, 2017.
  • Fred Brooks, "No Silver Bullet: Essence and Accident in Software Engineering", 1986.
  • myICOR method: Dr. Thomas Roedl (Tom Solid) and Paco Cantero.
  • Project documents: docs/PLAN-history.md, docs/CLOUD-MIGRATION.md, docs/ARCHITECTURE.md, and the specification versions in docs/specs/.

Join the discussion

This site uses Akismet to reduce spam. Learn how your comment data is processed.