All work

Salient

Private on-device AI for ADHD: research to store build in ten days

Role
Solo: product, brand, design and release
Team
Solo; Claude Code agents write the code
Timeline
Sep 19–28, 2026
Scope
on-device AI · solo
Status
In beta
Live
salient.day
  • All AI runs on the phone: no cloud, no analytics
  • Measured the AI on 44 hand-labelled dumps: 94% of items found, invented items cut by more than half
  • From research to a store build in ten days

I have ADHD, and every AI app that promised to empty my head sent it to a server and billed me monthly. Salient runs entirely on the phone. Talk, and it sorts the noise into tasks, shopping and notes, then shows one thing to do now. The hardest call was telling weaker phones “no” rather than shipping a small model that lost most items and made up others.

Salient logo on a dark background with the line “Private voice brain dump. One step at a time.”

People who give up on these apps leave without a word

I set the questions, and research agents collected about 1,900 App Store reviews of 14 apps and 40 Reddit threads. Price, trials and cancellation made up 43% of the low reviews.

  • Quitting is silent. Only 3 of 625 low reviews mention giving up, yet Reddit is full of it. So I built for coming back, not for streaks.
  • Complexity drives 12% of low reviews. So a plan is just tasks with dates, and a project is a task with steps.

Home screen with a single microphone button and one “Now” task card Recording screen sorting a spoken dump into a task and a shopping item while the person talks Calendar with tasks placed on dates taken from the person’s own words

Capture, sorting while you talk, and dates taken from your own words.

All AI runs on the phone, at the cost of a 3.2 GB download

Speech recognition (Parakeet) and sorting (Gemma 4 E2B) run on the device, and audio isn’t stored. A test fails the build if a transcript variable is logged or an analytics package appears.

Weaker phones get a “no” instead of a model that invents

I planned a lighter version on a small model, Qwen 0.8B. In a messy dump it found fewer than four in ten items; the large model found more than nine in ten (0.39 vs 0.92). It also ran twice as slowly and copied examples from its instructions into the results.

From my decision log: the first impression on a weak phone comes not from a caption but from milk appearing in the shopping list that the person never mentioned. That is worse than “can’t handle it”.

On phones with 3–4 GB of memory, the app now says it can’t sort, and recordings wait in the Inbox. My own 3 GB test iPhone lost sorting too.

“Now” went from eleven blocks to one card

Eleven blocks, the time shown twice, three ways to say “done”.

Now it’s one card with one button that changes with the moment (start → next step → done). Everything else is folded away.

An earlier version of the Now screen with a large timer ring, a step list and two buttons Come up for air screen: time in hyperfocus, water, food and stretch options, and where to pick up

Left: the earlier “Now” screen, before the rebuild. Right: the hyperfocus break.

A phone call ends the recording instead of recording through it

For an ADHD brain, losing a thought you already put into words is the worst failure. Yet dumps were saved only after “stop”, and a call paused the mic while the screen still said “Listening”. Now each phrase is saved as it’s spoken, and a lost mic means a clean stop. Recording through the call would have needed extra permissions and cost battery.

The companion helps with the task and never talks about you

The companion is a cookie with eyes. Research on Clippy gave it two rules: it never steps in on its own, and it never talks about the person.

“You can do it” judges the person; “open your email” helps with the task.

There’s no “streak”, “failed” or “missed”, and no red counters.

I stopped tuning the AI by eye and started measuring it

After two days of tuning by feel, I built a test set of 44 messy dumps in Ukrainian, mixed Ukrainian-Russian and mixed Ukrainian-English. I wrote and labelled them myself as text, so speech-recognition errors aren’t included, and the same set guided the tuning; a separate held-out set is next. Each lists the items a person would expect. The score weights what hurts most, starting with a lost item (40%) and a wrong type (25%).

Run (44 dumps) Score Items found Invented per expected item
Qwen 0.8B, Sep 23 0.458 0.39 0.14
Gemma 4 E2B, Sep 23 0.887 0.92 0.29
Gemma, Sep 24 0.904 0.92 0.15
Gemma, Sep 26 0.918 0.94 0.11

By September 26 the model found 94% of items and invented about one for every nine it should find. The last gain came from fixing two instructions that contradicted each other. The numbers also drove design calls:

  • Duplicates are flagged, not deleted. Of six matching rules, I chose the one that caught 20 of 27 duplicates with no false alarms.
  • I fixed the measure itself. A dump with an invented “tomorrow” on every item had scored a perfect 1.000.

Accessibility is measured, not eyeballed

A day with tasks used to be a 5×5 dot that a screen reader couldn’t see and a colour-blind person couldn’t tell apart. Now it’s a shape, and its label says “has tasks”. Tests run key screens at 1.3× and 2× text, where one button used to fall off the screen, and check every colour pair’s contrast in both themes. Calm mode stops the companion’s motion, as WCAG 2.2.2 requires for anything moving longer than five seconds. It also follows the phone’s reduce-motion setting. A pass with VoiceOver and TalkBack on a real phone is still ahead.

How I worked

I direct AI agents in Claude Code, and they write the code. I own the problem, the brand (logo, icon and splash are mine), the look and feel and the release.

The audit I set up on September 24 was built to argue with itself. Seventeen auditor agents raised 173 findings. Then “sceptic” agents, each told to disprove one finding, took on the 40 most serious, and 37 held up. The worst was that the app deleted its own model file right after install. All six critical findings are fixed and checked on a phone. It’s a code review, not a test: no agent ran anything on a device, and the lesser findings weren’t double-checked.

Where it is now

Salient is in Google Play internal testing and on TestFlight, and salient.day is live. There are no usage metrics, by design. It will be a one-time purchase.

What I’d do differently

  • Build the test set on day one. Two days of prompt changes went in unmeasured.
  • Test on the weakest phone earlier. The memory numbers come from a Mac, and the audit agents never ran anything on a device.
  • Check the companion’s job before designing it. Its “where to start” hint is on hold. The model called 24 of 30 real tasks “already one step”, and I learned that only after building it.

Open to work

Hiring a designer who can also run the backlog?

I’m looking for a senior or founding product designer role, remote across the EU — owning a problem from research to release.

B2B contract · available in 2–4 weeks · no visa sponsorship needed (EU)

Kraków · CET/CEST