SP//LOG

Raya: four days

I shipped a wake word with 96% recall that had never once fired. Four days of building a local agent, and every failure looked exactly like success from the outside.

A Raya terminal session. Wake fires at 0.91, the daemon captures 14.1 seconds of speech, transcribes a question about the raya-ai project, and replies that the repository has no remote and all 43 commits are local only.
Published
Reading
15 min
Difficulty
Intermediate
Status
Experimental

I shipped a wake word with 96% recall. It had never fired. Not once, not in any test, not ever.

Both of those things were true at the same time, and it took me most of a day to understand how.

That was day four. Here's the whole thing from the start.

What I was building

A local agent that reads all my projects, tells me what's rotting, writes patches, and answers when I talk to it. Runs on my laptop. Nothing about it is impressive on paper.

The starting point was a YouTube video about building a personal JARVIS. Six layers, spelling out J-A-R-V-I-S: job runner, agent harness, reasoning model, virtual connections, instructions and memory, skills and schedules.

The framing is good. The opening argument is better: the dashboard isn't JARVIS. The glowing circle does nothing. The system underneath is the thing.

But its target didn't fit me at all. It assumes a business with operations. Check the CRM, triage the inbox, pull the KPIs, build a morning briefing. I don't have a CRM. I have 29 project directories and a habit of leaving things half finished.

So I kept the six layers and threw out the target. New one: something that holds the state I can't, remembers decisions so I stop re-deciding them, and does the mechanical work between the thinking.

18 Aug: nine roles, and the first one lied to me

I built nine roles in a day.

A daily research digest. A sweep across every repo saying what moved and what's rotting. A design audit that checks my UI against my own anti-slop rules. Something that hands a coding task to an agent and gives me back a patch. A verifier. An archivist that harvests facts into memory. A sparring partner that argues with an idea before I build it. A researcher. A watchman.

Nine roles in a day. Hold that thought, because everything that goes wrong later goes wrong for that reason.

The first report was wrong and I nearly missed it

First run of the project-state sweep. It came back with a ranked list of what needed attention, and it looked great.

Two things in it were wrong.

It called devlog-website gitignore noise. It wasn't. Both dirty files were real source edits I'd made and forgotten about.

The worse one: it ranked sheetforge as my second biggest pile of unsaved work off a 4,766 line diff. I went to look.

4,501 of those lines were package-lock.json.

The actual work was 265 lines. The report wasn't lying, exactly. It counted the wrong thing and then sorted by it, which is somehow worse than lying, because everything downstream inherits the mistake and looks confident doing it.

Fixed both in the SOP. Lockfiles out of insertion counts, generated noise separated from real edits. Re-ran it. prism-ui jumped from fourth to second. sheetforge fell to third but stayed high priority for an entirely different reason: no remote. One disk failure and the whole project is gone.

I appended the correction to the original report instead of rewriting it. A report that quietly edits itself to look right isn't one you can trust next week.

Nine permission rules doing absolutely nothing

Late that night I found that every path permission I'd written was the wrong type. Write(digests/**), Write(state/**), nine of them.

Turns out Write(path) only matches the tool name. Edit(path) is what actually gates file paths, and it covers all the editing tools.

Nine rules. Silently inert. No error, no warning, nothing.

And I only found out by accident, because I'd turned on workspace trust and the allowlist finally started loading. Before that it had been ignoring all 21 entries anyway.

Which means my permission config had never once been enforced. Not on any run. Not from the beginning.

That should have worried me more than it did at the time.

The rule I wrote down, and then broke

At some point I stopped adding roles and wrote down a rule instead, because I could feel the roster getting away from me.

Two axes. Does this role ingest untrusted input, meaning the open web, a feed, a comment thread. Does it have hands, meaning can it run commands.

Three of the four combinations are fine:

  • untrusted input, no hands → a scout
  • trusted input, no hands → a thinker
  • trusted input, with hands → an operator, safe as long as git can undo it

Classify before you write. It sounds obvious on paper.

Remember this. It comes back on day three, and I'm the one who breaks it.

19 Aug: giving it hands, and finding out what they could reach

Seventy-six cents to do nothing

I added cost tracking because the timers were live and I'd just hit a credit limit.

It was the project-state sweep hitting its own guard, seeing today's file already existed, and exiting. Seventy-six cents to check whether a file exists.

Nearly all of it is cache creation. About 72k tokens of context loaded before the skill does anything at all. Most of that is my global config, which includes a 700 line design law that the digest pays for every single morning and never reads one word of.

That was annoying but fine. The next thing was not fine.

The morning it reported a quiet day

First real unattended run fired at 08:00. I got a digest with one item in it.

Exit code 0. File written. Looked like a slow news day.

It had swept 1 source out of 13.

Twenty-one permission denials. Every curl, every fetch, opencli, all blocked. I'd written the SOP with the exact commands the skill needs, then built an allowlist containing only git reads and file paths. The network access its entire job depends on had never been granted.

That exposed a second gap. The runner checked "did a file appear." A file appeared. So it reported success.

It now reads permission_denials out of the run JSON, and a blocked run exits 4 as DEGRADED. Because a run can produce its artifact and still have done almost nothing, and those two things look the same from outside.

The digest stopped costing money

Small win in the middle of all this.

The digest is fetch, then judge. That's one LLM call. It doesn't need an agent loop at all.

Rewrote it as a shell script that does the fetching directly and sends the pile to a free model on OpenRouter for ranking.

$1.76 a run down to $0.00. The daily research digest now costs nothing to run, forever.

delegate, and the command I handed over myself

delegate is the first role that writes code. Give it a bounded task, it runs a coding agent in an isolated git worktree, verifies the result, and writes a patch to pending/.

It never applies anything. That's not caution for its own sake. Every repo I'd point this at already has uncommitted work in it, personal-portfolio has 27 dirty files. If the agent edits into that tree, its changes and mine interleave and neither git diff nor git checkout can pull them apart again.

That's not data loss. It's review loss, which is worse, because you don't notice it happening.

It took three rounds of review before it was safe. The rounds are the story.

Round one. Review found the agent was running with my own repo as its working directory. My files were in edit scope. Fixed by pointing it at the worktree.

Round two. Review pointed out that this had moved the agent out from under my deny list completely, because Claude Code resolves project settings from the working directory. It was now running under whatever settings happened to live in the target repo. My fix: pin my own settings file explicitly.

Felt good about that one. It was the wrong fix.

Round three. Running this from inside a worktree:

run from inside the worktree
git branch -D feature-work

It deleted a branch in the real repository.

Worktrees share the parent's ref store, so writes from inside one land in the actual repo. My deny list didn't cover git branch. I'd enumerated five things I didn't want and that wasn't one of them.

And my own allowlist contained Bash(git branch:*).

So the fix I shipped in round two handed the agent the exact command used to destroy a branch in the repository it was supposed to be isolated from. I'd made it worse while believing I'd made it safer, and the only reason I know is that something adversarial went looking.

Fixed properly with a dedicated settings file that denies Bash(git:*) wholesale, plus --setting-sources user so the target's own config never loads.

I'd also written a comment in there claiming the source tree was "a place it cannot reach."

That's false, and I changed it to say so. node -e shells out to git. cp and mv take absolute paths. Bash(cat:*) walks straight past a Read(./.env) deny. It's a policy boundary, not a sandbox. Writing the honest version in the file matters more than the deny list does, because the next person to widen that allowlist will read the comment, not the rules.

Silence that depended on the model being polite

watchman is the only role that decides whether to interrupt me. Silence is the normal outcome, because a watchman that talks every day gets muted, and then the one channel that was supposed to matter is dead.

I built it so the model replies with the literal word QUIET when nothing qualifies, and the runner suppresses that.

Review pointed out this could never hold. The parser collapses newlines to spaces, so a fenced ```QUIET``` block misses the check. So does - QUIET. So does **QUIET**. So does any run where the model writes "nothing meets the bar" instead of the exact word.

Twice a day, forever, until I muted it.

Inverted it. Speaking now requires an explicit append to a log file, and the notification is built from whatever got appended. Silence is the default and no amount of model prose can override it.

21 Aug: the review said no

Ran a whole-branch review before merging. Verdict came back: NOT SAFE TO MERGE. Three criticals.

One. Every script sources .env with set -a, which exports my OpenRouter key and bot token into the environment. claude -p inherits that environment. My allowlist permits Bash(curl:*).

So one allowed command takes both keys, trips no deny rule, and logs no denial. The Read(./.env) deny I'd been relying on was theatre. The secret was already in the process before any rule got consulted.

Two. Here's the callback I promised.

The trust boundary I wrote down on day one, broken in my own config. The same allowlist granted WebFetch, curl, agent-reach and opencli alongside Bash(bin/raya-delegate:*).

Ingest and hands. Same file. The exact quadrant I said never gets built.

A fetched page with the right text in it could have started a coding agent inside a real repo. I wrote the rule. I put it in the plan. I wrote a section about it. Then I broke it in the one place where it's actually enforced, and didn't notice for two days.

Writing a rule down is not the same as implementing it. The document felt like the work.

Three. There was no spend cap anywhere that runs. The only one lived in the Telegram bot, which I never installed. Meanwhile my .env.example and the watchman SOP both told me a $5/day hard cap was in force.

All three fixed. The cap now lives in one place, fails closed if it can't read the ledger, and gets checked before anything reaches a model.

22 Aug: 96% on a model that never fired once

Back to where this started.

I wanted a wake word. Not "hey jarvis", which comes pretrained. "hey raya."

Generated 3,500 synthetic samples with piper across five spellings, because the TTS couldn't agree with itself on how to say "raya". Transcribing its own output gave me "hey Rhett", "hey Rhea" and "hey Raya". Trained against 391 hours of negative audio on the laptop GPU. Six minutes.

96% recall. Zero false wakes per hour.

Shipped it. It never fired.

Two bugs underneath.

Every training clip had the phrase starting at t=0 with silence padded after it. So the model learned "phrase at frame zero, then silence." In streaming, the phrase lands anywhere in the window. Randomising placement took it from 5.9% to 96%.

And no augmentation at all. Clean TTS isn't a microphone in a room. Added gain variation, noise at varying SNR, low frequency rumble, speed jitter.

Retrained. 96% on held-out synthetic.

Then I tested it against a recording of my actual voice.

0.007. The threshold is 0.2.

Still completely dead.

That's the domain gap, and no amount of augmentation fakes it. Synthetic speech and a real person through a real microphone are different distributions, and I'd trained entirely on one of them.

One more thing fell out of that. I ran whisper over my own recordings to check them, and it transcribed them as "hey dia" and "hey dio".

Nothing like any of the five spellings I'd guessed.

I'd spent 3,500 samples betting on pronunciations of my own daughter's name that I don't use.

And then it still didn't work

Wake fired at 0.98. Then:

raya session
captured 0.4sheard: [BLANK_AUDIO]

The recorder stopped instantly. Every single time.

My calibration samples the room at startup and sets a silence threshold. It returned 0.3426.

My speech peaks around 0.19.

Every word I said was below the threshold that exists to detect speech. The formula was median times three with no ceiling, so one noisy moment at startup poisons the entire session.

Fixed two ways. The gate is now relative to the loudest audio in the current utterance, so speech can't be quieter than the thing meant to catch it. And calibration is capped at 0.04, because nothing above that is plausible for a desk mic.

Then it worked.

raya session
wake (0.98)captured 5.0s, peak 0.379heard: Can you tell me the state of my projects?reply: You have 35 project directories with seven rotting. personal-portfolioleads at fifty days and fifteen hundred uncommitted changes, while sheetforgehas no remote at all.

All of that is correct. It read my actual repos and told me the truth about them, out loud, on my own machine, for zero rupees.

First round trip was 21 seconds. Capture ate 12 of them, because my room's noise floor sat right at the gate. The model ate 28 in a later run. Tightened the gate, swapped a 550B model for a 120B, gave whisper more threads. Down to about 12 seconds now, and most of what's left is whisper.

Where it is

Nine roles on timers plus a filesystem watcher. It reads 29 repos, writes patches in isolated worktrees, keeps a memory, logs every rupee it spends and refuses to go past a cap. Daily running cost is zero, because the parts that don't need an expensive model don't use one.

I say "hey raya" and it answers out loud. All local. Nothing leaves the machine except transcribed text.

I still can't reach it from my phone. The Telegram bridge is built and tested and I haven't turned it on, which is a choice rather than a bug.

What I'd tell you if you're building one of these

Measure the path that actually runs. Not a convenient proxy for it. 96% recall and 0.059 were the same model on the same audio, measured two ways, and only one of those measurements described the system I had.

Writing a rule down isn't implementing it. I wrote the trust boundary on day one, put it in the plan, and violated it in the config two days later. The document felt like the work. It wasn't.

A permission rule that doesn't apply produces no error. Nine of mine did nothing for two days. Nothing failed. Nothing warned. The only signal was a message I saw by accident.

Anything that reports success needs something separate checking whether it did the work. A digest reporting a quiet day it never checked. A run exiting 0 having been blocked 21 times. A cap documented in two files and implemented in none.

esc