ETA·BRAINa Katyella project

ETA Brain · How it’s made

How ETA Brain finds an answer

Ask a question and you get back an answer assembled from half a dozen podcast conversations, every claim carrying a receipt. This page is the machinery in between, in two parts: a tour of the search pipeline as it runs today, then the journey to it and the well-regarded ideas that didn’t survive the measuring.

July 2026 · Built on the Acquiring Minds archive · 121 episodes, ~187 hours of interviews as of today

Part one

The search pipeline today

This part is the tour. Every station in it is live on this site right now.

01The big pictureYour question’s journey

The best way to think about this is as an assembly line. Quality control tidies up the incoming question without changing what it’s asking. Then the work splits between two librarian scouts running in parallel: one hunting exact wording, like an index, the other hunting by topic, more like a map. They come back with a heap of research that might apply. A sorter works through the pile and keeps what it knows the writer will find useful. On the bigger questions, it also checks for gaps in the research. If it finds any, it sends questions back to the librarians for more citations. Finally it lands at copy, where the writer drafts the article from what survived and pins every claim to the exact moment in the episode where somebody said it.

your questionquality controlquery scrub · Haikuthe index scoutfull-text search (FTS)the map scoutvector search (embeddings)both comb 48,577 shortmoments of conversationthe pile~50 · RRF fusionthe sorterLLM reranker · picks 12big question? the gap check reads the sorter’s picksand sends the scouts back out, up to twice, before writing startsthe writercomposer · Sonnetthe filing cabinetanswer cache · repeats are free
swipe to see the full figure →
The assembly line. Violet stations are AI doing judgment work. Slate stations are plain machinery. A fresh factual answer puts its first words up in about 7 seconds, finishes near 11, and costs about 3¢. A researched deep-dive starts writing around 20 seconds, finishes near 29, and costs about 6¢. A repeat of either is instant and free. If an AI station fails mid-answer, the question still gets through on a simpler path. The whole line runs under a hard daily budget.

That’s the line, end to end. The rest of this part walks it slower: the paperwork between stations, one real question going through, the one station that isn’t a language model, and the loading dock where episodes become the archive in the first place.

Reference · The paperwork

The assembly line is the metaphor. What follows is the paperwork. Each card shows the exact data one station hands the next, in the shape the code actually passes it. Skim if you like: the tour picks back up below.

What actually travels between the stations: the data, stage by stage

question: "Whats a good way to negotiate a personal gaurantee"
plain text, capped at 500 characters
quality controlquery scrub · claude-haiku
scrubbed: "What is a good way to negotiate a personal guarantee?"
the ORIGINAL is kept too: the writer answers your words,
not our cleaned-up version
the two librarian scoutsFTS tsquery + embedding
tsquery: negotiate & personal & guarantee ← for the index scout
vector(384): [0.021, -0.118, 0.054, …] ← a 384-number "meaning
fingerprint" for the map scout
the pileRRF fusion, no gate
candidates[~50]: { segmentId, episodeId, startS, endS,
text, score, … }
both scouts' lists merged: score = Σ 1/(60 + rank)
reciprocal-rank fusion (RRF): the higher a passage ranks on
either scout's list, the more points it collects
the sorterLLM reranker · claude-haiku
{ "ids": [7, 23, 41, …] } ← just 12 excerpt numbers, in order
structured JSON: what it knows the writer will find useful
the research packetfull snippet + surrounding conversation
evidence[12]: full snippet text + surrounding conversation,
each tagged for citation: [E21@23:06]
(gap-check follow-ups can grow this to evidence[20])
the writercomposer · claude-sonnet
AskResponse: { answer: "…negotiate the size down [E21@23:06]…",
citations: [{episodeId: 21, mmss: "23:06", url: "/episode/21?t=1386"}],
followUps: [2-3 suggested next questions, drafted in the same pass],
sources: [12 short excerpts], cached: false }
the filing cabinetanswer cache · Postgres kv_store
key: sha256("whats a good way to negotiate a personal gaurantee")
← your ORIGINAL words, lowercased, typo and all
value: the whole AskResponse · next asker gets it instantly
swipe to see the full figure →
The pipeline as data. Dark pills are transformations, labeled with their technical names. Cards are the exact shape of what gets handed on. The payload contracts as judgment gets applied to it, from 500 characters to 50 rows to 12 ids, then expands again into a full cited answer. (The gap check’s detour is omitted for legibility.)

02One real questionWatch a real question go through

Everything above is the machine in the abstract, and diagrams have a way of flattering their subject. So here’s one real question, asked on 2026-07-06 and captured station by station exactly as the pipeline saw it: the ranked pile, the sorter’s cuts, the finished answer. Nothing below is mocked up.

The question

What are some ways buyers have softened or limited the personal guarantee when financing a deal?

Quality control read it and changed nothing. No typos to fix, so the question passed through as written.

The pile: top of 50 candidates, as ranked by the two scouts

#1 · Good Bones: Saving a $3m Business in Decline✗ cut
[E4@58:02] · fused score 0.0315 · Himmat Singh: So I was, and pointing that out to my partner and to the operator that look, I see a slowdown in your business for s…
#2 · No SBA Loan and $75k Out of Pocket✓ kept by the sorter
[E18@1:01:53] · fused score 0.0303 · Megan McGee: So I paid $75,000 out of pocket and the rest of it was seller equity rollover and a seller note. And I also had some …
#3 · How to De-Risk the Personal Guarantee✗ cut
[E21@57:28] · fused score 0.0282 · Ryan Conner: But if you've got our policy in place and the actual lenders coming after you for your personal assets in that period…
#4 · How to De-Risk the Personal Guarantee✗ cut
[E21@27:11] · fused score 0.028 · Brendan Burdette: And we're really encouraged to help people out in this space because it's such an incredible thing people can do…
#5 · How Business-Buyer Fit Led to 2.2x and No PG✗ cut
[E9@57:52] · fused score 0.0167 · Will Smith: We still haven't heard about the structure. We've heard about the purchase price, but not the structure of the transac…
#6 · The Origin Story of a Compounder ($80m and Counting)✗ cut
[E22@1:22:55] · fused score 0.0164 · Zach Cooper-Vastola: You know, there's no shortage of brokers and people representing million dollar EBITDA businesses that are re…
#7 · SBA Acquisition to $9m Cash Exit in 5 Years✗ cut
[E62@1:23:55] · fused score 0.0161 · Andy Rougeot: So cutting off the tail ends of the amazing outcomes because I think you miss those sadly if you're buying secondari…
#8 · Acquisition Unlock: €210m in 5 Years✓ kept by the sorter
[E80@48:17] · fused score 0.0159 · Nick Keegan: So pretty much everybody we spoke to and this is actually, this is something Shane introduced, which looking back, I …

…plus 42 more candidates below these. The sorter kept 7 in total. 5 of them came from further down the pile than what you see here.

The writer’s answer (opening)

Buyers have used several negotiating tactics to soften or limit personal guarantee exposure:

Full answer: 7 citations across 6 different episodes, each one a link to within seconds of the quoted moment.

swipe to see the full figure →
Rank is not relevance. Look at candidate #1: the scouts put it at the very top of the pile, and the sorter cut it anyway, because it sounded related and wasn’t. That single row is the whole argument for LLM reranking over a score threshold. A formula can rank. Judging takes something that actually reads.

03The strange stationHow a machine matches meaning, not words

The index scout is easy to picture. It matches the exact words you typed, the way a book’s index would. The map scout needs explaining, because its map is not a figure of speech. The scout reads a piece of text and turns it into a point in space: an actual list of 384 numbers (the “meaning fingerprint” from the paperwork above), coordinates on a map with 384 axes instead of two. The map is arranged so that distance means difference in meaning. Two sentences that say the same thing in different words land in almost the same spot. Two that share words but not meaning land far apart.

Every snippet in the archive got its coordinates once, at ingest. Your question gets its coordinates the moment you ask. Then the search is pure geometry, and the database runs it with no model at all: find the snippets sitting nearest your question. Nearest on the map means most likely to answer you, even when a passage shares not one word with what you typed. It’s why a search for “letting a manager go” surfaces a guest talking about “having to fire someone.”

your questionsame meaning, different wordsunrelated topics
swipe to see the full figure →
The meaning map. Each dot is a passage placed by what it means (384 real axes, two drawn here). The search is just “which dots are closest to the question’s dot,” which is why the map scout finds the right passage even when the wording doesn’t match.

This is the one station on the line that isn’t a language model, which raises a fair question: why not Claude here too? Because Claude is trained to take text in and write text out. It never hands back coordinates, and its internal sense of a sentence was never tuned so that distance equals meaning. Drawing the map is a separate craft with its own training, done by a different kind of model (a dedicated embedder), in our case a small open-source one we run ourselves. Every other AI station on the line is Claude doing a language job. This one earns its place by turning meaning into geometry the database can search in milliseconds.

04The loading dockHow episodes become the archive

None of this works until the episodes are on the shelves, and that happens at the loading dock, long before anyone asks a question. New episodes arrive from each show’s public feed. If the publisher posts a transcript on the episode page, we take it as is. If not, a Mac mini in the workshop listens to the audio and types the transcript itself, using a speech-to-text model (WhisperX) plus a second model that tells the voices apart (the term of art is diarization). It works through one episode at a time and takes about half an hour each.

Then every voice gets a name and a role, host or guest, so an excerpt can say who said it. The audio is deleted the moment the transcript exists. We never host it: when you want to listen, every link on this site sends you to the publisher. We keep short excerpts, they keep the show.

Next the transcript is cut into small units, 2 to 3 sentences each, stored with a wider stretch of the surrounding talk and a timestamp accurate to the sentence. (Part two explains why that size wins.) Each unit gets its coordinates on the meaning map once, right here at the dock. Finally the finished shelves are pushed, one way, to the hosted database the site reads. The working copy never leaves our own machine, and a timer re-checks the published copy every 15 minutes. So far the dock has processed 121 episodes, about 187 hours of conversation, into the 48,577 searchable moments the scouts comb.

a new episodetranscript postedno transcriptpublisher transcripttaken as isthe Mac mini types itWhisperX · tells voices apartnames and roleshost or guest · audio deletedsmall units2-3 sentences · coordinatesone waythe shelvespushed every 15 min
swipe to see the full figure →
The loading dock, feed to shelves. Violet stations are models at work. Slate is plain machinery. The audio never makes it past the second station, and the site only ever reads the pushed copy on the right.

That’s the tour: the line, the paperwork, one real question, the strange station, and the dock that feeds them all. Everything above is running right now. What’s left is how the line got this way.

Part two

The journey to the current pipeline

You can stop after the tour and have the full picture. This part is the build story: what we launched, what broke, and what the measurements kept.

05The original problemThe inspector who demanded two stamps

The sorter’s station used to belong to an inspector with one rigid rule and no judgment. How he lost the job starts on launch day, when the answers were bad in a way that took real effort to pin down: reasonable questions came back empty while outright gibberish sailed through to a paid answer. The culprit was the inspector, standing between the scouts and the writer (technically a relevance gate on the fused RRF score, the merged points from both scouts’ lists). The rule he enforced: a passage moves on only if both scouts stamp it.

The rule sounds sensible until someone misspells a word. The index scout can only match strings that actually appear in the transcripts, so a typo sends him back empty-handed. Once his stamp is missing, the arithmetic is final: no passage can clear the bar, however confidently the map scout stamps it. One wrong letter shut down the entire question.

Here’s the thing: the inspection failed in the opposite direction too. A rescue clause, written to spare near-miss questions from an empty result, waved lone weak hits through whenever the pile came up short. A short pile is exactly what gibberish produces. So nonsense reached paid AI calls while honest questions starved: strict exactly where it should’ve been lenient, lenient exactly where it should’ve been strict.

a clean questionAatwo stampsthe checkpointanswer ✓“personal gaurantee” · one typoAaone stamp missingnothing passes“The podcasts haven’t covered that.”
swipe to see the full figure →
The two-stamp rule. The checkpoint demanded a combined score no single scout could reach alone, so the moment the index scout struck out, the outcome was already decided. The map scout had found the right passages. They never made it past the checkpoint.

Replacing him took a day. Choosing the replacement took measurements.

06The fixWe didn’t argue. We held an exam

The exam: 44 tough questions (straightforward, big-picture, typo’d, reworded, pure nonsense) for which we knew in advance exactly which moments in which episodes held the answer. Seven retrieval designs sat it, from carefully tuned classical search to the fashionable techniques in the field’s playbook.

The field deserves naming, because a couple of the entries are the going advice for exactly this problem. There was a cross-encoder, a small model trained for the narrow job of scoring one passage against one query. There was HyDE, where the AI dreams up an ideal answer first and the scouts then hunt for real passages that resemble the dream. There was an episode-first design that ranks whole shows before the moments inside them. And there was the humblest entry on the sheet: keep the launch-day inspector, just rewrite his rules to be reasonable.

The winner, by a wide margin, was none of the fashionable techniques: hire a sorter (the term of art is LLM reranking). Let everything plausible into the pile. Then pay a fast AI to actually read all ~50 candidates and keep the twelve it knows the writer will find useful. A sorter that reads understands the intent behind a typo. It spreads its picks across different episodes. And it’s allowed to conclude that nothing in the pile helps, which is how nonsense finally gets turned away. It found 2½× more of the known-right moments than the launch design.

The most instructive loser was the kitchen-sink entry, which stacked all the clever techniques at once and scored worse than the sorter alone. Turns out improvements don’t add. Each stage deposits its own noise into the pile, and every stage downstream has to fight through it.

07The ladderClimbing, one honest step at a time

With the sorter hired, we kept climbing, under one house rule: every new idea takes the same exam, and it keeps its place only if the numbers improve. We call the rule keep-if-wins, and the result gets written down whichever way it goes. Seven ideas have earned a place so far. Four didn’t, among them some of the field’s favorites and one we had asked for ourselves. The three newest wins went up in July 2026, after we surveyed what other AI search systems were doing and ran the promising ideas up the same ladder.

Quality control: query scrubbing. Fix the question’s spelling before anyone searches, so every station downstream inherits a clean query. One tiny Haiku call, $0.0003: typo’d questions went from working 57% of the time to 71%. (Fixing the cause upstream beats compensating downstream.)
The search box got the reasonable inspector. The type-as-you-go search box, it turned out, had been quietly running the launch-day rules all along. A free fix and a night-and-day difference on that page: recall 15→27, typo survival 0→43%. (The exam's rewritten-rules entrant: it lost the main event to the sorter, then won the one page where a paid sorter is too slow for type-as-you-go.)
Smaller pieces: sentence-window re-chunking. The library’s “moments” used to average 79 seconds, and some ran to seven minutes: far too coarse to match or cite precisely. We re-cut the corpus into 48,577 snippets of 2–3 sentences, each carrying its surrounding conversation, and it proved the best single improvement of the project: how close to the top of the pile the right moment lands (a score called MRR) jumped 43→47. The diagram below shows why.
The gap check: a deeper research pass. A triage step sorts each question while the scouts are already out, so simple questions never wait on the sorting. Factual ones actually got faster. The bigger ones earn a deeper pass (the term of art: an agentic tier). The gap check reads the sorter’s picks, spots what’s missing, and sends the scouts back out, up to twice, before writing starts. Judged first impressions rose two-thirds of a point out of ten, source weaving half a point, and groundedness (how much of the answer is actually backed by the tape) went up, not down. Answers now draw on ~7.4 distinct conversations. (Costs ~6¢ and runs 21–30 seconds, spent only on the bigger questions.)
Ask next: the writer suggests follow-up questions. While drafting the answer, the writer now also jots down two or three questions worth asking next, grounded in the research it just read. They appear as chips under the answer. On the exam, answer quality rose 8.01→8.43 with groundedness holding, and 88% of the suggestions were judged answerable from the archive, against an 80% bar. (Drafted in the same pass as the answer, so they cost nothing extra.)
An audit of the receipts: support verification. A checker that reads every claim in an answer and asks whether the cited excerpts actually back it up, built ready to trigger a correction step if too many failed. The line passed an audit it didn’t know it was taking: 116 of 116 claims held, on both exam answers and fresh production ones. Nothing new shipped, because nothing needed fixing. (The term of art: entailment checking. The checker stays on the line and reruns weekly.)
Streaming: the writer now types in front of you. Answers used to arrive in one piece after a blank wait. Now the sorter’s picks appear the moment they’re chosen, and the writer’s words show up as they’re written: first text at 6.9 seconds on a fresh factual question, where the old path showed nothing until 12.3. The final answer is identical either way. If the stream breaks, the one-piece path still runs. (The prediction it broke is written up below.)
A fancier map scout: bigger embedding models. Three bigger, newer meaning-fingerprint models for the hunting-by-topic search: two measured, one disqualified before the exam as too big to deploy. Best case: +3.7 recall points, under the +5 bar we’d set, and it arrived with a multi-episode regression. The other candidate was outright worse. The fingerprint model was never the bottleneck. Money saved.
Adding context notes: contextual retrieval. A celebrated 2024 technique, the field’s favorite recipe: label each snippet with what it’s about before indexing. It found slightly more (recall +1.8) and ranked noticeably worse (MRR −5.4), because the generated labels made everything in an episode sound alike. Beaten by our actual data.
Using the topic tags. Boost passages whose topic matches the question. We swept 48 different configurations. Not one helped. The tags are fine for browsing, useless for ranking.
Scholarly citations: Chicago-style verbose attribution. Our own request: name every episode and speaker, quote verbatim, and measure it like everything else. Groundedness held, but answers fell from 7.4 to 5.1 distinct conversations, and a follow-up adjustment meant to restore the breadth only made things worse. We compared samples side by side and kept the leaner voice. Even the changes we asked for ourselves obey keep-if-wins.

Two more additions from that run take no exam, because they change no answer. The line now keeps a log of what real visitors ask, so future exams can use real questions instead of only ours. And the answer pages now carry the markup that lets search engines and AI assistants read them properly.

The streaming row hides the run’s best lesson. We predicted first words on screen in 2 to 3 seconds, because we assumed the writing was the slow part. The measurement said no. The stations before the writer, quality control through the sorter, set a floor of about 7 seconds, and streaming only removes the wait for the writing itself. This page argues for measuring instead of trusting a good story. This time the good story was ours.

Why smaller pieces mattered so much

before: search a 7-minute ramble, cite its very beginningeverything blurs together · “listen from 46:27” might mean 51:40citation off by ±1 minafter: match a snippet, hand the writer its surroundingsthe match…but the sorter and the writer still see the whole exchangecitation lands within ~13s
swipe to see the full figure →
Match small, read wide. One structural change improved finding, ranking, typo survival, and citation precision at once, which no model upgrade we measured came close to doing. The sleeper win of the whole project.

One thread through the whole climb: questions with typos that still get answered

launch day
0 in 10
+ the sorter57%
~6 in 10
+ quality control71%
7 in 10
+ smaller piecestoday · 86%
~9 in 10
05 in 1010 in 10
From never to nearly always. Of the three steps on the chart, only one touches spelling at all, and even that one is a model reading for intent, not a dictionary. The rest of the gain came from better judgment and smaller pieces.

08What we learnedJudgment beats scorecards

Every idea that earned its place improved the question (quality control), the pieces (smaller snippets), or the judgment applied to what was found (the sorter, the gap check). Every idea that lost tried instead to make the library’s description of itself smarter (the representation, in the jargon): bigger meaning-fingerprints, context notes, tag boosts. Judgment won every match-up it entered, and even our own pet idea, the scholarly citations, lost by the same rule: measured, then cut. Scoreboard, not vibes.

The net effect: a typo’d question went from never working to working nine times in ten; answers that once leaned on a single episode now weave together about seven; citations that pointed within a minute of the quote now land within thirteen seconds, and an audit of every claim against its cited excerpts came back 116 for 116; and judges score the answers 7.2 → 8.4 out of ten since the new engine shipped. If someone has already asked your question, the answer arrives instantly. The line built it the first time.

So: ask something big and slightly misspelled. The line is running →


Built July 2026 by Katyella. Every number on this page comes from a frozen, repeatable test (44 golden questions with verified answer locations, plus a 14-question answer eval judged twice per run) and traces to the project’s evaluation records. The full ledger, launch config to today: right moments reaching the writer 15 → 39 of 100; typo survival 0 → 86%; nonsense rejection 0 → 100%; episodes woven into each answer 1.3 → 7.5; citation precision ±1 minute → ±13 seconds; cited claims backed by their excerpts 116 of 116; judged answer quality, measured from the day the new engine shipped, 7.22 → 8.43 of 10. A fresh factual question costs about 3¢ and starts answering in about 7 seconds; a researched big-picture one costs about 6¢ and starts near 20; repeats are free.

Build your ownThe same line, over your archive

Everything above is a recipe, and recipes are for cooking. If you have an archive of your own, podcasts, documentation, support tickets, meeting notes, research calls, the same assembly line will run over it. Below is this page as a prompt. Paste it into an AI coding agent like Claude Code, fill in the three brackets, and it will build the line stage by stage.

One honest warning before you copy it. The prompt buys you the architecture and the measuring discipline. It does not buy the weeks of measured iteration that tuned this line to this archive. Your data will teach you different lessons. The scoreboard is how you learn them.

I want you to build me a search and question-answering tool over my own archive, modeled on a measured build (etabrain.vercel.app/how-its-made). Work in stages and show me each one running before you move to the next.

MY ARCHIVE
- Sources: [DESCRIBE YOUR N DATA SOURCES, e.g. two podcasts, a docs site, and a folder of support tickets]
- Where the text lives: [FILES / URLS / DATABASE / RSS FEEDS]
- What people will ask: [THE KINDS OF QUESTIONS, WITH 3 EXAMPLES]

STAGE 1, INGEST. Normalize every source into one table of text units. Each unit needs a stable id, the source it came from, and a deep-linkable location (a timestamp for audio, a page for PDFs, a URL for web pages). Store text, ids, and locations; leave the original files where they are.

STAGE 2, CHUNKS. Cut the text into small retrieval units of 2 to 3 sentences. Store a wider context window next to each one (roughly a paragraph, or 45 seconds of talk). Search matches the small unit; every later reading stage gets the wide window. Do not chunk by fixed character counts.

STAGE 3, RETRIEVAL. Build two searchers that run in parallel: keyword full-text search (Postgres FTS or SQLite FTS5) and vector search with a small local embedding model (bge-small or similar). Fuse both lists with reciprocal rank fusion into roughly 50 candidates. Do not buy a bigger embedding model; it is rarely the bottleneck.

STAGE 4, THE SORTER. This stage matters most. Send all 50 candidates with the question to a small, cheap LLM that picks and orders the 12 or so a writer would find useful. It must tolerate typos, prefer diverse sources, and return an empty list when the archive cannot answer the question. An empty list means the tool says "I don't know" instead of guessing.

STAGE 5, THE WRITER. A stronger LLM drafts the answer from the picks, citing each claim with the unit id and location so every citation becomes a deep link. In code, not in the prompt, strip any citation that does not resolve to real evidence. Cache finished answers by a hash of the question. Add per-IP rate limiting and a hard daily spend cap that fails closed.

STAGE 6, OPTIONAL. For big thematic questions, add a router that sends them to a research pass: it reads the picks, notices gaps, and runs up to two more retrieval rounds before the writer starts.

THE DISCIPLINE, which matters more than any stage above. Before changing anything, write about 40 golden questions whose answers you can point to in the archive as verbatim quotes, plus a judged answer eval of 10 to 15 questions. Measure the baseline. Then change one thing at a time, re-measure, keep only what wins, and record every result in an append-only scoreboard file. Expect popular techniques to lose to this process. Keep them out unless they pay.

If you’d rather have it built with you, that’s the business we’re in.

Every station you just read about is live on this site. Ask something and you can watch the line work: the sorter’s picks land first, then the writer starts typing, and when it finishes it suggests what to ask next. Typos welcome. You know who’s catching them now.

Ask it something →