On retrieval
RAG Is Plumbing. I Wanted to Build a House.
Everyone talks about RAG as chunks, embeddings and re-ranking. Ask the Ancients is what happened when I made it the foundation of a paid product, with four different ways to explore the same texts.
Most writing about RAG is about plumbing. Chunk sizes. Embedding models. Cosine similarity. Re-ranking. I have written some of it myself. All of it matters, and you still cannot sell anyone a vector index.
The more interesting question is what you build on top. What does a product look like when retrieval is the foundation, and every feature has to stand on it?
Ask the Ancients is my answer. One retrieval pipeline pulls the right passages out of a philosopher’s own writing. On top of it sit four different ways to explore the same texts: a one-to-one Chat, an Agora that puts two philosophers side by side, a Council where three of them review each other’s answers, and a Quorum that is an openly synthetic group discussion. Same pipeline underneath, four very different experiences on top.
It started as a weekend RAG demo in March. In September I rebuilt it as a paid product, which meant I suddenly had opinions about credit ledgers and refund rules. This is what that took.
What I built and why
Ask a chatbot what Marcus Aurelius thought about anger and you get a confident paragraph. Ask where he said it and the paragraph gets vague. The model has read about him. It has never read him.
Most philosophy apps land in one of two places. Trivia, or roleplay: a voice that sounds suitably ancient and cites nothing.
I wanted a third option. Every reply is built from passages retrieved out of one philosopher’s published texts. No web search, no general knowledge, no blending. Seneca cannot borrow from the Buddha.

Every claim links to its passage, with the book, section, translation and year.
The people I built it for arrive with a question instead of a reading list. Someone deciding whether to quit a job, who wants to know what the Stoics said about control. A student comparing Advaita and Vishishtadvaita on the self. A reader who has heard of the Dhammapada and has never opened it.
The goal is to make people more curious about philosophy. There are plenty of therapy apps, journaling apps and Stoicism apps out there, and some of them are good. I did not want to build another one. The citations are there to send you back to the text.
Four ways to ask
Chat. One philosopher, or let the Ancients pick one from your question.
Agora. Two philosophers on the same question, and a neutral note on where they part ways.

Seneca says anger is never justified. Aristotle says yes, in the right measure. The note underneath explains why they split.
Council. I stole this one from my own workflow. In agentic development, an LLM council has several models answer independently, review each other’s work, and then revise. Council does the same with philosophers. Three of them answer on their own. In a full Council, each then responds to the other two with one of three moves: concede, hold or sharpen. A deep Council adds a round where everyone reconsiders their own answer. Then a judge writes up where they converge, where they differ, and what is still unresolved.
The judge is instructed never to rank them or pick a winner. Deciding who is right is your job.
Quorum. A synthetic group discussion among several philosophers. It wears an “AI-generated dialogue” label on its sleeve, because that is exactly what it is.
The (AI) economics of it all
Every answer is a call to a language model, and language models do not accept payment in wisdom.
For a single chat, a generous free tier is easy. The fun parts of the product are the ones that bring several philosophers into the room, and a full Council makes seven model calls to answer one question. A one-person product that gives that away has a short and very philosophical life.
So Ask the Ancients is free to start and paid to go deeper. It is also my chance to run the economics of an AI-native product end to end, where every number is mine and so is every mistake.

Every plan answers from the same texts with the same passages attached. Paying buys depth, never better grounding.
The price table decided the chunk size
The credit plans assumed a chat turn would cost about half a cent. Then I measured one. It cost nearly double. The culprit was the passages: each chunk ran 2,400 to 5,100 characters, so a paid answer carried around 15,000 tokens of text, most of it about something else.
I could have raised prices or shrunk the plans. I cut the text instead, re-chunking the corpus into 2,560 passages of about 1,200 characters, split on sentence boundaries. The average prompt halved. Retrieval recall on my test set jumped from 71% to 88% the same day, because a smaller chunk is about one thing. I would like to say I planned that second part.
Credits come from a cost model
A full Council makes seven model calls. An Agora makes two and a summary. Charging the same for both would be simple and wrong. Every mode’s credit cost is computed from measured token usage and the current model price, so when I swap a model or change the chunking, every price updates itself. I do not have to remember to. Failed runs refund automatically. Nobody should pay for an answer that never arrived.
All of this lives in a versioned manifest, outside the code: which model plays which role, how many credits each plan gets, what each mode costs, the daily allowances, and the welcome grant for new accounts. When I want to rebalance a plan, run a promotion or hand out trial credits, I publish a new version as an admin. Nothing gets redeployed, and every trace and eval run records the version it ran under. Staging is on version twenty-five, which tells you how often I changed my mind.
The frugality rule
One rule keeps the costs honest: no new paid service until there are paying customers. The whole rebuild, evals included, has run on a five-dollar Cloudflare plan and about forty dollars of inference.
How it got built
I ran Ask the Ancients like a small product team, and I led it. I wrote the requirements, designed the system and the UX, set up the infrastructure, worked out the economics and made the calls. Claude Code and Codex were the engineers. They took my specs and wrote code against them.
Every feature started as a written spec that I approved. A plan broke it into tasks. One model implemented each task, a second model reviewed it, and the two argued until the review passed. Then it came to me. The agents could commit as often as they liked, but nothing was pushed without my sign-off, and since a push deploys staging, my sign-off was also the only way into CI. Weeks of that loop produced more than 200 commits and 935 unit tests.
Measuring it, because vibes are not a metric
Evals
A golden set is only as good as its questions, and my own guesses were not a sample. So I sent a handful of research sub-agents through Reddit and Twitter to collect the philosophical questions people actually put to AI. From that research I wrote 49 questions across the Stoic, Indian and Western texts, each paired with the passages a good answer should draw on. Retrieval, generation and the multi-voice modes each have their own eval. Every generation run is logged in Langfuse as an experiment on that dataset, so a prompt change or a model swap gets compared run against run instead of by feel.
Where it stands today:
- Retrieval recall on the golden set went from 71% to 93%.
- Across 30 generation eval runs, the number of invented citations is zero.
- An LLM judge finds 84 to 89% of cited claims supported by their passage.
- A standard answer costs about $0.0013 in model fees.
The evals also caught things I would never have found by reading answers:
- citation markers written in five different bracket styles.
- a reasoning model that spent its entire budget thinking and returned nothing.
- a Council review cut off after six words.
CI
GitHub Actions runs type checks, lint, the unit tests and a secret scan in parallel on every push. Only when all four pass does it deploy staging, and then a Playwright suite runs against the live staging site. A retrieval gate fails the build if recall drops below 80%. Twice a week, on Monday and Thursday mornings, the full eval suite runs against staging on its own. It costs about twenty cents a run, which is the cheapest insurance I have ever bought.
Observability
Every request gets an id at the door, and that id follows it everywhere. The same id names the Langfuse trace, tags the call in Cloudflare’s AI Gateway and appears in the Worker logs that Sentry collects. When an answer looks wrong, I can pull up the prompt, the passages, the model and the cost for that exact request.
- Langfuse holds the traces, and scores every answer: how much of it is grounded, whether any citation was invented, whether the persona slipped, and how the reader rated it.
- Sentry catches errors from the browser and the Worker, with message content stripped out before anything leaves.
- PostHog tracks product events, from a new thread to an opened citation, with user input and philosopher text masked.
- Better Stack checks production every three minutes, including a health endpoint that does a real database read, and runs a public status page.
- AI Gateway enforces a daily spend limit, so a runaway loop stops at three dollars. My rent is safe.
The part I am unreasonably proud of
To most people, a philosopher card that opens smoothly is not worth a section in a write-up. I once spent an entire day aligning images on a single page, so I am not most people. Design matters to me more than it probably should, and this is the part of Ask the Ancients I am proudest of. Nobody has to notice. I would still like them to.
Open a card and it grows in place into a full profile: the thinker’s works, a line from their writing, and a way to ask them something. The rest of the grid gives up space and reflows around it. On the Philosophers page, a timeline above the grid moves to that thinker’s century as you go.
An agent will happily build a card that expands. Getting it to expand well is a different job. The first version opened and closed almost instantly: a card went from 203 by 280 pixels to 432 by 500 in eighty milliseconds, which is less an animation than a teleport.
I spent half an hour writing out exactly how the motion should feel. Eleven rounds of feedback followed, six of them in a single day. Now the whole grid resizes over 500 milliseconds on a custom easing curve. The chosen card’s row and column grow while its neighbours shrink and stay visible. The face fades out before the profile fades in, so the two never fight. Text keeps its final size during the resize, so nothing reflows mid-motion. Interrupt it halfway and it continues from where it is. If your system asks for reduced motion, all of it switches off.
You can tell the difference in about two seconds, so here are two videos.
The card matrix on the landing page.
The Philosophers page, where the timeline follows the card you open.
What comes next
The open question is who pays. Four kinds of people use it, and I have not yet learned which one opens a wallet. The launch is the experiment, and the analytics are already in place to read it.
After that: more voices. The pipeline can add a philosopher in a day. Finding a good public-domain translation takes much longer.
If you found this useful, or have opinions, I would love to hear them. EmailLinkedIn