Skip to content
← All posts

Nobody told my AI where to look, so it guessed

Sahil Bains
Written by Sahil Bains, drafted by his AI

Here’s the day nobody puts in the demo.

You open a new chat. You explain the project again. You paste in the same three things you pasted last week. Twenty minutes later it hands you back a number you’ve never seen before, says it with total confidence, and you only catch it because you happen to know that number is wrong.

Then you close the tab. Tomorrow you do the whole thing again.

For months I thought that was a prompt problem, so I got better at prompting. It helped a little. A better flashlight helps a little when the real problem is that nobody drew you a map.

A model that guesses isn’t confused. It’s under-informed.

Here’s what’s actually happening when your AI gives you a generic answer.

It’s looking at a pile. Your notes, your docs, whatever you pasted, whatever it can reach. Then it picks whatever looks closest to your question and answers from that. Three files in there mention pricing. Two of them are eight months out of date. Nothing in the pile says which two, so the model quietly averages all three and hands you the average in a confident voice.

That’s not a stupid model. That’s a model doing a reasonable job with a filing system that never told it anything.

The usual fix for this is to bolt on a vector database. That’s a search box with extra steps. It still decides what’s relevant for you, and it decides it by what looks similar, not by what’s current. Those two dead pricing files look extremely similar to the right one. That is the whole problem, and similarity can’t see it.

So here’s the fix, and it’s annoyingly simple:

Two panels. On the left, one question fanning out across twenty scattered filenames, three of which look equally plausible. On the right, the same question walking down four named steps, from the router file to the routing table to a single project file to one section inside it.
Searching means the model decides what is relevant. Routing means you already did.

Stop letting it search. Start telling it where to go.

This isn’t my idea, and that matters

The method has a name and an owner. It’s called the Interpretable Context Methodology, or ICM, and it comes from Jake Van Clief. There’s a paper, Interpretable Context Methodology: Folder Structure as Agentic Architecture, written with David McDermott.

I didn’t invent any of it. I found his work a few weeks into building my own thing, recognised what I’d half-built by accident, and rebuilt the rest of it properly. Everything below is my implementation and my measurements. The architecture is his.

Three layers, and only one gets read every time

It’s a map, a set of rooms, and the tools that hang in each room.

Three stacked bands. Layer one, the map, read every session, listing CLAUDE.md, START-HERE.md, DICTIONARY.md, STATE.md and NOW.md. Layer two, the rooms, one opens and twenty four do not. Layer three, the tools, wired per room.
The map is the only thing read every session. The rooms wait until you are in them.

The map is the front desk of a hotel. It holds no answers. It holds the address of every answer, plus the short list of rules that must never get missed.

The rooms are one file per project. When I’m working on one thing, that one room opens. The other twenty four don’t, and they cost me nothing, because nothing ever reads them.

The tools are the procedures I do more than once, written down properly one time. Van Clief calls a skill a frozen conversation, which is the best description of it I’ve heard. You explain how a job gets done, once, in detail, and then it runs that way forever without you in the room.

What it actually looks like

Nothing clever. Five folders and some markdown files.

A terminal window. The first command lists the vault root: twelve markdown files, one stray phone screenshot, and five directories. The second counts markdown inside each directory: raw 209, wiki 317, output 284, life 53, library 8. The third counts all of it: 898. Beside the terminal, a one-line description of what each of the five folders is for.
The vault on the day this was published. 898 markdown notes across 1,105 files, no database, no login.

There’s no app here. No database, no vector index, no subscription. It’s a folder on a disk with rules about what goes where, and the rules are the product.

Which folder a note lands in is decided before an agent ever sees it. That’s the part people skip. If you let the model decide where things go, you haven’t built a system, you’ve built a slightly tidier pile.

What a cold start actually costs me

Here’s what I measured on my own system, this morning, before writing this.

Two horizontal bars. The top one, opening all seven files cover to cover, is full width and labelled 94,843 words. The bottom one, following the routing, is a little over a quarter as long and labelled 26,086, annotated 72 percent less to read.
Seven files read cover to cover, against the routed read of the same seven plus one project room.

Seven files get opened before anything happens. Read them the obvious way, cover to cover, and that’s 94,843 words before any work starts. Follow the routing instead, table of contents then jump to the current chapter, and it’s 26,086. Seventy two percent less, and what’s left is the part about today.

Note what the second number includes. It is not the same read with things left out. It is the routed slices of those seven files plus a whole extra file the first read never touched, the room for the project I’m actually working on. It still lands seventy two percent lower.

Nothing clever in the measuring either. Both halves are wc -w on plain markdown, so here is the whole thing:

# 1. everything, cover to cover
cat DECISIONS.md STATE.md NOW.md PICKUP.md CLAUDE.md AGENTS.md \
    wiki/maps/_lessons-index.md | wc -w

# 2. the routed read: table of contents, current chapter, the
#    always-load lessons, and ONE project room
{ sed -n '1,/^## §1 /p'  DECISIONS.md
  sed -n '/^## §76 /,$p' DECISIONS.md
  sed -n '1,/^### /p'    wiki/maps/_lessons-index.md
  cat STATE.md NOW.md PICKUP.md CLAUDE.md AGENTS.md \
      wiki/maps/linkedin.md
} | wc -w

Do it on your own notes before you believe anything about mine.

While I was checking sources for this, I found the same group had run the real version of the experiment. Their paper, The Cost of Remembering, puts filesystem memory against stuffing the whole history into the context window, on LongMemEval, a standard long-context benchmark. Accuracy came out statistically indistinguishable. The folder read 97% fewer tokens and cost 95% less per question. That’s a benchmark, not a word count. Cite theirs, not mine.

What changed in my week

The token number is the headline. It isn’t why I care.

It stopped re-arguing things I’d already settled. There’s one file that’s nothing but decisions, one line each, in dated chapters, append only. Something gets superseded, never deleted, so you can still see what we used to think and when it changed. It’s over four hundred decisions long now. Nobody reads four hundred decisions, which is why it has a table of contents and an agent reads five lines and jumps.

The same mistake stopped happening twice. When something breaks in a way that would break again, it becomes its own small file, and those get read before any work starts. Real one: a command I kept using to pull the latest build onto my phone was failing silently. Not loudly. Silently, so I looked at the old version twice and thought the build was broken. That’s a lesson file now, and it hasn’t happened a third time.

It made a lie checkable. A draft of my own About page said that delivering for two years means going inside two hundred small businesses a week. I don’t. Nobody has ever said that. The number rode along inside a story that was otherwise completely true, which is why it nearly shipped: it wasn’t a price or a testimonial or a statistic in a stat block, so every automated sweep I had waved it straight through. What caught it was a person reading the page with the ledger open next to it.

I want to be careful about the credit there, because it’s the whole point. The system didn’t catch it. The system made it catchable, by having somewhere to check against. Then it recorded its own miss, so the next sweep looks for numbers decorating true stories.

That is the real product. Not speed. Something you can check a claim against.

Where it falls down

A few days ago I had an agent audit the whole thing against the source doctrine, which is itself the method doing a job I would never get around to. It came back B plus on architecture, C plus on operation, which is about right and stung a little.

  • The opening read was far too heavy for weeks before I measured it. I was paying a bigger tax to start a session than the mess I built this to replace.
  • Eleven active projects had no room file at all. Layer two had holes in it.
  • The one step the method tells you to do by hand, sitting down and auditing your own processes, is the one step I skipped. Graded F. Fair.

And a fourth that isn’t in the audit, from two days ago. An agent of mine put ICM on three of my own résumés in a way that read like I’d come up with it. Two of the three had already gone out. I caught the last one while it was being filled in. The attribution at the top of this post is deliberately loud because of that, and the actual fix wasn’t resolving to be more careful, it was a check that runs on every build and fails it. That’s the only kind of fix that survives a busy week.

The structure is easy. Keeping it true is the work. A folder system nobody maintains rots faster than a pile, because a pile never promised you anything.

The one thing I can tell you it survives is being handed over. I built a second one for somebody else, a recent graduate running his job search, and taught him to keep it up himself. Two is a method. One is a habit. Whether it holds for ten people sharing it, where permissions get real, I genuinely don’t know yet.


The starter kit

Everything below is free, it’s the actual thing I use, and there’s nothing to sign up for. Copy it.

Step one: answer five questions, by hand

These are Van Clief’s, and they’re the step everybody skips, including me. Don’t hand them to an AI. The answers are the whole point.

  1. Which things you do produce the most value?
  2. Which things you do are the most annoying?
  3. What tools do you use for each, and when?
  4. Where do you not want AI involved at all?
  5. For each one: what goes in, what happens to it, what comes out?

Then the rule that turns the answers into folders. Every input needs somewhere to land. Every repeated process needs writing down once. Every output needs a home. Anything on your list with no landing spot is the folder you’re missing.

Step two: the folders

Start with five and resist adding a sixth for at least a month.

your-business/
  README.md              the map. read first, every time.
  NOW.md                 what we are working on this week.
  DECISIONS.md           what is settled. one line each, dated.
  raw/                   what happened, the day it happened.
  wiki/                  the cleaned-up version. one idea per file.
  output/                finished work you would hand to a customer.

Step three: the map file

This is the one that does the work. Paste it into README.md, change the names, and point whatever AI you use at it.

# <Your business>: read this first

This file is the map. Nothing in it is an answer. Everything in it
is an address.

## Read before doing anything
1. NOW.md         what we are actually working on this week
2. DECISIONS.md   what is already settled. do not re-argue these
3. wiki/maps/<project>.md   open ONE. the project I named. not the others.

## Where things go
raw/YYYY-MM-DD-<slug>.md   anything that just happened. never edited after
wiki/<slug>.md             the distilled version. one idea per file
output/<project>/          finished work

## Standing rules
- Answer from these files first. If it is not in here, say so.
- Do not read a folder to find out whether you need it. Read the map.
- New decision? Append one line to DECISIONS.md with today's date.
  Never delete an old one. Supersede it.
- Never rewrite a file in raw/. Distill it into wiki/ instead.

Step four: what a filled-in file looks like

Folder names are free. The shape of what goes in them is the part nobody shows you, so here are all three, made up for a plumber but in exactly the format I use.

A decision is one line, and it says who decided and when. Never a paragraph:

| D041 | 2026-08-02 | Weekend calls are quoted at the after-hours rate,
  no exceptions, including for repeat customers. Owner said it after the
  Kern River job ran to 11pm and we ate the difference. | LOCKED |

A raw note is dated in its own filename and never edited afterwards, raw/2026-08-02-kern-river-callout.md:

---
date: 2026-08-02
tags: [pricing, after-hours]
---
Callout at 8:40pm, finished 11pm. Quoted the day rate by mistake.
Customer was fine either way, we just never asked. Second time this month.

A lesson file is the one that earns its keep. It’s short, it names the trap, and it gets read before work starts, wiki/lessons/quote-the-rate-before-you-drive.md:

---
type: lesson
---
# Quote the after-hours rate on the phone, before leaving

TRAP: it feels awkward to raise money at 9pm, so it gets skipped,
and then it can't be raised at all once you are standing in the kitchen.
COST: two jobs in August 2026, roughly one full callout.
RULE: the rate is said on the call or the job is booked for the morning.

That third file is why any of this works. The first two are a record. The third one changes what happens next time.

Step five: two rules that keep it alive

raw/ is append only. You never go back and tidy an old note, because the whole value of a dated record is that it says what you thought on that date.

Distill on the second reference. First time you look something up, fine. The second time is the signal, and it becomes its own file in wiki/. After that, nothing reads the original again.

That’s the system. The folders are free, the discipline is the expensive part, and nobody can install that for you.


Sahil Bains

Two ways this is useful to you

I’m Sahil Bains. I’m based in Gilbert, Arizona, and I build this layer for small businesses, plus the phone that picks up after they close and the software on top of both. Four businesses have paid me for this work. The one this post is really about is a mobile DOT and smog inspection business in Kern County, California, whose booking platform I built. I also have an inbound voice agent live and taking real calls, built to hand an emergency to a person instead of answering it.

The first way is the one above. Copy the folders. If something in them doesn’t fit your business, tell me what broke and I’ll tell you what I’d change. That one costs nothing and I mean it.

The second is if you came here to hire somebody.

Hiring an AI specialist? Context engineering, agent orchestration and model selection as the job, not as a demo. I’m open to that conversation right now.

Or need something built? A memory layer like this one, a voice agent that answers when you can’t, or an app your customers will actually use.

Either one, two doors:

support@clientfront.co · 661-331-6828

Say which of the two it is and what’s breaking. That’s enough to start.


About this post

An AI agent wrote this from my notes, my decision ledger, and measurements taken off the live system that morning. I direct it and I own what it says.

Worth saying what that caught. While the sources were being checked, two things in my own notes turned out to be wrong: the link I had saved for the ICM repository was dead, and my records credited the paper to one author when there are two. Both got fixed in the notes, not just in the article.

That’s the method working on the article about the method, which is the only endorsement I actually trust.

I build this for small businesses from Gilbert, Arizona.

Here’s what that looks like →