Why it works like this.
The front page says the harness is the product. This is what that means in practice: a handful of decisions, each one made because something went wrong first.
You can see what an answer will cost before you send it.
Context is measured slot by slot — system prompt, tool schemas, memory, documents, images, history — and shown as a breakdown rather than a bar creeping toward a limit. A number you cannot break down is not information, it is a warning light.
Long conversations compress instead of ending.
Older turns are folded into a summary and a short knowledge base that the model keeps reading, so nothing hits a wall. What you can scroll back through is untouched — the compression is for the model, not for you.
It can look up your old conversations mid-answer.
Before this, anything worked out weeks ago was re-derived from scratch every time it came up — only a person could search, and only from the sidebar. Now the model can go and look, mid-sentence, the way you would.
When a tool needs something from you, it asks in a card.
Answering resumes the turn from where it stopped rather than starting it over, and the question keeps its own state — so a card answered an hour later still works, and one whose code changed underneath it renders inert instead of breaking.
What it learns is written after the answer, not during it.
Extraction used to sit in the tool loop, and a memory saved mid-answer made the model re-synthesise from the top — you watched a half-written reply get wiped. It runs once the answer is finished now, which also means it decides what to keep having seen the whole exchange.
The pet
Leo came from Silvermage, the command-line tool this shares a house with. Seventeen faces, straight from the original spec. Pick one and he will make it.
Hello. I live at the top of an empty conversation.
default, friendly
Four people use this. It runs on one machine at home, it is built most nights, and there is no way to sign up for it — the front page has the short version.