Say it out loud.
Guftgu is Urdu for conversation. You write to one of six voices — Marcus Aurelius, Krishna, the Prophet, the Buddha, Lao Tzu or Kabir — and it answers in character, in plain modern English, from its own public-domain book. Every line it quotes is checked word for word against that book before you see it; when nothing in the book fits what you said, the voice says so instead of making something up.

The quotes are checked, not trusted
The voice is a language model, and language models are fluent liars about books. So the reply is taken apart before it is shown: every span in quotation marks has to appear, word for word, in one of the passages that were handed to the model for this turn. A quote that does not is removed, and the ones that survive are shown in the margin with their book, chapter and verse, so you can keep them.
When someone is not okay
Some messages are not a request for philosophy. A crisis check runs before anything else — before the rate limit, before the model — and when it fires, the voice stops quoting, speaks plainly and puts Indian helplines in front of the person (Tele-MANAS, 14416). It costs no model call and never waits in a queue. Its tests have to fire on crisis fixtures and, just as strictly, must not fire on ordinary sadness, because a helpline number in answer to a bad day is its own kind of not listening.
The apostrophe that looked like a quotation
Vasuki learns to talk from tens of thousands of generated conversations, and the filter that keeps invented quotes out of that data was dropping more than half of a test batch. The model was not the one inventing: the filter treated single quotes as quotation marks, so in "I don't think you're failing" it read "re failing. You" as a quote, could not find it in the Gita, and threw the whole conversation away. Every reply with two contractions died — a filter that was quietly teaching the model to talk like a textbook. It was caught reading the drops by hand, before the full run.
A model of its own
Vasuki is a 60M-parameter Mamba — a state-space model rather than a transformer — with a transformer twin of the same size as its baseline, a 1,791-token vocabulary learned from scratch, and its own fused Metal kernel for the selective scan, which took a step that ran out of memory at 47.7 GB down to 7.6. Both are pretrained on an M4 Max and then fine-tuned on persona conversations; whichever quotes more faithfully on held-out conversations is the one that speaks.
