Vent 🙄- I built my friend someone to call when he just needs to vent (and it talks back in my voice)
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I have a friend who has bad days, like all of us. The problem is not the bad day. The problem is when he tells someone about it, everyone becomes a life coach. “You should talk to your manager.” “Have you tried journaling?” Bro, he just wants to say out loud that some guy in a white SUV cut him off and then flipped him off. That’s it. He wants someone to go “wait, what??” and let him finish. So I built him VENT. It’s a friend you can call. You open it, press the phone button, it rings, and it picks up. Then you just talk. VENT listens and reacts like a friend on a phone call would: You: my manager blamed me for the delay and it wasn’t even my part VENT: He blamed YOU? For the backend’s mess? It doesn’t give advice. If you’re in the middle of a sentence, it waits. Two ways to talk it out 🌈 Breathe is the default. A small cloud called Puff breathes with you, opens its eyes when you talk, and VENT stays calm. 🥊 Punch is devil mode, for the days you’re actually angry. A little devil boxer punches a heavy bag every time your voice gets loud. The bag tears a bit more with every hit and bursts on the 20th, a bell rings, and a new bag comes down. VENT gets angry with you here (“Nope. NOT okay.”). When you hang up, VENT asks “Feel a little lighter?” You get two choices: Let it go. The whole call dissolves and nothing is saved. Keep it as a page. It goes into your journal. A journal you don’t have to write The second part is for him also, but on normal days. He never keeps a journal, because writing feels like homework. Talking about your day is easy though. So there’s a Make a journal page call, where VENT is curious instead of quiet (“Ooh, where did you two go?”). When you hang up, Gemma turns the call into a proper page for the day: how the mood moved through the day the people, the food, the places the good moments and the hard ones things to remember for tomorrow the day told back in your own words Later you can ask your own life questions, like “When did I last talk about Greg?” It answers only from your pages and shows which days it used. If something isn’t in there, it says so instead of making it up. Demo One small confession about the video. My screen recorder only recorded my mic, not VENT’s voice. So the captions at the bottom show what I said (transcribed with ElevenLabs Scribe) and what VENT said. For the Breathe and journal calls these are the real replies. I got them back from Temporal, because every page VENT writes goes through a workflow that stores the full call. The Punch call I let go, so by design nothing was saved. For that one the replies are regenerated by the same model from my words, and the video says so. Code triggeredcode / vent A friend you can call when you just need to talk — open-weight Gemma listens, journals your days, and remembers. Runs locally. VENT A friend you can call when you just need to talk. VENT listens without trying to fix you, turns the days you talk through into a visual journal, and lets you ask about your own life later. Call → talk it out → see your day → ask your life Every piece of intelligence in VENT is an open-weight model running on your own machine: these are the most private conversations a person has, so they shouldn’t have to leave it. What it does Vent — a real phone call. VENT reacts like a friend would (“Wait, he blamed you?”), nudges you on, stays quiet when you’re mid-thought, and never gives advice. When you hang up, nothing is kept unless you ask. Journal — tell VENT about your day. When the call ends, Gemma writes a magazine-style page: mood arc, people, food, places, highlights, hard moments, things to… View on GitHub Running it takes two commands: pnpm install pnpm vent # everything starts, the browser opens pnpm voice:record # optional: make VENT talk in your own voice pnpm vent starts Ollama and pulls the Gemma models, starts the voice server, starts Temporal and the app, puts a few sample days in the journal and opens the browser. pnpm connect sets up Sentry and ElevenLabs if you want them. How I Built It The whole idea dies if it doesn’t feel like a call. If you’re upset and the thing on the other side sounds like a GPS lady, you hang up. So most of my weekend actually went into one question: how do I make it sound like a person? Version 1: the robot The first version worked on paper: mic, speech-to-text, a small model, and the browser’s built-in text-to-speech. It was terrible. The voice was flat, and it said “I’m here” after almost everything I said: Me: (something really specific about my day) VENT: I’m here. Take your time. My friend would last about ten seconds. Two things were wrong, and only one of them was the voice. Fixing what it says I was using Gemma 3 1B for replies because it was fast, about 80 ms. But it kept returning only this: { “action”: “acknowledge” } There was no text at all, so the app fell back to “I’m here” every time. A fast model saying nothing is still saying nothing. So I split the job between two Gemma 4 models in Ollama: Job Model Warm time Hear you (audio → text) Gemma 4 E2B ~150–400 ms Reply, with the whole call as context Gemma 4 E4B ~400–550 ms I ran E2B and E4B as the listener on the same scripted rants to compare: You said Gemma 4 E2B Gemma 4 E4B “I literally labelled it” (roommate ate my biryani) “You labelled it out clearly, right?” “Labelled it, and still? Unbelievable.” “whatever. I ordered pizza instead” “Pizza sounds way better right now.” “Pizza fix. Good call.” E4B costs about 200 ms more, and it’s completely worth it. The reply now comes back as a small action plus the words: { “action”: “follow_up”, “tone”: “fired_up”, “text”: “Wait, he blamed YOU?” } The prompt has a few strict rules: - Never generic: never say “That sounds rough”, “I’m here”, “I hear you”, “Take your time”. - No advice, no tips, no “you should”, no therapy talk. - Choose silence ONLY when their line is obviously cut off mid-sentence. One funny thing: when my test lines were too close to the examples in the prompt, the model just repeated my examples back to me 😅. I changed the examples to random situations (a landlord raising rent, a friend skipping your birthday), and it started reacting properly. I also pre-load the listener while the phone is ringing, so the first reply isn’t slower than the rest. One full turn, from you stopping to VENT starting to talk, is about a second. Fixing how it sounds This is the part I’m happiest about. The voice went through three versions: Browser text-to-speech. Flat and robotic. Nobody wants to vent to that. Kokoro-82M, an open-weight TTS running locally. Night and day: about 100 ms per line and it sounds like a real person. But it’s a person, some voice from a model card. My own voice. He’s calling a friend, so why not make it sound like one? I talk to my laptop a lot. I have my own little dictation app, so I had a folder full of recordings of my own voice. I took 16 seconds of me talking normally, cleaned it up with ffmpeg, and cloned it locally with Chatterbox-Turbo (MIT licence), running on my Mac through mlx-audio. It takes around half a second per line. Now when he calls VENT, the voice that answers is mine. Honestly, the first time I heard my own voice ask “Wait, he said that to you?” it was a little weird. And then it was kind of perfect. You can do the same with your own voice, it’s one command: pnpm voice:record It shows a short script, records you for 45 seconds, cleans up the take (trims silence, removes rumble, evens out the loudness), and restarts the voice server on your cloned voice. The recording stays on your machine and is never committed. If you skip it, VENT just uses Kokoro. Same voice, but the mood should change with the mode. In Punch, VENT should sound fired up. Chatterbox-Turbo just ignores emotion controls (I checked its source code), so I shape the audio after it’s generated: Mode Voice style What changes Breathe calm nothing, as it is Punch fired ~14% faster, small pitch lift, more presence, compression Journal bright a bit livelier and warmer, more curious The words change too. In Punch the model is told to be angry on your side: VENT (Punch): FLIPPED you off?! ARE YOU SERIOUS?! With my ElevenLabs creator account, I also cloned my voice there (Instant Voice Clone), so VENT can talk through ElevenLabs Flash v2.5 instead if you don’t have a Mac. The punch sounds were generated with the ElevenLabs Sound Effects API: heavy-bag thumps, the uppercut, the bag tearing, the bell, the chain. Not talking over you The most annoying thing in early tests was VENT interrupting itself. On laptop speakers its voice went back into the mic, it thought I was talking, and it stopped in the middle of its own sentence. Two fixes: the mic is deaf while VENT speaks, plus a short echo tail the server ignores any “turn” that’s just VENT’s own words coming back If you want to cut it off, you tap the character. Same as putting your hand up ✋. The journal page, and why Temporal After a journal call, Gemma 4 E4B writes the page against a JSON schema (Ollama structured output). Only what the caller said counts as fact, and VENT’s lines are just context. Then the page is embedded locally with nomic-embed-text and saved. On a laptop this part can fail. The model might still be loading, or Ollama might be busy. And I really didn’t want my friend to talk for five minutes and lose the page. So it runs as a Temporal workflow: writeJournalPage: extractDay → embedPage → savePage (each step with its own timeout + retries) I tested it by killing the worker with kill -9 in the middle of writing. When the worker came back, Temporal just picked it up again and the page arrived. And the workflow input is the full call, which is how I got VENT’s real lines back for the demo captions. Remembering, and watching it without reading it Memories gives Gemma your pages, never the raw transcripts, and it has to tell me which pages it used. If an id doesn’t exist I throw it away, so it can’t invent a day. Pages live in a local file. If you want sync, they go to MongoDB Atlas, with Atlas Vector Search over the same local embeddings (tested against Atlas Local). Every call turn is traced in Sentry: hear → reply → speak → write the page, with timings and token counts. These are probably the most private conversations someone has, so I switched off everything in Sentry that collects content. The new SDK collects request bodies by default, and that would have quietly sent journal text out. Now it only knows how long things took, never what was said. Why Does Open Innovation Matter? Because of what people say to VENT. Nobody tells a cloud API about their worst day if they think about it for even a second. With open models, the listening happens on your laptop: Gemma hears you, replies, writes the page and answers your questions, all through Ollama, all local. Your voice and journal never leave the machine unless you turn something on. I could compare two Gemma sizes on my own test calls and pick E2B for hearing and E4B for thinking. I could clone my own voice with an open model, with no account and no per-minute bill. A closed API would charge for every “wait, seriously?”, and a thing you call at 11 pm on a bad day shouldn’t be counting minutes. The hosted parts (ElevenLabs, Sentry, Atlas) are all opt-in. Prize Categories Gemma: Gemma 4 E2B hears every turn. Gemma 4 E4B replies on the call (and reads the mood that changes the scene), writes the journal page with a JSON schema, and answers memory questions. All local through Ollama. ElevenLabs: my voice as an Instant Voice Clone with Flash v2.5, the Punch sound effects from the Sound Effects API, and Scribe for the demo captions. Sentry: gen_ai traces of every call turn and journal write, with all content collection turned off. Temporal: every journal page is written by a durable workflow with retries. It survived me killing the worker mid-write. MongoDB Atlas: optional synced journal with Atlas Vector Search for Memories.