
Sept 14, 2026 • By M Gilang Januar
VoiceTaking started life as a voice notes app. You pressed record, spoke, and got back a transcript you could shape with an AI writing assistant. People used it for brainstorms, for journaling, for dictating drafts on a walk.
Then we looked at what they were actually recording. Over and over: meetings. Propping a laptop next to a speakerphone. Recording a Google Meet call on a phone and uploading the file afterwards. Typing frantically through a client call because the transcript only ever caught one side of it — theirs.
So we rebuilt the product around that. This is VoiceTaking v2: a meeting note taker that hears the whole room, and does it without a bot joining your call.
Three things define v2:
Most meeting note takers work the same way: you connect your calendar, they send a bot into your call, and that bot sits there as a participant. It works, technically. But it comes with a tax that everyone has quietly agreed to pay:
You are already in the meeting. Your computer is already playing the audio. Asking a second, remote participant to dial in and listen on your behalf is a strange way to solve the problem.
VoiceTaking captures the meeting from the browser you are already using.
When you start a capture, you pick the tab your call is running in and tick "Also share tab audio". That gives us what the other participants are saying. Your microphone covers your side. Both are transcribed live, on separate channels, so the notes always know who was talking.
Nothing joins your call. Nothing appears in the participant list. It works with Google Meet, Zoom in the browser, Teams, Whereby, a Discord call, a webinar you're only watching — anything that makes sound in a tab. There is no integration to set up, because there is no integration.
Two audio sources means two channels, and two channels means attribution that doesn't depend on guessing who is who inside a single mixed recording. Your side is your side. The call is the call. Within the call audio, speakers are separated too, so a three-way conversation doesn't collapse into one wall of text.
It's multilingual by default. If your team switches language mid-sentence, the transcript follows.
The live transcript sits next to an editor while the meeting is happening. Type what matters and let the transcript catch the rest. Anything you write before the call is kept and treated as fact when the summary is written, so an agenda you paste in beforehand genuinely shapes the output.
If you'd rather watch the call full-screen, pop the live transcript out into a floating window and keep it on top of everything else.
A technical interview needs a scorecard. A strategy review needs decisions and the metrics behind them. A standup needs blockers. "Summarise this meeting" gives all three the same shapeless paragraph.
So you pick a template first and the summary takes that shape, with its own sections — plus action items pulled out with an owner and a due date. The templates are editable, and you can write your own.
Capture saves and closes when you stop sharing the tab, or when nobody has been audible for a few minutes. The summary is written server-side afterwards, so you can close the tab the moment the call ends and come back to finished notes.
Two smaller things that came out of using it ourselves. Muting yourself inside Zoom is invisible to the browser, so there's a Mute mic control in VoiceTaking that actually keeps your side out of the notes. And the recording is kept and plays back against the transcript, for the thirty seconds you need to hear again.
Nothing you have is going away. Your recordings, transcripts and notes are where you left them, and the AI writing assistant still works the way it always did. Meetings are simply what the product is built around now.
Capturing another tab's audio is a browser capability, not a trick, and not every browser has it:
We'd rather say that here than have you discover it during a call that matters. It is also the reason we are building a desktop app.
Tab capture is a browser capability, and that ceiling is the browser's, not ours. So VoiceTaking v2 ships with a desktop app in progress.
Capturing at the system level instead of the tab level changes what's possible:
Mac first, Windows next. Your account, recordings, transcripts and templates are the same on both — the desktop app is another way into the same workspace, not a separate product.
Want to be in the first round? Start using VoiceTaking and we'll email you when the build is ready, or tell us which platform you need at [email protected].
Start a meeting — no credit card, nothing to install, no bot to explain. Open your call, start capture, pick the tab.
If it doesn't behave the way you expect, tell us: [email protected].