backchannel

Your meetings, transcribed live.
Your next move, surfaced mid-call — on your own hardware.

A self-hosted, open-source AI meeting assistant: live speaker-attributed transcripts and real-time insight agents, with no bot in the call.

Free desktop app for Windows, macOS & Linux · self-host anytime with Docker Compose · works with any meeting — Zoom, Meet, Teams, or in the room · MIT licensed

Backchannel live during a fictional recovery-readiness review: a quiet listening bar, the live strategic-signal strip, 125 insights, an answered mid-call question, a speaker-attributed transcript, and the ask bar along the bottom.
FIG. 1A high-stakes recovery review, handled remotely: signals across the top, a question asked and answered mid-call, the attributed transcript running beside it. 125 insights from one 46-minute call, distilled into a single briefing.
During the call

Everything a second set of ears should do

Built for high-stakes sales, discovery, and service-delivery calls — especially when the account team is remote. Provider-flexible, centered on a transcript you can trust.

Live diarized transcription

Silero VAD and WeSpeaker ResNet152 speaker embeddings attribute every line to a speaker, with interim text streaming seconds ahead of the final transcript. Microphone and tab/system audio are captured as separate tracks, so remote participants get their own identities.

Fictional recovery-readiness transcript with every line attributed to Me, Leah, Owen, or Maya and shown with timestamps.

Agents that work the call, not just the recap

Five live agents -- an analyst, a fast objection handler, a synthesizer, an opportunity specialist, and a strategic-signals scanner -- each run on their own trigger and push questions, objections, opportunities, and action items mid-call. Signals raised early are kept rather than overwritten, so the point someone made twenty minutes ago is still there. At call end, two independent briefing lenses draft in parallel and an arbiter reconciles them into one summary.

Speaker-attributed action items for the board-ready evidence outline and Tuesday technical working session, each badged with the agent that produced it.

Provider-routed models, on your keys

Mix Google Gemini and OpenAI per agent, transcribe fully offline with local ONNX Whisper and Parakeet, or point the agents at any number of your own OpenAI-compatible servers -- Ollama, LM Studio, vLLM, LiteLLM -- where every model an endpoint serves appears by name in each agent's picker, and run the whole pipeline with no key from anyone. Privacy First judges the destination, so an endpoint on your own machine or LAN keeps the agents running with the switch on; only cloud providers stay blocked.

Admin Connections tab: the Privacy First switch and Google and OpenAI credential cards above the Self-Hosted Models card for LM Studio, Ollama, vLLM, and LiteLLM endpoints.

Mid-call directives

Drop a note while the call runs and the analysis agents factor it into what they surface next.

Import and re-transcribe

Bring in existing transcripts or audio files, and replay any recorded call through a different model later.

Exports, chat, and token accounting

Transcript, insight, and summary exports, cross-session chat over your past meetings, and a per-session Tokens tab that breaks usage down by agent and model.

New in v0.5.0

Ask the call a question while it is still running

Twenty minutes ago someone named a number and you did not write it down. Type the question into the bar under the transcript and the answer comes back from the conversation as it stands — without stopping the recording, opening a second tool, or waiting for the post-call briefing. The answer is saved with the call, starred, and exported alongside every other insight.

The call's command bar: a Chat and Directive mode toggle, the question “What did Owen commit to sending tomorrow?” typed in, and a Gemini 3.6 Flash model chip marked Recommended.
The whole interface: two modes, one input, and the model that will answer.
A live call with two answered mid-call questions pinned at the top of the insight feed and a third being typed into the ask bar at the bottom of the screen.
FIG. 2Asked mid-call, answered from the transcript so far — and kept, so the answer is still there afterwards.

Two modes, one bar

Chat asks the call a question. Directive steers what the agents look for next. An answer you want the crew to act on converts to a directive in one click.

You choose who answers

The model that answers is picked in the bar itself — Google, OpenAI, or a model on your own hardware. Nothing is selected on your behalf.

Grounded in the room

Answers draw on the live transcript, the insights raised so far, strategic signals, your directives, and the summaries of documents you attached.

Live

Questions answer themselves

The synthesizer keeps reading the room. When a question gets answered mid-call, it marks the card Answered, summarizes the answer, and spins off the follow-up you still owe. No bookkeeping from you.

The live feed filtered to 34 questions: the leading card is marked Answered with the synthesizer's one-line answer summary, and carries the follow-up question it spun off.
FIG. 3Mid-call — the synthesizer hears the answer land, marks the question Answered, and hands you the follow-up you still owe.
After the call

When the call ends, the wrap-up runs itself

Stop the call and Backchannel drains the last segments, runs the final analysis, and saves the session, then rolls the whole conversation into one reviewable record.

Alderwake recovery-readiness session header showing a 44-minute 12-second call and a two-minute resumed segment.
FIG. 446:12 across two segments — the resumed call and the original settle into one reviewable record.

The briefing does the reading for you

A 46-minute call produced 125 grounded signals. You never wade through them. Backchannel distills the whole conversation into one briefing that opens with an at-a-glance strip and leads with the top outcomes — risks, actions, and open questions each color-coded, owners named as people rather than identifiers — so you start from the point instead of the pile.

125 1 One briefing of outcomes, objectives, and follow-ups — with every underlying insight kept, attributed, and exportable.
Alderwake recovery-readiness briefing: an at-a-glance strip of three outcomes, four actions, three risks and three open questions, the kept strategic-signal history, and the top three outcomes with named owners.
FIG. 5The settled briefing — two independent lens drafts reconciled by an arbiter, opening with the whole meeting in five seconds.

Nothing raised during the call is quietly dropped. Signals the strategic scan surfaced are kept with how often each came back, so a theme that ran through the conversation reads as a theme rather than one line you happened to catch.

Strategic Signal History expanded in the briefing: six kept signals labelled Action, Discovery Question, Opportunity, Signal and Risk, each showing how many times it was seen and when it was first and last raised.
FIG. 6Durable signals, with first and last sighting — “seen 6 times” is the difference between a passing remark and the thing the meeting was actually about.

The detail is all still there: every one of the 125, each attributed to who said it, ready to take with you as one enriched Excel workbook, HTML, or text.

Insights tab: 125 total, with 2 asked, 24 action items, 16 objections, 18 opportunities, 31 observations, and 34 questions above the answered mid-call questions.
FIG. 7All 125, kept and attributed — 24 action items, 16 objections, 18 opportunities, 31 observations, 34 questions, and the 2 you asked yourself.

Diarization gets close. You get it right.

Automatic diarization is fast, but it mislabels: splitting one person across several voices, or blending two into one. Rename, merge, and tag speakers by hand, then re-run the analysis so every insight reflects who actually said it.

Rename, merge, and tag

Silero VAD and WeSpeaker embeddings assign a voice to every line before anyone is named, and they don't always match reality. Map each auto ID to a real person, tag your side and theirs, and merge the duplicates a long call inevitably throws off.

Speakers tab: auto-detected speakers with name mapping, team and external tagging, merge controls, and an Enhance Insights button.

Re-run, and every insight updates

Hit Enhance and Backchannel replays the analysis over the corrected speakers — “Me to draft the SOW” instead of “Speaker 4” — so every action item and observation is re-attributed to the right person.

Action items after speaker mapping: each card names the person it belongs to, quotes the line it came from, and flags the one still needing a follow-up.

Ask across every meeting

Every transcript stays local, and stays queryable. Ask a question against one call or all of them at once, and get a grounded answer with the meetings it came from.

Cross-session chat answering what was committed and what is blocking the timeline, with a grounded answer and session scope pickers.
FIG. 8One question asked across every saved meeting, answered with the sessions it came from.

Install the way you want to run it

Use the desktop executable when you want the shortest path, or Docker Compose when you want the full self-hosted stack, development hooks, and GPU options.

Docker Compose

Self-host the full stack

Compose keeps every service isolated and is still the most flexible path for local development, server installs, and optional NVIDIA GPU diarization.

git clone https://github.com/talberthoule/backchannel.git
cd backchannel
cp .env.example .env   # optionally set GEMINI_API_KEY here
docker compose up --build

# app at http://localhost:3000 -- add API keys any time in Admin -> Connections

The first start builds images and downloads models, so give it a few minutes. Have an NVIDIA GPU? Add the override for GPU diarization: docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d --build

Built-in local batch transcription and diarization need no API key. For insight agents, connect Google, OpenAI, or a self-hosted OpenAI-compatible service, then choose explicit models under Admin. Local recommendations appear after the connected model passes Local Fit.

From spoken word to actionable insight

The live call path, end to end.

  1. Capture. The browser records mic (and optionally tab) audio and streams PCM16 16 kHz chunks over a WebSocket.
  2. Diarize. Voice activity detection cuts speech into segments; speaker embeddings assign each segment to a voice.
  3. Transcribe. Each segment is transcribed by the selected Google, OpenAI, or local ONNX model, filtered, and saved in speaking order.
  4. Analyze. Text agents read the growing transcript on their own schedules and propose questions, objections, opportunities, and action items.
  5. Deliver. Deduplicated insights stream to the call view instantly; a synthesizer refines them as the conversation evolves.

Three services, two AI paths, one transcript

A React SPA, a FastAPI backend, and PostgreSQL: a low-latency interim path for feedback and a durable batch path for the record.

Backchannel architecture diagram: client layer, backend layer, AI layer, and persistence

A crew, not a monolith

Each agent has one job, its own trigger, and a configurable model and prompt. Agents start as Not selected; provider-aware Recommended tags offer starting points without forcing a choice.

AgentTriggerPurpose
Audio Bridge
audio_gateway
Continuous audio stream Listens silently via Gemini Live, OpenAI Realtime, or the on-device Parakeet captioner and streams interim transcription
Consolidated Analyst
consolidated_analyst
Every 40s + final pass Runs configurable lenses for questions, observations, opportunities, and action items in one pass
Objection Handler
objection_handler
Every 10s over the last 90s Flags objections with an immediate response and the underlying strategic concern
Principal Agent
synthesizer
New or updated insights; 75s cooldown, 120s fallback Quality-checks insights and connects them into broader objectives, initiatives, and patterns
Opportunity Specialist
opportunity_specialist
New opportunities; 55s cooldown + final match Matches opportunity insights against the configured knowledge sources without creating new insights
Strategic Signals
strategic_signals
Every 45s during the call Refreshes the live signal, risk, next question, opportunity, and action cue while linking supported cards to saved insights
Briefing Meeting Lens
brief_meeting_lens
At call end or on demand Drafts the meeting record: outcomes, decisions, blockers, commitments, and follow-ups
Briefing Discovery Lens
brief_discovery_lens
At call end or on demand Drafts the broader signal: objectives, pains, gaps, opportunities, risks, and open paths
Briefing Arbiter
brief_arbiter
At call end or on demand Reconciles the two independent lens drafts into the settled briefing

Privacy, data flow, and what it takes to run

The questions that decide whether you self-host, answered plainly.

Is Backchannel free?

Yes. Backchannel is open source under the MIT License, with no hosted tier, no seat pricing, and no feature paywall. Your only costs are your own hardware and any Gemini or OpenAI API usage you configure.

Does my meeting audio leave my machine?

Only where you route it. Voice activity detection and speaker diarization always run locally, and a fresh install uses built-in local batch transcription. Agent models stay unselected until you choose Google, OpenAI, or a self-hosted OpenAI-compatible server such as Ollama or LM Studio. Recordings and transcripts stay on your server.

Does it work with Zoom, Google Meet, Teams, and in-person meetings?

Yes -- any meeting, digital or in person, because it never joins the call. Backchannel captures your microphone and, optionally, tab or system audio directly in the browser, so there is no bot participant and no per-platform integration to install. For a conference room or an in-person conversation, the microphone alone carries the meeting, with diarization separating the voices. Cross-platform coverage is no longer unique: Google's take-notes-for-me now reaches meetings hosted on other providers, and Zoom's assistant can join Meet and Teams. Both do it as a visible bot in a vendor cloud, and neither can sit in a room.

How is it different from Otter.ai and other cloud note-takers?

Real-time help is no longer rare: Otter shipped Live Assist in July 2026 and Zoom's Sales Assist reached general availability a day later. Both are Enterprise-priced and tied to their own platform, and like Fireflies.ai and Granola they run in a vendor cloud. Backchannel is self-hosted, MIT licensed, and joins no call as a bot, and its objection handler generates a response to the objection actually raised, every 10 seconds over the freshest 90 seconds of speech.

Can I ask a question during the meeting?

Yes. The call's command bar answers from the conversation as it stands — the live transcript, the insights raised so far, strategic signals, your directives, and the summaries of attached documents. Recording continues while it answers, the answer is saved with the call and starred, and it exports alongside every other insight. You pick which model answers, including a model running on your own hardware.

Do I need a GPU?

No. CPU-only Docker Compose is the default. An NVIDIA GPU can accelerate diarization via a compose override, and AMD GPUs on Windows are supported with a native backend setup script.

What do I need to run it?

Use the desktop app for the easiest start: download the build for Windows, macOS, or Linux, unpack it, and run the app. Built-in local batch transcription needs no API key. Agents start as Not selected: connect Google, OpenAI, or a self-hosted service if you want its models, then choose explicit models in Admin. Recommended marks a good starting point. Use Docker Compose when you want the full self-hosted stack, local development, or GPU diarization.

Stay in the loop

Follow the project on GitHub

The source, issues, and release tags are all public. Star the repository, watch it for new releases, and read the release notes as they land — every desktop build is a free download, no account needed.

Prefer email? No weekly drip campaign — just release notes and product updates worth opening.