← Back to Case Studies

Question Triage Time Cut By 95% In Live Sessions

Hosts who run their live sessions on Queuena.com spend 95% less time filtering and selecting audience questions, and every important question still reaches the set they answer. ML LABS builds and runs it for calls, webinars and conferences, and the design decisions described below are what produce that result on every session.

The problem it solves is familiar to anyone who has opened a channel for an audience to ask anything. Within minutes there is a wall of questions that overlap, repeat and contradict each other in wording while asking the same thing underneath. Nobody can read that wall live, and a host who tries spends the session sorting instead of answering.

In a glass-walled event room, a speaker answers raised hands beside a screen ranking three questions and a QR code for new ones
A room's questions, answered as three.

95% Less Triage, No Question Lost

  • Time spent filtering and selecting questions drops by 95%.
  • Every important question stays in the set the host answers.
  • Calls, webinars and conferences run on the same flow.

What the host sees is a short ranked list: a handful of umbrella questions, each standing for every question it has absorbed. The host answers the list, and the room hears its own questions answered, including the ones asked in words nobody else used.

Questions Merged By What People Meant

The first decision is what "the same question" means. Two questions phrased nothing alike are often the same question, and two questions sharing most words are often different ones. Collapsing on surface text produces false merges and false splits at the same time, so Queuena.com merges on intent instead: on what the asker actually wants to know.

It merges hard: roughly five umbrella questions, never more than ten, however many arrive. That cap is not a performance limit. It is the summary's spine: a system with room for forty umbrellas will produce forty, and forty umbrellas is the flood with better formatting.

Four well-merged umbrella questions beat eight narrow ones, because the presenter can hold all four in their head at once and act on them.

Stuffing the whole accumulated set of questions into one long prompt would not rescue fidelity either. Long contexts bury information in the middle (Liu et al., 2023), so the fidelity that looks bought is spent on the tokens a model reads least carefully.

No Asker Left Silently Unanswered

A wrong merge and a wrong split fail differently, and that asymmetry sets the policy. A split that should have merged shows up on screen as two near-identical umbrellas, and the host answers both at once. A merge that should have split is invisible: the absorbed question never appears, and its asker sees it go unanswered while the summary looks done.

So the model's job is deliberately narrow. It proposes the umbrella structure; the cap, the ordering and the merge policy are code, and none is negotiable by the model. Keeping a model's scope tight and its authority explicit (Anthropic, 2024) is what makes the output predictable enough to put in front of a live audience without a person checking every refresh. That is how the time drops without the important questions dropping with it.

Summaries Refresh As Questions Arrive

A stale summary is worthless, so Queuena.com regenerates it every five new questions. The trigger is arrival rather than a clock: a live talk that draws a burst of questions gets a fresh list immediately, while a quiet channel with a slow trickle does not pay to recompute a summary that no new question has changed since the last refresh.

That cadence runs the summarizer constantly during the moments that matter most, and a naive build re-sends the entire system prompt and context on every call. Continuous re-summarization then becomes a feature to ration, and a rationed summary is a stale one. The fix is to cache what does not change: prompt caching that reuses a fixed prefix across calls (Anthropic, 2024), applied to the system prompt, makes each regeneration roughly ten times cheaper after the first call. The audience set the cadence, not the budget.

Raw Questions Kept As Research

Every question an audience asks is a demand signal: an unfiltered record of what people do not understand, want or fear. At low volume that record is anecdote. At scale it is market research that the audience volunteers for free, and discarding the raw questions after summarizing them would destroy the most valuable output the system produces.

The collapse is lossy by design, which is right for the moment on stage and wrong for the quarter that follows. Queuena.com keeps the raw questions and merge decisions as a corpus of their own, so a host can later ask what an audience missed. Retrieval before generation (Lewis et al., 2020) is as useful over a question history as over any other corpus.

Three Decisions Before Going Live

  1. Set the hard cap before writing a prompt. Pick the number of umbrella questions a person can act on, and force the system to prioritize into it rather than expand past it.
  2. Cache the stable prefix before scaling. Constant re-summarizing stays affordable.
  3. Trigger regeneration on arrival, not a clock. A live burst refreshes the list immediately, and an idle channel costs nothing at all in the time between questions.

One Host Answers The Whole Room

The 95% comes from treating summarization as a decision aid, not a compression exercise. Its job is to let one person act on a whole crowd within seconds: collapse hard, cap it tight, and cache so the collapse is cheap enough to run continuously. The merge policy and the cache strategy are architecture, settled well before the first question arrives.

That is one contained workflow, exactly the kind of work an engineering partnership is scoped to deliver — acceptance targets written down before work starts, the system is run against them before you accept it, and then kept running by the person who built it. The flood is not the problem. The flood is the raw material, and the collapse is the product.

References

  1. Anthropic. Prompt Caching. Anthropic Documentation, 2024.
  2. Liu, N. F., et al. Lost in the Middle: How Language Models Use Long Contexts. arXiv, 2023.
  3. Lewis, P., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv, 2020.
  4. Anthropic. Building Effective Agents. Anthropic Engineering, 2024.

NEXT · TO PRODUCTION

Could this work for you?

Two minutes. Find out where you stand.

15 minutes · no charge · with Omar