← Back to Articles

Turning A Crowd's Questions Into A Few Answers

You open a channel for an audience to ask you anything, and within minutes you have a wall. The questions overlap, repeat, and contradict each other in wording while asking the same thing underneath. You cannot read them live, and nobody wants to scroll them.

What you actually have is a handful of questions wearing hundreds of costumes. The job is not to display the flood or to answer it item by item. The job is to collapse it — down to the few questions that, once answered, cover everything anyone asked.

A moderator sorting audience question cards at a conference table
A thousand questions are a handful in disguise.

Merge Intent, Not Wording

The first decision is what "the same question" means. Two questions phrased nothing alike are often the same question, and two questions sharing most words are often different ones. Collapsing on surface text produces false merges and false splits at once.

Merge on intent — on what the asker wants — and merge hard. Target roughly five umbrella questions and never exceed ten, no matter how many arrive underneath. That cap is not a performance limit. It is the summary's spine: a system with room for forty umbrellas will produce forty, and forty umbrellas is the flood with better formatting.

Merge aggressively. Four well-merged umbrellas beat eight narrow ones, because the presenter can hold four in their head and act on them.

A summary that faithfully preserves every nuance has failed at the only job that it had. Stuffing the whole accumulated set of questions into one long prompt does not rescue it either — long contexts bury information in the middle (Liu et al., 2023), so the fidelity that you thought you were buying is spent on tokens the model reads least carefully.

A wrong merge and a wrong split fail differently, and the asymmetry decides how the policy should lean. A split that should have merged shows up on screen as two near-identical umbrellas, and the presenter fixes it by answering both at once. A merge that should have split is invisible: the question that got absorbed never appears, and the person who asked it watches their question go unanswered while the summary looks complete.

So the model's job is deliberately narrow. It proposes the umbrella structure; the cap, the ordering, and the merge policy are code, and none is negotiable by the model. Keeping a model's scope tight and its authority explicit (Anthropic, 2024) is what makes the output predictable enough to put in front of a live audience without a human checking every refresh.

Regeneration Cost Decides Viability

The summary is worthless if it is stale, which means that it has to regenerate as questions arrive. Regenerating every five new questions is arrival-triggered rather than clock-triggered, and it is the right trigger for this workload: a live talk that draws a burst of questions gets a fresh summary immediately, while a poster in a lobby collecting a slow trickle over a month does not pay to recompute a summary that nothing new has changed.

Now look at what that cadence costs. Regenerating every fifth question means the summarizer runs constantly during the moments that matter most — and a naive implementation re-sends the entire system prompt and the whole accumulated context on every single call. The arithmetic here is not subtle. Continuous re-summarization becomes a feature you have to ration, and a rationed summary is a stale summary.

The fix is to cache what does not change. Prompt caching that reuses a fixed prefix across calls (Anthropic, 2024), applied to the system prompt, makes regenerations roughly ten times cheaper after the first call. That ratio moves continuous regeneration from a cost problem to a non-issue, and it is why the cadence can be set by the audience rather than by the budget. Whether the prefix is cached is an architecture decision made before the first question arrives, and cheaper to make then than to retrofit under load.

When The Questions Are The Asset

There is a boundary where the summary stops being the point. Every question an audience asks is a demand signal — an unfiltered record of what people do not understand, want, or fear. At low volume that record is anecdote. At scale it is market research the audience is volunteering, and the moment it crosses that line, throwing away the raw questions after summarizing them destroys the most valuable output the system produces.

The collapse is lossy by design, and lossy is correct for the moment on stage. It is wrong for the quarter that follows. A system built to serve only the live moment will discard the corpus that answers the strategic question — what this audience does not understand about the thing you are selling them — and the discard is invisible, because the summary still looks great. Keep the raw questions, keep the merge decisions, and treat the accumulated corpus as a retrievable asset in its own right: retrieval before generation (Lewis et al., 2020) is as useful over your own question history as it is over any other corpus.

First Steps

  1. Set the hard cap before you write a prompt — pick the number of umbrella questions a human can act on, and force the system to prioritize into it rather than expand past it.
  2. Cache the stable prefix of every regeneration before you scale, so continuous re-summarization is affordable rather than rationed into staleness.
  3. Trigger regeneration on question arrival rather than a clock, so a live burst refreshes immediately and an idle surface costs nothing between questions.

Summarize For Decision, Not Fidelity

Summarization is a decision aid, not a compression exercise. Its job is to let one human act on a crowd in seconds, which means collapsing hard, capping tight, and caching so the collapse is cheap enough to run continuously. Fidelity to every distinction and usefulness under a hard cap pull in opposite directions, and the product has to pick a side out loud.

The merge policy and the cache strategy are both architecture, not prompt tweaks, and both are settled before the first question arrives. The difference between a wall of text and a decision a presenter can act on is a merge policy, a hard cap, and a cached prefix. That is one contained workflow, exactly what an engineering partnership is scoped to deliver — acceptance targets written into the SOW before work starts, the system run against those targets before you accept it, and kept running by the person who built it. The flood is not the problem. The flood is the raw material, and the collapse is the product.

References

  1. Anthropic. Prompt Caching. Anthropic Documentation, 2024.
  2. Liu, N. F., et al. Lost in the Middle: How Language Models Use Long Contexts. arXiv, 2023.
  3. Lewis, P., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv, 2020.
  4. Anthropic. Building Effective Agents. Anthropic Engineering, 2024.

NEXT · TO PRODUCTION

Find the real blocker.

What is slowing delivery, and the fastest path through.

15 minutes · no charge · with Omar