Everyone pitched. Two weeks later, nothing had moved.
Marco van Hout · 4 min read
The final pitch round selects for the best presenter, and now that every team has a model writing the deck, polish carries no information. Score the last block on evidence and sponsor commitment instead: a runnable last 90 minutes, two concepts carried forward, one check on day 21.
The applause is the problem
Last afternoon of a three-day sprint. Five teams, five demos, a jury, a winner, drinks. Two weeks on, the winning concept sits in a shared folder and nobody can say who owns the next step, because nobody was asked.
Multi-day is the biggest single duration segment in the sessions designed on Metodic over the last 90 days, and a steady group of those designers build make-and-pitch formats across two and three days. Several of them already write owners, experiments or 30-day follow-ups into the session goal itself. Most agendas still treat the pitch as the finish.
What the pitch slot selects for
A final pitch round is a selection mechanism, and it selects on delivery. The calmest presenter and the tightest slides win, because a jury has eight minutes per team and delivery is what eight minutes can judge.
Generative AI has made this worse. In an August 2026 Harvard Business Review article, De Freitas, Israeli, Nave, Timoshenko and Toubia report that AI acts on the human bottlenecks inside the innovation process and often reinforces them. In ideation, models steer teams toward familiar ideas and make them more fixated on them. In screening, polished AI-generated pitches can be mistaken for better ideas. When every team has a model writing the script, polish carries no information. A slot that scores on polish is scoring noise. The authors also note that simulated customers miss the irrational behaviours that shape real adoption, and advise asking first whether the bottleneck is informational, judgment-based or incentive-driven, while keeping direct contact with real customers. For the last block: the jury scores evidence a model cannot fake, and the organisation commits something a model cannot deliver.
Do not let a model rank the concepts
The tempting shortcut is to have a model rank the five. An August 2026 arXiv preprint on LLMs as judges of novelty found their judgments diverge significantly from human consensus, with a systematic bias toward rating ideas "medium novel". It studied research ideas, not workshop concepts, so hold it lightly. Still, a judge that pulls everything toward the middle cannot pick two and stop three.
The last 90 minutes
These numbers are defaults, not research findings. Change them before the day, not during the block. Assumes five teams and a sponsor with budget authority present.
- 0-10 min — Evidence cards. Each team fills one A4 card, silent, no slides: the user problem in one sentence, how many real people they spoke to (minimum 3 to qualify), the one thing that surprised them, the cheapest next experiment, and its cost in days and euros.
- 10-40 min — Pitches, 4 minutes each plus 2 of questions. The first 2 minutes must cover evidence, not solution. Maximum 3 slides; anything AI-generated is labelled on the slide. Jury questions may only start with "What did you observe" or "What would it cost".
- 40-55 min — Scoring on three lines. Evidence strength (0-5), cost of the next experiment (cheaper scores higher, 0-5), and the sponsor's willingness to fund that experiment this month (0-5). Delivery is not a line. Scores are written before any discussion, then read aloud.
- 55-70 min — Sponsor commitment, live. Two concepts carry forward, no more. For each, the sponsor names an owner who is in the room, a ceiling for the first experiment (default 5 working days and 2,000 euros) and a check date 21 days out. All three go on the wall. No owner, no carry-forward; say so before the block starts.
- 70-80 min — Parking, with reasons. The other three concepts get one written sentence each from the jury on why they stop. The team signs the card. This stops the corridor lobbying of the following week.
- 80-90 min — Close. Each owner says in one sentence what will be done by the check date. The facilitator photographs the wall and sends it to everyone within the hour, check date in the subject line. No awards.
The first check: day 21
Twenty-one days, not thirty. Three weekly cycles: long enough for a cheap experiment, short enough that the sprint is still in memory. The check is a 30-minute call: sponsor, both owners, facilitator.
Default pass mark: the experiment ran with at least 5 real users or customers (not colleagues, not synthetic personas), and the owner can state one number they did not know on the last day of the sprint. Pass gets a second experiment and a second date. Fail is stopped in writing the same day. An experiment that never started is the most useful result: the bottleneck was incentive, not information, and that is a different conversation with the sponsor.
If you are designing a multi-day jam or sprint, describe the programme in Session Studio and let it draft the timed agenda.
Sources
- Research: The Innovation Problems AI Can't Solve — Harvard Business Review, August 2026
- Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty — arXiv preprint, August 2026
Design your own session
METODIC turns ideas like these into a complete session agenda with activities, timing, and materials — for workshops, meetings, offsites, and team sessions.