Can Gemini annotate our vocal bursts?

60 DramaBox clips · gemini-3.8-flash, asked blind — the model was never told which burst the scene was written to produce · full 83-class taxonomy, 1–3 labels per span, confidence per span · our detector's verdict beside it on the same time axis.

What you are judging. The statistics below say how often Gemini's labels agree with the requested class and with our detector. They cannot say whether the labels are right. That is a listening question, and this page exists so you can answer it. Play a clip, read the words, look at where Gemini put its spans and what it called them, and decide whether you would be happy training a locator and a classifier on that.

53/60
clips where any of Gemini's 1–3 labels is the requested burst
3/60
clips where our detector said the requested burst
68
times Gemini said Shriek (detector: 0)
31
distinct labels used (detector: 8 of 83)
3.18
Gemini bursts per clip (detector: 1.22)
48.3 %
clips whose loudest moment falls inside a Gemini span (detector: 23.3 %)

Agreement with the requested burst

scene classGemini top labelGemini any of 1–3 Gemini scream-familydetector strictdetector family
nratenratenratenn
shriek scenes14/3046.7 %23/3076.7 %29/3096.7 %0/306/30
scream scenes28/3093.3 %30/30100.0 %30/30100.0 %3/303/30
all 6042/6070.0 %53/6088.3 %59/6098.3 %3/609/60

Family = {Scream, Shriek, Mournful Wail, yell, wail, shout…}, the same table the detector's family reading uses, so the last two columns are directly comparable. Shriek and Scream share a family, so the family column cannot tell the two apart — that is what the strict and top-label columns are for.

Two things the agreement table does not say

Its spans are wide. There is no ground-truth onset for these clips, so the timing check is detector-independent: does the clip's own loudest 300 ms window — computed from the waveform, no model involved, drawn as the pink line on every card — fall inside an annotated span, and how much of the clip did the annotator cover to catch it?

coverage of the cliploudest moment inside a spanlift over chance
Gemini, all spans38.2 %48.3 %1.27×
Gemini, scream-family spans only28.5 %46.7 %1.64×
our detector6.5 %23.3 %3.57×

Gemini is far better at what; our detector is nearly three times more precise about where, per second of audio it commits to. Good enough to retrain a classifier on; too coarse, as it stands, to retrain a locator on.

It almost never says “nothing here”. Across these 60 clips Gemini set no_burst zero times — which proves nothing on its own, because all 60 were written to contain a burst. So 20 hard negatives were cut from the longest stretches Gemini itself left unannotated, and re-submitted with the identical question. It answered “no burst” on 1 of 20, and 11 of 20 came back with Scream or Shriek as the top label — in audio the same model had just declared empty. With the real scream present it reserves Scream for it and ignores the shouted dialogue; with the scream cut away, the shouted dialogue becomes the scream. The labels are context-dependent and not stable under re-segmentation: a re-annotation must run on whole clips and cut training rows out of the resulting spans, never annotate short rows directly.

The clips

Blue = Gemini, amber = our detector, pink line = the clip's own loudest 300 ms window (computed from the waveform, no model involved). Default order puts the biggest disagreements first.