60 DramaBox clips · gemini-3.8-flash, asked blind —
the model was never told which burst the scene was written to produce ·
full 83-class taxonomy, 1–3 labels per span, confidence per span ·
our detector's verdict beside it on the same time axis.
What you are judging. The statistics below say how often Gemini's labels agree with the requested class and with our detector. They cannot say whether the labels are right. That is a listening question, and this page exists so you can answer it. Play a clip, read the words, look at where Gemini put its spans and what it called them, and decide whether you would be happy training a locator and a classifier on that.
Shriek (detector: 0)| scene class | Gemini top label | Gemini any of 1–3 | Gemini scream-family | detector strict | detector family | |||
|---|---|---|---|---|---|---|---|---|
| n | rate | n | rate | n | rate | n | n | |
| shriek scenes | 14/30 | 46.7 % | 23/30 | 76.7 % | 29/30 | 96.7 % | 0/30 | 6/30 |
| scream scenes | 28/30 | 93.3 % | 30/30 | 100.0 % | 30/30 | 100.0 % | 3/30 | 3/30 |
| all 60 | 42/60 | 70.0 % | 53/60 | 88.3 % | 59/60 | 98.3 % | 3/60 | 9/60 |
Family = {Scream, Shriek, Mournful Wail, yell, wail, shout…}, the same
table the detector's family reading uses, so the last two columns are directly comparable.
Shriek and Scream share a family, so the family column cannot tell the
two apart — that is what the strict and top-label columns are for.
Its spans are wide. There is no ground-truth onset for these clips, so the timing check is detector-independent: does the clip's own loudest 300 ms window — computed from the waveform, no model involved, drawn as the pink line on every card — fall inside an annotated span, and how much of the clip did the annotator cover to catch it?
| coverage of the clip | loudest moment inside a span | lift over chance | |
|---|---|---|---|
| Gemini, all spans | 38.2 % | 48.3 % | 1.27× |
| Gemini, scream-family spans only | 28.5 % | 46.7 % | 1.64× |
| our detector | 6.5 % | 23.3 % | 3.57× |
Gemini is far better at what; our detector is nearly three times more precise about where, per second of audio it commits to. Good enough to retrain a classifier on; too coarse, as it stands, to retrain a locator on.
It almost never says “nothing here”. Across these 60 clips Gemini set
no_burst zero times — which proves nothing on its own, because all 60
were written to contain a burst. So 20 hard negatives were cut from the longest stretches Gemini
itself left unannotated, and re-submitted with the identical question. It answered
“no burst” on 1 of 20, and 11 of 20
came back with Scream or Shriek as the top label — in audio the same
model had just declared empty. With the real scream present it reserves Scream for it
and ignores the shouted dialogue; with the scream cut away, the shouted dialogue becomes the scream.
The labels are context-dependent and not stable under re-segmentation: a re-annotation must
run on whole clips and cut training rows out of the resulting spans, never annotate short rows
directly.
Blue = Gemini, amber = our detector, pink line = the clip's own loudest 300 ms window (computed from the waveform, no model involved). Default order puts the biggest disagreements first.