Why these moments — This Escalated FAST… He Pointed a Fake Gun at Cops and Regretted It!.mp4
Generated 2026-08-15T14:52:39
6 clips, the best scoring 50% above the weakest and drawn from 89% of its length.
6segments kept
30stotal length
65.4%of the source
What this run found
The video
Mostly person on screen throughout.
Movement
Movement bursts and then settles in about half the clips.
It settles as the loudest point arrives, not before it.
Sound
The kept moments sit about as loud as the video usually is — loudness is not what marks them out.
Face expression
A face was readable through only part of the video, so this describes that part and not the rest.
The cut
What was kept, where it came from, and what the scoring did to get there. Measured, all of it.
Where the clips came from
0:000:220:45
Loudness across the whole video, on the same time axis as the strip above — shown for context whether or not audio contributed points.
scorekeptscored well, not kept
Across the clips In 3 of 4 clips the loudest point and the motion peak land within 2s of each other.
The cut, end to end
The same clips as above, but on the output's clock rather than the source's — bar width is each clip's share of the highlight, height is what it scored. Use the left column to find a moment while watching the rendered file; the right column is where it came from, which is what every other timestamp in this report uses.
The video itself rather than the cut — how it divides, how it sounds, and what the expression reading does across it.
Shoot sessions
1 lighting setup(s) separate at silhouette 0.00, giving 1 sessions.
Span
Length
Setup
0:00:00 – 0:01:01
1.0 min
setup 0
What this does not establish. That a session is a different day. Two setups may be two locations, two cameras, or one space relit — nothing in the picture separates those from two dates. What it does establish is that the frame-wide colour reference differs between sessions, so any comparison that crosses a boundary compares capture conditions as much as content.
The video in chapters
Each stretch below is described by ollama/qwen2.5vl:7b, reading frames and the transcript — those paragraphs are a reading, not a measurement, and everything under them is what they were written from. The boundaries and every figure were computed before any model saw them, and the words beside each title are arithmetic over the transcript rather than anybody's reading of it.
The video was not divided — nothing in its shot structure separated one stretch from another. 82 words transcribed, 99% of the runtime, spread across every chapter.
share of the cut, against the chapter's share of runtime
In this stretch, a video is presented as if taken from a dashcam, showing the aftermath of a tense confrontation outside a house. The scene depicts a person with tattooed arms holding a black object, which appears to be a replica gun, near an open door. The individual is seen switching the replica gun between their hands, while the background shows a residential setting with a clear blue sky and some greenery. The video seems to be documenting the incident following a police shooting of a person named Baldenado, who was later seen smiling and holding the replica gun. The footage also captures a woman inside the house yelling that the object was a toy, a warning that the officers did not hear. The individual's history of developmental disabilities and prior police encounters is noted, and the scene suggests that the incident is under investigation to determine if proper procedures were followed. The video appears to be a compilation of different moments from the event, possibly showing different angles or perspectives of the incident.
Mostly person — on screen for 100% of this chapter's detected seconds.
99% of it is speech, 107 words a minute.
One voice throughout.
0:00:00SPEAKER_01Officer Dominic Sanchez, perceiving the weapon as a real threat, fired three shots at Baldenado,
0:00:22SPEAKER_01Later body cam footage revealed Baldenado smiling and switching the replica gun between
everything said here (8 lines)
0:00:00SPEAKER_01Officer Dominic Sanchez, perceiving the weapon as a real threat, fired three shots at Baldenado,
0:00:19SPEAKER_01striking him in the arm, leg, and abdomen.
0:00:22SPEAKER_01Later body cam footage revealed Baldenado smiling and switching the replica gun between
0:00:27SPEAKER_01hands just before the shooting, while a woman inside yelled that it was a toy, a warning
0:00:32SPEAKER_01the officers didn't hear.
0:00:34SPEAKER_01Baldonado, who has a history of developmental disabilities and prior police encounters,
0:00:38SPEAKER_01is currently recovering in the hospital.
0:00:41SPEAKER_01The incident is under investigation to determine if proper procedures were followed.
How the expression reading moves
A face was readable in 17% of the video (8s). Across those seconds the classifier reported happy 75%, neutral 12%, surprise 12%.
0% of the read seconds carry a negative-valence label and 75% a positive one. A further 12% read as surprise, which carries no direction and is counted on neither side.
Caution: only 17% of the video had a readable face, so every share here describes that part of it and not the rest.
All of the above describes labels a five-class classifier assigned to faces it could see — not what anyone felt. It has no notion of intensity, degrades on profile, occlusion and motion, and cannot distinguish a performed expression from a felt one. Treat it as a map of where to look, not as a finding about a person.
positive-readingnegative-reading
The advisor
Everything above is what this run measured. Everything here is what to do about it — including what the run could not have measured however it was configured.
Said here, measured nowhere
Lines this page already quotes that share no word with any class or event this run produced. The report prints them and says nothing about them, which reads exactly like a claim that was checked and held — so it is said here instead. Ranked by how much vocabulary unusual for this video each line carries; nothing below knows what any of them mean.
0:00:22 SPEAKER_01: “Later body cam footage revealed Baldenado smiling and switching the replica gun between”
chapter 1 · 13 words · picked out by · the detector was labelling person there
What it would take to measure one of them. Which route fits depends on what the thing actually is, which you can see and this report cannot — so each one carries the condition it holds under, and the case where it will waste your time.
Costs nothing to rule out
Compose it from classes this video already produced
Only possible when what was said is an arrangement of things this detector already finds — one inside another, several at once, none of them present. Ask the advisor to draft the rule — “Check something that was said…” in the AI summary menu. When the claim cannot be built from the classes this video produced it says so, and that answer is one model call.
Then, before committing
Before spending a session on labels, spend five minutes finding out whether you have to. Run the open-vocabulary detector over a short window with your query and a control query for something ordinary you know is in the shot. If the control scores well and yours does not, no threshold will save it and the trained class is the answer. If both score, you have just avoided the session.
minutes. Point at a few frames and name them; no dataset, no labels, no GPU. The first search over a video pays for one embedding pass, and every later query is free. Gives you a score per sampled second — how much this moment looks like the frames you pointed at. No box, so nothing to count.
Right route when you can point at frames that show it. This is the only route that works when you can recognise something on sight but cannot put it into words — not when it is small in the frame. The embedding describes the scene, so a few percent of the picture is drowned by everything around it — that case needs a detector.
Most reliable
Train a class of your own
a session, and a GPU. Collect frames from the conditions it fails in, label them, train, export. A third of the set should be frames containing whatever it will confuse for the target, with nothing boxed. Gives you boxes at frame rate: countable, usable in composition rules, usable live, and reusable on every video you analyse after.
Right route when it matters enough to justify the labelling, or nothing cheaper could see it — not when your labels do not contain the conditions it fails in. Ten frames of the case that breaks it beat a thousand more of what already works.
Meanwhile, for nothing
Score the moments where it is talked about
instant — a transcript keyword weight and a re-score, no re-analysis. Gives you the seconds the transcript says it, which is not the same thing and must not be reported as if it were.
Right route when you want the moments about it now, while deciding whether one of the routes above is worth the time — not when you need to know when the thing itself was on screen.
What to try next
Worked out from this run's own numbers — each point below is backed by the figures shown with it, not by a guess about what you meant.
mediumThe highlight came out shorter than you asked for
You asked for up to 46s and got 30s. In MAX mode the cut stops when it runs out of moments that scored anything at all — not when it runs out of budget.
Try: Lower the detector thresholds so more moments score, give another signal a weight, or accept the shorter cut: padding it means including moments nothing was detected in.
mediumSomething was said that this run has no measurement for
At 0:00:22, in chapter 1: "Later body cam footage revealed Baldenado smiling and switching the replica gun between" No class or event this run produced shares a word with that line, so the report quotes it and says nothing about it — neither confirming nor contradicting it. The detector was busy in that stretch — it labelled person there — so this is not a gap in coverage but a gap in vocabulary.
Try: First, for nothing: ask the advisor to draft a rule for the line. It answers that the claim cannot be built from the classes this video has, when it cannot, and that rules out the cheapest route in one call. Fastest that would actually measure it: teach a category from example frames — minutes. Point at a few frames and name them; no dataset, no labels, no GPU. The first search over a video pays for one embedding pass, and every later query is free. Right route when you can point at frames that show it. This is the only route that works when you can recognise something on sight but cannot put it into words; not when it is small in the frame. The embedding describes the scene, so a few percent of the picture is drowned by everything around it — that case needs a detector. Most reliable: train a class of your own — a session, and a GPU. Collect frames from the conditions it fails in, label them, train, export. A third of the set should be frames containing whatever it will confuse for the target, with nothing boxed. Worth it when it matters enough to justify the labelling, or nothing cheaper could see it. Before spending a session on labels, spend five minutes finding out whether you have to. Run the open-vocabulary detector over a short window with your query and a control query for something ordinary you know is in the shot. If the control scores well and yours does not, no threshold will save it and the trained class is the answer. If both score, you have just avoided the session.
mediumScene changes are switched on but never fired
Scene changes are worth 5 points, yet contributed nothing to any kept moment. The detector either found nothing or found it below its confidence threshold.
Try: Check that this detector is actually running and that its threshold is not set above what your footage produces. If it genuinely has no class for what you are looking for, teach a category from examples instead of raising the weight.
mediumTranscript keywords are switched on but never fired
Transcript keywords are worth 2 points, yet contributed nothing to any kept moment. The detector either found nothing or found it below its confidence threshold.
Try: Check that this detector is actually running and that its threshold is not set above what your footage produces. If it genuinely has no class for what you are looking for, teach a category from examples instead of raising the weight.
mediumActions are switched on but never fired
Actions are worth 5 points, yet contributed nothing to any kept moment. The detector either found nothing or found it below its confidence threshold.
Try: Check that this detector is actually running and that its threshold is not set above what your footage produces. If it genuinely has no class for what you are looking for, teach a category from examples instead of raising the weight.
mediumMotion events were found but count for nothing
81 motion events were detected across this video and are weighted at zero, so none of them could influence the 6 moments that were chosen.
Try: Give motion events a small weight — 3 to 5 points is usually enough to break ties without taking over — and re-run. Detection is cached, so it costs a re-score, not a re-analysis.
lowNearly every clip contains the same thing
'person' appears in 5 of 6 clips (83%). The highlight is likely to feel repetitive even though each moment scored well on its own.
Try: Selection has no notion of variety — it takes the highest scores, and if they all look alike, so will the cut. Spread the cut with the coverage slider, or swap individual clips for the alternatives that lost to them.
The record
What every claim above was computed from, clip by clip. Nothing here is a reading; it is the arithmetic, kept so any of it can be checked.
What decided the cut
Audio peak110
Motion peak25
Objects25
Expression10
The moments, in order
One row per clip that was kept. The bars are the points that second earned, broken down by signal; the tags are what was detected there, grouped by what produced them — objects come straight from the detector, events are combinations the composition rules recognised, actions come from the action model with their confidence.
The paragraph inside each card was written by ollama/qwen2.5vl:7b from that clip's own frames, 2026-08-15T15:04:35. Those paragraphs are a reading, not a measurement — every figure beside them was computed before any model saw the footage.
person
boxes as detected at the peak second
Clip 10:04 – 0:09 · 24 points
peak at 0:06 · 5s long · from Chapter 1
Audio peak15
Objects5
read from this clip
The video starts with a view from inside a house, looking out through a glass door. A person is standing outside, holding a gun, and appears to be aiming it towards the house. The scene is viewed from a slightly elevated angle, with a hand partially visible in the foreground, holding the camera. The outdoor environment is bright and sunny, with a clear sky and some greenery visible in the background. The person outside seems to be moving slightly, adjusting their stance and aiming the gun more directly. The overall scene suggests a tense situation, with the person outside seeming to be in a confrontational or threatening stance.
Sound
Loudest at 0:07, a little above the video's usual level (+6 dB). Nothing was labelled at that second.
Face expression
Expression here is mostly neutral over 1 readable second, reading more negative than the video's own average (+0.00 against +0.67).
Summary
One of the weaker clips that still made the cut, chosen on a rise in sound and what was on screen, though not at the same instant. A little louder than this video usually runs and with more movement than this video usually carries.
vs the video+loudness +6 dB on the video's median-expression reading -0.67 valence on the video's mean, mostly neutral-person on screen 6th percentile for size in this video
show the measurements
Both landed together — movement stopped and the loudest point arrived within 1s of each other.
Peak 0:07 at -13.0 dBFS · +6.0 dB vs the video's middle — nothing labelled there
scored above 60% of the video · peaked at -7 dBFS (100th pct) · signals 4s apart · best detection 0.57
objectsperson
said here
SPEAKER_01Officer Dominic Sanchez, perceiving the weapon as a real threat, fired three shots at Baldenado,
Loudness through the clip — | marks 0:06, the second that scored highest. Volume is drawn for context only and contributed 15 points, so a louder stretch elsewhere in the clip did not move it.
plays the source file in place
Multi-signal boost — 2 signals agreed, ×1.2 (+4)
person
boxes as detected at the peak second
Clip 20:12 – 0:17 · 48 points
peak at 0:14 · 5s long · from Chapter 1
Motion peak10
Audio peak20
Objects5
Expression5
read from this clip
The video transitions from an outdoor scene where a person is holding a gun and aiming it, to a close-up of a tattooed arm gripping a handgun. The frames show a steady hand with a detailed tattoo, emphasizing the action of handling the firearm. The setting appears to be a residential area, with trees and a house visible in the background. The scene then shifts to a still image of a man in a dark uniform, with text identifying him as Officer Dominic Sanchez, suggesting the context of a police incident captured on DashCam. The image is static with no movement or sound.
Movement
Movement spiked at 0:13 and dropped away after — the burst-then-stillness this signal scores. 2 of them fall inside this clip.
Sound
Loudest at 0:14, about as loud as the video usually is. person was on screen at that second.
Face expression
Expression here is mostly happy over 3 readable seconds, reading more positive than the video's own average (+0.98 against +0.67).
On screen
Person fills 30.6% of the frame — more than in 76% of this video's other 5s stretches containing one, though frame share also rises when the camera simply moves closer.
Summary
The strongest clip in this highlight, chosen on a burst of movement, a rise in sound, what was on screen and a facial expression, though not at the same instant. With more movement than this video usually carries.
vs the video=loudness +1 dB on the video's median+expression reading +0.30 valence on the video's mean, mostly happy+person on screen 76th percentile for size in this video
show the measurements
Both landed together — movement stopped and the loudest point arrived within 1s of each other.
Peak 0:14 at -18.3 dBFS · +0.7 dB vs the video's middle — person
scored above 99% of the video · peaked at -10 dBFS (99th pct) · signals 2s apart · best detection 0.86
objectsperson
said here
SPEAKER_01Officer Dominic Sanchez, perceiving the weapon as a real threat, fired three shots at Baldenado,
Loudness through the clip — | marks 0:14, the second that scored highest. Volume is drawn for context only and contributed 20 points, so a louder stretch elsewhere in the clip did not move it.
plays the source file in place
Multi-signal boost — 4 signals agreed, ×1.2 (+8)
person
boxes as detected at the peak second
Clip 30:18 – 0:23 · 30 points
peak at 0:18 · 5s long · from Chapter 1
Audio peak15
Objects5
Expression5
read from this clip
The video shows a tattooed arm gripping a handgun, with the hand steady and the focus on the detailed tattoo. The frames capture the arm from slightly different angles, emphasizing the tattoo's intricate design. There is no movement of the arm or gun, suggesting a close-up shot of the person holding the weapon. The frames are steady, with no noticeable action or changes occurring between them.
Sound
Loudest at 0:18, about as loud as the video usually is. person was on screen at that second.
Face expression
Expression here is mostly happy over 1 readable second, reading more positive than the video's own average (+0.97 against +0.67).
Summary
Middling for this highlight, chosen on a rise in sound, what was on screen and a facial expression, all landing on the same second.
vs the video=loudness +2 dB on the video's median+expression reading +0.30 valence on the video's mean, mostly happy=person on screen 55th percentile for size in this video
show the measurements
Peak 0:18 at -17.2 dBFS · +1.8 dB vs the video's middle — person
scored above 82% of the video · peaked at -11 dBFS (97th pct) · signals landed together · best detection 0.86
objectsperson
said here
SPEAKER_01Officer Dominic Sanchez, perceiving the weapon as a real threat, fired three shots at Baldenado,
SPEAKER_01striking him in the arm, leg, and abdomen.
SPEAKER_01Later body cam footage revealed Baldenado smiling and switching the replica gun between
Loudness through the clip — | marks 0:18, the second that scored highest. Volume is drawn for context only and contributed 15 points, so a louder stretch elsewhere in the clip did not move it.
plays the source file in place
Multi-signal boost — 3 signals agreed, ×1.2 (+5)
person
boxes as detected at the peak second
Clip 40:25 – 0:30 · 36 points
peak at 0:27 · 5s long · from Chapter 1
Audio peak25
Objects5
read from this clip
The video shows a close-up of a tattooed arm holding a replica gun. The tattoo is prominently displayed on the arm, with intricate designs. The arm is steady, and the focus is on the detailed tattoo and the grip on the replica gun. In the background, there are trees and a building, indicating the setting is outdoors. The frames capture the arm from slightly different angles, emphasizing the tattoo's detail. The frames also show text indicating that the person is switching the replica gun between hands just before a shooting incident. The context suggests this is part of a larger narrative involving a warning about the replica gun being mistaken for a real one.
Sound
Loudest at 0:29, about as loud as the video usually is. person was on screen at that second.
On screen
Person fills 31.9% of the frame — more than in 94% of this video's other 5s stretches containing one, though frame share also rises when the camera simply moves closer.
Summary
Middling for this highlight, chosen on a rise in sound and what was on screen, though not at the same instant.
vs the video=loudness +1 dB on the video's median+person on screen 94th percentile for size in this video
show the measurements
Peak 0:29 at -17.5 dBFS · +1.4 dB vs the video's middle — person
scored above 94% of the video · peaked at -10 dBFS (98th pct) · signals 2s apart · best detection 0.74
objectsperson
said here
SPEAKER_01Later body cam footage revealed Baldenado smiling and switching the replica gun between
SPEAKER_01hands just before the shooting, while a woman inside yelled that it was a toy, a warning
Loudness through the clip — | marks 0:27, the second that scored highest. Volume is drawn for context only and contributed 25 points, so a louder stretch elsewhere in the clip did not move it.
plays the source file in place
Multi-signal boost — 2 signals agreed, ×1.2 (+6)
person
boxes as detected at the peak second
Clip 50:34 – 0:39 · 36 points
peak at 0:36 · 5s long · from Chapter 1
Motion peak5
Audio peak20
Objects5
read from this clip
The video shows a close-up view of a tattooed arm holding a replica gun, with the tattoo prominently displayed. The camera remains steady, focusing on the detailed design of the tattoo. In the background, a person appears, standing outside a door, partially visible through a glass pane. The camera then shows a hand holding a device, possibly a remote or a tool, against the door. The scene suggests a tense or cautious interaction, likely related to the individual mentioned earlier in the video. The person appears to be looking or reaching through the door, and the camera captures this action from a slightly different angle, with the person's hand and the device being more prominent.
Movement
Movement spiked at 0:36 and dropped away after — the burst-then-stillness this signal scores.
Sound
Loudest at 0:34, about as loud as the video usually is. Nothing was labelled at that second.
Face expression
Expression here is mostly happy over 1 readable second, reading more negative than the video's own average (+0.51 against +0.67).
On screen
person appears without a cut, so it came into a frame that was already running.
Summary
Middling for this highlight, chosen on a burst of movement, a rise in sound and what was on screen, though not at the same instant. In order: the loudest point arrives, person comes on screen in the same moment and movement drops away in the same moment — with less movement than this video usually carries.
vs the video=loudness +2 dB on the video's median-expression reading -0.16 valence on the video's mean, mostly happy-person on screen 19th percentile for size in this video
show the measurements
In order: the loudest point arrives at 0:34, person comes on screen 2s after and movement drops away in the same second.
Peak 0:34 at -17.5 dBFS · +1.5 dB vs the video's middle — nothing labelled there
scored above 94% of the video · peaked at -10 dBFS (98th pct) · signals 2s apart · best detection 0.39
objectsperson
said here
SPEAKER_01the officers didn't hear.
SPEAKER_01Baldonado, who has a history of developmental disabilities and prior police encounters,
SPEAKER_01is currently recovering in the hospital.
Loudness through the clip — | marks 0:36, the second that scored highest. Volume is drawn for context only and contributed 20 points, so a louder stretch elsewhere in the clip did not move it.
plays the source file in place
Multi-signal boost — 3 signals agreed, ×1.2 (+6)
Clip 60:40 – 0:45 · 30 points
peak at 0:40 · 5s long · from Chapter 1
Motion peak10
Audio peak15
read from this clip
The video shows a view through a glass door, focusing on a person standing outside on a driveway, partially visible through the reflection. The individual is wearing a black t-shirt and blue shorts and appears to be reaching towards the door. The scene is calm, with a bright blue sky and some trees visible in the background. The door itself has a white frame with decorative glass panels. The camera remains stationary, capturing the interaction without any sudden movements.
Movement
Movement spiked at 0:40 and dropped away after — the burst-then-stillness this signal scores.
Sound
Loudest at 0:44, about as loud as the video usually is. Nothing was labelled at that second.
On screen
person appears without a cut, so it came into a frame that was already running.
Summary
Middling for this highlight, chosen on a burst of movement and a rise in sound, though not at the same instant. In order: movement drops away, person comes on screen in the same moment and the loudest point arrives a moment after that — with more movement than this video usually carries.
vs the video=loudness -1 dB on the video's median-person on screen 10th percentile for size in this video
show the measurements
In order: movement drops away at 0:40, person comes on screen 1s after and the loudest point arrives 3s after that.
Peak 0:44 at -20.3 dBFS · -1.4 dB vs the video's middle — nothing labelled there
scored above 82% of the video · peaked at -11 dBFS (97th pct) · signals 3s apart
said here
SPEAKER_01is currently recovering in the hospital.
SPEAKER_01The incident is under investigation to determine if proper procedures were followed.
Loudness through the clip — | marks 0:40, the second that scored highest. Volume is drawn for context only and contributed 15 points, so a louder stretch elsewhere in the clip did not move it.
plays the source file in place
Multi-signal boost — 2 signals agreed, ×1.2 (+5)
Scored well, but not included
The highest-scoring moments that did not make the cut — usually because the highlight was already full, or a neighbouring second scored higher. Raise the weight of a signal below to pull moments like these in.
What the run observed, what was only asserted, and what it could not determine — kept apart, because running them together is how the third quietly becomes the first.
What this run observed. This run observed: person (object) from 0:00:04 to 0:00:39, in 5 of the kept moments.
What was said, and by whom. Everything else below was said by one speaker over 99.2% of the runtime. It is that speaker's account of the material, not something this run observed, and it may describe events outside the footage entirely.
Time
Said
0:00:00
Officer Dominic Sanchez, perceiving the weapon as a real threat, fired three shots at Baldenado,
0:00:19
striking him in the arm, leg, and abdomen.
0:00:22
Later body cam footage revealed Baldenado smiling and switching the replica gun between
0:00:27
hands just before the shooting, while a woman inside yelled that it was a toy, a warning
0:00:32
the officers didn't hear.
0:00:34
Baldonado, who has a history of developmental disabilities and prior police encounters,
0:00:38
is currently recovering in the hospital.
0:00:41
The incident is under investigation to determine if proper procedures were followed.
What this run could not determine. Listed rather than left to silence, because a report that stops at what it found invites its gaps to be read as absence.
what was said within the material — the audio is one speaker over 99.2% of the runtime, which reads as narration about the material rather than sound from inside it. Anything it asserts is that speaker's account, not an observation of this run
what any person knew, intended, perceived or felt — no detector observes an internal state, and a label for an appearance is not a reading of one
whether anything seen is what it appears to be — a detector reports appearance, and what a thing actually was is a fact about the world
why one observation followed another — the order and the interval are measured here, and neither is a cause
anything outside the frame, or outside the span that was analysed
anything the silent detectors would have covered: action, keyword — each ran or was weighted at nothing and found nothing, which is not evidence that there was nothing to find
anything in the 34.6% of the source that was not kept — the order below is the order within the selection