Qwen, Gemma and Marlin Captioned My CCTV Footage

Qwen, Gemma and Marlin Captioned My CCTV Footage

I’m building Triloch, a search engine for my home CCTV recordings that runs entirely on my own machines. The funny thing is that the search never watches any video. A vision model watches each eight-second clip and writes a caption, and the search only ever reads those captions. So if no caption mentions the coffee cup, no amount of clever searching will ever find me drinking coffee.

How Triloch searches a clip An eight-second clip becomes 16 frames. A vision model reads the frames and writes a caption. A search box points at the caption, because search reads only the caption text and never the video. ONE 8-SECOND CLIP 16 frames, two per second vision model CAPTION (EXAMPLE) A person lifts a white cup to their mouth and drinks, then sets it back on the desk. When did I drink coffee? search reads only the caption How Triloch searches a clip An eight-second clip becomes 16 frames. A vision model reads the frames and writes a caption. A search box points at the caption, because search reads only the caption text and never the video. ONE 8-SECOND CLIP 16 frames, two per second vision model CAPTION (EXAMPLE) A person lifts a white cup to their mouth and drinks, then sets it back on the desk. When did I drink coffee? search reads only the caption
How search sees a clip. The caption wording is an example, not real model output.

That makes the captioning model the most important choice in the whole system, because it quietly decides what I can find later. In part 3 I built a set of questions with marked answers, and learned that a search score can look decent while telling you very little about whether a caption described what actually happened. So this time I wanted to skip the scores and read the captions themselves.

I’d been running two small models, Qwen and Marlin, and I was curious whether a newer one, Gemma, would do better. I gave all three the same clips and the same three prompts, then checked every caption against what really happened in the room. When the results came back shaky, I kept going, first with a bigger Qwen and then with a brand new prompt.

Here’s the ending up front, so you’re not waiting for a winner: nobody won. Gemma got dropped, the bigger model never managed a single caption, and the new prompt made the other two worse. After two weeks of this, I stopped testing models and went back to building the actual search.

Same clips, same prompts

For a fair fight, every model had to see exactly the same thing. I picked twelve eight-second clips from September footage I’d already recorded, and earlier runs had cut each one down to 16 frames, two per second. Every model got those exact frames.

Everything ran locally on my M3 Pro through MLXApple’s open-source machine learning framework for Apple silicon. It uses the Mac’s shared memory for both the CPU and GPU, so a laptop can run a model without a separate graphics card. I run the vision models through MLX-VLM, a package built on it.. Qwen3.5-2B and Marlin-2B used 8-bit weights, while Gemma 4 E4B used 4-bit weights, stored at lower precision to save memory, so the fight wasn’t perfectly level. There was supposed to be a fourth contender too, Molmo 2 8B, but it never wrote a word. It wanted to run code from its own download, and the library I use only hands it still images anyway.

Each model then got all three prompts, with temperature at zero and a 256-token cap. None of them hinted at anything I’d later search for:

  1. A short one, the prompt I’d been using: describe visible actions and objects in time order.
  2. A structured one with four headings: ongoing states, actions in time order, objects handled, and unclear details.
  3. A narrative one: describe the clip from beginning to end, including changes in posture.

I checked the 54 captions from six clips I knew well: me walking in, a trip to the bookshelf, a phone call, my head coming up off the desk, and some fidgeting with my glasses and face. I watched all six in full and answered questions about what happened in each, and a coding agent turned my answers into 25 facts. It froze that list before reading a single caption, so no fact could quietly bend to fit a model.

Then it went through the captions one by one and marked every fact correct, missed, contradicted, or unknown when the caption was too vague to tell, and a separate review pass checked those calls and fixed a few. A miss only counts against a model if the action was actually visible, so the agent also pulled the frames back out and looked. The phone at my ear, the stand and walk, my head coming up: several of the missed actions were sitting right there in the frames the models saw.

What each model wrote

The phone clip shows the difference best. I’m sitting at my desk, I lift my phone to my ear, and then I stand up and walk away. Here’s what each model wrote with the structured prompt:

Marlin

Scene: A high-angle, wide-shot captures a man in a grey t-shirt and dark shorts sitting in a black gaming chair at a wooden desk … The man is holding a white smartphone to his right ear, appearing to be in the middle of a conversation. …

Events:
<0.0 - 4.0> The man sits in the chair holding a phone.
<4.0 - 5.5> The man stands up from the gaming chair.
<5.5 - 7.0> The man stands still while holding the phone.
<7.0 - 8.0> The man walks toward the right side of the room.

Qwen

Actions in time order: The man remains seated, holding the phone to his ear. He then stands up from the chair while continuing his conversation.

Gemma

Actions in time order: The person is looking down at the laptop. The person appears to be typing or interacting with the laptop.

Unclear details: No other actions or objects are clearly visible or in motion.

Marlin got the whole sequence, phone, stand and walk, and it did that while completely ignoring the four headings I’d asked for. That turned out to be a habit. All 18 of Marlin’s captions, under all three prompts, came back in its own Scene and Events layout, complete with timestamps nobody had requested. Qwen politely used my headings and stopped at standing up. Gemma never mentioned the phone at all.

The other clips split the same way. Marlin caught my head coming up off the desk, and Qwen, which missed more overall, was the one that noticed me adjusting my glasses.

Unfortunately, that phone caption was typical of Gemma. Most of its captions said some version of a person sitting at a desk using a laptop and stopped there, skipping actions I’d confirmed. When it did add specifics, some of them were wrong: things I wasn’t doing, objects in the wrong place, people who weren’t there. Two of its short-prompt captions even described events after the eight seconds had ended. On top of that, while it wrote its very first two captions, the Mac’s swap grew by 2,744 MiB. Swap is disk space macOS borrows when RAM runs out, and it’s slow. That was enough for me, and I dropped Gemma.

The overall tally shows why the vague captions bothered me so much. Across all 54 captions, the agent made 225 checks: 60 correct, 72 missed, 9 contradicted and 84 unknown. Unknown came second only to missed, and it’s the sneaky kind of failure. A caption like “the person continues working at the desk” never contradicts anything, but it never gives search anything to match either.

With the structured prompt, Marlin got 10 of its 25 facts right and missed 3, Qwen got 8 and missed 5, and both left 12 unknown. I thought I had my shortlist, until a later review of the other six clips turned up one where I’m alone at my desk. Marlin described a second person in it with every prompt I gave it, and Qwen did too with the structured prompt, the very one I’d picked for it. Qwen’s other two prompts left the phantom out, but each of those had made up something else I’d rejected. Not a single model and prompt I tested came out clean.

A bigger model, then a new prompt

When small models struggle, the obvious move is a bigger one, so I tried Qwen3.5-4B with the structured prompt on the same clips. I’d told the test script to stop the run if swap grew by more than 512 MiB, and the 4B model blew past that 76 seconds in. By the time the worker was shut down, swap had climbed 1,582 MiB and I had zero captions to show for it. Swap counts everything running on the machine, so I can’t pin all of that on Qwen, but it wasn’t a great start.

The agent suggested closing everything else on the laptop and trying again, and I said no. This is the laptop I work on, and Triloch has to run in the background while I’m working.

What followed was two days of runs that went nowhere, six in a row without a single caption. The first two were the 4B model: a path bug in my own test script, then that swap limit. Next I wanted to see whether the small Qwen could run while I worked, but the script wanted 6.7 GiB of free memory before starting a model and came up 106 MiB short. A 4-bit version of the small Qwen, which would need less memory, didn’t get far either, because its download ran into my 10-minute limit. Then the same memory check refused to start a prompt test, and when I tried that test again, a bug in the script that records each run crashed it.

Four of those stops were limits I’d set myself and two were my own bugs, and none of them taught me a thing about captions. The funny part is that the small Qwen and Marlin had already been running on this laptop alongside my work without any trouble. So I stopped measuring memory and went after the captions directly, with a better prompt.

Remember Marlin’s stubborn Scene and Events layout? Its model card says that’s exactly the format it was trained on. So I wrote a new prompt in that shape for both models, and added a rule aimed at the phantom: mention another person only when a separate body is clearly visible, and tell reflections and people on screens apart from real ones. I also told them not to invent timestamps. Both models ran it on four clips I’d already checked, eight captions in all, and the agent scored them against the ten frozen facts from those clips:

Model Structured prompt New prompt Second person
Marlin 6 of 10 4 of 10 Still there
Qwen 5 of 10 4 of 10 Still there

I was hoping for at least a small win, and got the opposite. The new prompt did worse for both models, and both still put the second person in the room, rule or no rule. Marlin even wrote timestamps in all four captions, right after being told not to.

Where I stopped

Before calling it, I went back to what the captions are actually for, which is search. Two results from September had stuck with me, and the first one is a caption that was right when I was wrong.

In the first full run, Qwen and Marlin captioned all 3,156 clips from two days, and a keyword search ran over the captions. For “When did I drink coffee?”, my answer key had one answer at 13:00, the sip I’d acted out from the script. Qwen’s top result was at 13:16, so the scorer counted it as a miss. But when I opened that clip in the review viewer, there I was, raising a white cup to my mouth. It was coffee. Apparently one sip wasn’t enough, and my answer key only knew about the one the script asked for. I added the second sip and rescored the same saved results. Qwen’s first-result hits went from 3 to 4 out of 12 and its top-ten hits from 5 to 6, while Marlin’s stayed at 7.

The coffee answer key A timeline from 12:55 to 13:20 in three steps. One: my answer key has one sip at 13:00. Two: Qwen's top result is at 13:16, where the key has no sip, so it is scored as a miss. Three: I watched that clip and it showed a second sip, so I added 13:16 to the answer key. The same result is now scored as a hit, and Qwen's hits in the top ten go from 5 of 12 to 6 of 12. 1. My answer key had one sip, at 13:00. 2. Qwen's top result was 13:16. Nothing marked there, so a miss. 3. I watched that clip: a second sip. Same result, now a hit. MY ANSWER KEY QWEN'S TOP RESULT 12:55 13:00 13:05 13:10 13:15 13:20 13:00 13:16 scored: miss 13:16 scored: hit Qwen, hits in the top ten: 5 of 12 6 of 12 The coffee answer key A timeline from 12:55 to 13:20 in three steps. One: my answer key has one sip at 13:00. Two: Qwen's top result is at 13:16, where the key has no sip, so it is scored as a miss. Three: I watched that clip and it showed a second sip, so I added 13:16 to the answer key. The same result is now scored as a hit, and Qwen's hits in the top ten go from 5 of 12 to 6 of 12. 1. My answer key had one sip,at 13:00. 2. Qwen's top result was 13:16.Nothing marked there, so a miss. 3. I watched that clip: a secondsip. Same result, now a hit. MY ANSWER KEY QWEN'S TOP RESULT 12:55 13:00 13:05 13:10 13:15 13:20 13:00 13:16 scored: miss 13:16 scored: hit Qwen, hits in the top ten: 5 of 12 6 of 12
“When did I drink coffee?”, in local time. Qwen's result never changed. My answer key did.

Ever since, I watch a strange top result before I score it. The script told me what I meant to do in front of the camera. The footage had the rest of the afternoon.

The second result went the other way. Two of my real questions were about the night before: when were the lights turned on, and when were they turned off? Someone else in the house did both, at 22:06 and 22:18. I took the top three keyword search results for each across 12 hours of Marlin’s September captions. For off, none of the three came anywhere near the answer. For on, the third result did overlap the answer, but none of the top three captions said a light came on. The ones that mentioned lighting just described a room that was already lit. And the one caption that did say the lights went off ranked first for both questions.

That’s when it hit me that I’d spent two weeks testing models and still had nothing I could type a question into. So I called it. Marlin with the structured prompt is the default now, and Qwen stays around for the details it catches. Both still put a second person in that one clip, and I’m building on them anyway. A search box I can actually use will show me which of their mistakes really matter.

The tests did leave me something I’ll keep using. The 25 facts are frozen and the frames are saved, so any new model’s captions get checked the same way without me rewatching a single clip. If I were starting over, that’s the part I’d copy first: write down what happened before you read any captions. What the tests can’t tell me is how these models do on a day I haven’t already watched, so I’ve set October 2 aside for that and haven’t looked at it yet.

Next up is a search box in the review viewer: type a question, get three clips back. I can’t wait to point it at the lights.

Full size image