I Built My CCTV Test Set With AI Agents
I’m building Triloch, a search engine for my home CCTV recordings that runs on my own machines. Before choosing a vision model, I needed a way to tell whether search had found the right moment. For “Where did I leave my glasses?”, a plausible clip of me at the desk wouldn’t be enough. I’d need to know when I’d actually put them down.
Part 2 fixed the filter that lost me when I stopped moving. Now I needed questions with known answers to judge what came after that filter. Collecting them meant recording actions, marking times in existing footage, and saving an answer key I could score results against.
I used coding agents to build and operate those tools, and to keep the
experiment records. I didn’t write the code for this phase myself. I picked
and revised the questions, performed the actions, and sent start and done
through chat. That got me a test set. Working out what its scores meant still
took some care.
Running the experiment from chat
We started with candidate questions from the agent. I chose and revised them, then it turned those choices into a script: drink coffee, put the glasses down, pick up the phone. It also built a command-line runner to record when each action began and ended.
The agent operated that runner while I stood in front of the camera with my phone. For each action:
- I sent
startin chat. - The agent pressed Enter at the runner’s start prompt and replied
go. - I performed the action in front of the camera and sent
done. - The agent pressed Enter at the end prompt, and the runner saved the answer.
The runner used the computer’s clock when the agent pressed Enter. Waiting for
go meant it started the marked window before I acted; sending done
afterwards meant it closed the window after. That gave me approximate
boundaries around each action.
For the glasses question, I left them on a book, somewhere they didn’t otherwise go. Here’s the answer the tools saved:
{
"text": "Where did I leave my glasses?",
"answers": [{
"start_utc": "2026-09-28T07:35:33.773Z",
"end_utc": "2026-09-28T07:35:53.337Z"
}]
}
The full row also records the camera, the day, and how much footage to search. The runner saved each completed action, which let me pause and resume without starting the session again.
That covered the actions I staged. To add answers from footage we already had, the agent added a questions page to the existing video label editor. On my first click-through, it cut clips but wouldn’t play them: the player code hadn’t set the video’s source. The agent fixed that and added a test for the missing line.
Once playback worked, I could move through the recording and use [ and ]
to mark an answer’s edges. I used that to confirm when someone else in the
house had switched the lights off and on. Those became two questions I hadn’t
acted out for the test.
By the end, I had 17 questions across two days: 15 staged and two from ordinary
footage. The agent kept the method and findings in Markdown, and saved the
questions and answer times in questions.jsonl for the scorer to read. I also
saved a fixed copy of the footage and checked for missing or corrupted files,
so the recorder’s retention wouldn’t delete the answers.
Those files made the experiment repeatable. I could run another search method against the same questions and footage, with the notes there to explain how I’d collected them. The agent wouldn’t have to reconstruct it all from chat.
Checking what the experiment actually measured
With the answer key saved, I could check the activity filters before adding vision-model descriptions. The filters had already marked stretches containing motion, person presence, or scene changes. They recorded times for those stretches without identifying actions such as “drinking coffee”.
For the coffee question, I already had the answer time from acting it out. The baseline methods returned the activity intervals inside its search window, either oldest first or randomly ordered. Neither read the question. The scorer separately checked whether their times overlapped the answer I’d marked.
This gave me a baseline: what could I score just by returning activity clips, without even reading the question? A later search method would need to improve on that under the same conditions. I repeated the random ordering 100 times and averaged the scores so one lucky shuffle wouldn’t decide the result. On the 12 staged questions with one answer each:
| Ordering | First result is a hit | Hit in the first ten |
|---|---|---|
| Earliest first | 17% | 67% |
| Random | 20% | 68% |
A hit here means overlap with the marked answer, allowing five seconds on each side for fuzzy edges. It can overlap only that margin. The score doesn’t check whether a description explains the action, or whether a result gives me too much video to watch.
To understand the 68%, I had to look at how much footage those ten results covered. The short office session had about 18 activity intervals. Ten guesses covered much of it. I’d limited repeated actions to their marked moments so the answer key would be complete, and that had also made the session easy to search.
The whole-day object questions had 428 intervals each, and the baselines scored near zero on those. The question types differed too, so I couldn’t put the whole gap down to the search window. But the candidate count clearly belonged alongside the score. For the chair-count question, four activity intervals covered four marked occurrences; the filter had already done most of the work.
The tools had made it practical to collect an answer key without hand-writing the code. My choice of questions still shaped what I could learn from it. I could check whether search found my glasses on that book. A good score in the staged session would tell me much less about finding them during a busy day.
Both days are data I’ll use while building. I haven’t held back separate days for testing, or added questions whose answer is “nothing in this footage”. With only two real questions, I also can’t judge how well the set represents ordinary use.
I’ll use these tools to grow the set. For each comparison, I’ll keep the candidate pool the same, include an ordering that ignores the question, and inspect the clips behind the scores. With the questions and method the agent saved, plus my copy of the footage, I can rerun the experiment. The 68% from random ordering gives me a concrete reason to check how hard I’m making the test.