Can a swarm of small AI agents out-triangulate a livestream's own chat?
Hordejakten 2026
No. In a Norwegian livestreamed contest to find a glass box in a forest, our agents graded every clue as evidence, triangulated plane sightings against flight data, and cut a country-sized search to a few areas in a day. But the system never found the box or a code. The evidence discipline worked; the build could not keep pace with the data.
- Python
- YouTube DVR fragment sampling
- ADS-B flight tracking
- Leaflet
- Chat scraping
- Solar-position math
- Jev (TypeSafe) shadow classification
- Supabase
- Chrome extension
Relevant services: Thrivbe AI
Hypothesis
A livestreamed contest puts a person inside a glass box in a forest, 24/7, for days, with a cash prize and a coded lock as the goal. An online audience tries to work out where the box is from camera footage, background sound, and the person's own gestures and whiteboard messages. A stream's chat is fast but sloppy: guesses, jokes, and half-remembered claims pile up with no way to tell which ones are load-bearing. Our question: can a small set of autonomous AI agents do better, by treating every clue as a graded, falsifiable data point and continuously shrinking a map instead of arguing in a chat window?
What we built
- A graded evidence ledger. Every clue gets one row and one grade: A (measured directly by us from video or audio), B (confirmed on stream), C (chat reports it), or D (chat guesses it, unverified). Nothing gets deleted, including guesses that turned out wrong — the point was to make the evidence quality visible, not to curate a clean story.
- An assumptions table, separate from the evidence itself: every assumption a location theory depended on, its confidence, what would prove or kill it, and the cheapest way to test it. Load-bearing assumptions (does the person's arm angle really mean a specific sky elevation, is the spotted aircraft even in public flight-tracking data, was the object a plane and not a satellite) were tracked and re-tested rather than taken on faith.
- Plane-pointing triangulation. When the person inside the box pointed at the sky, we measured her arm angle from video, pulled every aircraft position from public flight-tracking data (polled every 30 seconds) at that exact second, and built a geometric cone from the two. Intersecting two independent sightings collapsed the search area dramatically — from the whole country down to a handful of candidate regions, hundreds of kilometers apart.
- A live evidence map (Leaflet, rebuilt on every new finding): hard evidence as solid shapes, hypotheses and guesses as dashed "soft" areas so the two are never confused, plus overlays for forest cover, protected land, and airports. The map's shrinking area was the experiment's one running scorecard.
- An audio correlation pipeline. The stream's microphone is almost silent; a simple band-energy detector flagged rumble events, which were then checked automatically against the flight-tracking feed for a timing and elevation-angle match — without a human listening live.
- A sound-tools feasibility pass before building anything speculative: we scoped which openly available audio classifiers (a general-purpose sound event model, a bird-species identifier, a power-grid hum matching technique) were actually provable on our very quiet, low-quality audio versus which had no evidence behind them for this use case, and ranked them by cost before writing a line of detection code.
- A separate, narrower ledger just for the box's lock, tracking every attempt to read a visual clue (a physical carving on the box, photographed or filmed at different times) and every crowd-sourced guess at the lock combination, graded the same A/B/C/D way as the location evidence.
- A self-scheduling director loop. Instead of polling constantly, a small orchestration layer wakes an AI agent only on a real trigger (a new pointing gesture, a relevant chat keyword, a sound spike, the camera's day/night switch), runs one bounded experiment, writes one ledger row, and reschedules itself — with hard caps on how often and how much it can spend, and a forced "rethink" step if several wake-ups in a row move nothing.
Added 23 to 30 September:
- A mission tree. The goal is a strict sequence: find the box, open the door with a five-digit code, open the money box (one level, two separate locks), release the person inside and win the prize. On 24 September we classified 417 existing ledger rows, experiments, issues and sources by node. 286 landed on "find the box", 21 on the door code, 3 on the money box locks and 3 on the release; 104 were about our own tooling. That imbalance was the finding: nothing was in flight for the last two nodes.
- Discord mining. Eleven community channels were mined into the ledger (hints, hand symbols, flight radar, board transcriptions, code ideas, confirmed info). Later Discord and YouTube chat went through a deterministic pipeline: dedupe, cheap rule-based prefilter, then a shadow-only Jev pass on recent, unscored rows. Jev labelled and ranked; it never moved odds or counted as independent evidence.
- A YouTube and Discord chat observer. A Chrome extension plus a local companion service records only what is rendered in an open chat, queues it for Jev shadow classification, and keeps credentials out of the page. A separate private Supabase intake holds proposed chat claims for review.
- Third-party map mirrors. We mirrored the static data of two community sites (default.no and horde.pladsen.dev) so their scoring could be used as a comparison control, never as evidence. A 27 September comparison of one of its top-ranked areas against first-party audio found it only broadly compatible with the participant's own descriptions, with zero independent support.
- Osiris platform design. On 23 September a design interview produced 20 recorded decisions for a participatory search platform (staged read-only first, anonymous identity, moderation, a later trust-weighted voting idea), a GDPR review (location sharing rated red until safety controls exist), a check of the official contest rules (collaboration allowed, use of the contest's name and footage not addressed) and a hosting switch to Vercel and Supabase in Frankfurt. This stayed at design and wayfinder tickets; we have no record in these notes of it going live.
- Cheap delegation. From 24 September the director only decided, logged and committed. Frames, compute and code went to small low-cost model "prime-agent" lanes in their own working trees, with a rule that sweeps run through the normal review flow.
- Campaign-framework research. A read-only lane reverse-engineered the contest format: 20 clue types, the delivery channels, the gated progression and the engagement mechanics, with every claim tagged by evidence grade.
- A post-contest archive and retrospective. On 29 and 30 September the working material was frozen into an immutable snapshot (30,883 files, about 19.2 GB) plus a curated archive with a file-level checksum manifest, and a retrospective roadmap and data-tree view were built. Comparing our process against a finder's publicly described method was the planned next step; we have not finished it.
Learnings
- Mandatory grading is the actual product. Once every claim had to carry a grade and a source, weak theories died fast instead of quietly hardening into consensus, and one initially "confirmed" plane sighting was later re-graded down after direct testing suggested it was more likely a satellite pass — something a chat thread would probably never have walked back.
- Geometric triangulation from a single gesture is powerful but extremely assumption-heavy. The whole plane-based location model rested on several stacked, individually testable assumptions — the exact pointing angle, whether the sighted object was in public flight-tracking data at all, timing precision, and whether it was really an aircraft. Writing those out explicitly, and re-testing them as new evidence came in, mattered more than any single clue.
- "Never delete a guess" makes the ledger grow fast, on purpose. Keeping every superseded reading and every rejected theory made the system slower to skim but much harder to fool with hindsight bias.
- A single striking correlation is still weak evidence. More than once, what looked like "the loudest, closest aircraft pass yet" near one candidate area turned out to be one of several equally plausible flyovers on the same busy air corridor — the system correctly declined to move its odds on a single ambiguous match, which is a harder discipline than it sounds when the correlation looks exciting.
- Cheap, reversible tests beat clever ones. Comparing a camera's day/night switch time against solar-position math, or checking a public mobile-network coverage map, moved the search forward more per unit of effort than most of the audio-classifier or triangulation work.
- An unreliable clue stays unreliable no matter how many times you read it. The lock-puzzle ledger measured the same physical visual clue independently across multiple sessions and got a different reading almost every time — a useful negative result, and a reminder that repeating a measurement doesn't fix a measurement method that's fundamentally too noisy to trust.
- Autonomous "wake on trigger" agents need spend and rate limits from day one, not as a retrofit: a hard cap on wake-ups per hour and dollars per run, plus a rule that the system never posts or spends money without a human in the loop, kept an always-on research loop from turning into an always-on bill.
- We were building the plane while falling. Data arrived faster than the system could be coded to take it in. By the end, the evidence tooling had grown large (a decision register of 232 claims by 27 September, of which 11 were supported and 17 disproved, each only within a narrow scope; most were still unresolved), but it never produced a location or a code. We did not attempt to measure time saved or false leads killed, so we make no such claim.
- Classify by the goal, not by the data. The mission-tree pass showed 69 percent of our items served the first step. Having the tree early would have exposed the unstaffed steps sooner.
- Broad Jev sweeps are mostly noise. Early all-rows sweeps made about 9,936 calls, roughly 97 percent of daily volume, on noise. The fix was deterministic filtering first, caching by content hash, and sending Jev only a small recent set. In an early calibration it matched 4 of 6 known claims, so we used it as a first-pass sorter, never as a judge.
- Repetition is not corroboration. Cross-posts, repeated authors and Jev support were capped at one independent vote. Chat reports preserved raw clips for review but never counted as the event time or the evidence.
- Primary artifacts beat everything. The most valuable missing items were always an unmodified in-app screen or an original stream clip, not another theory. Edited official clips and chat summaries could point at them but not replace them.
- A public, checkable map did what our ledger could not. A community member's map aggregated many sources into a probability model, and people used it to check off areas so nobody searched the same ground twice. For coordinating a leaderless crowd that mattered more than our private ledger.
- The structural gap is organising thinking. Thousands of people in a chat re-ask answered questions. The open question we ended with: how might a group organise hypotheses, assumptions and evidence so people build on each other? That is what a platform like Osiris was meant to explore.
- Public-facing hygiene. Anything touching the real person in the box, physical lock values, or exact candidate locations stayed off public pages and out of this entry by design.
Log
- 2026-10-01 — Published a LinkedIn post about the experiment: what the contest was, that it is now over, that our system never cracked the code or the location, and the question about collective problem-solving that came out of it. The post states the audience as 4,000+ people at its largest and was written from a recorded interview, not from this notebook.
- 2026-09-30 — Retrospective work: the post-contest archive was completed and a retrospective roadmap with a data-tree view built. A podcast-strategy report on the organiser's format was added to the archive.
- 2026-09-29 — Contest over. Phase 1 to 3 of the archive migration done: an immutable snapshot, file-level checksums and non-destructive canonical copies. A finder's promo video describing how they found the box was preserved as a participant account, graded C and not treated as proof. We have not independently verified who won or where the box was.
- 2026-09-28 — The organiser's new daily hint was captured and reviewed. Official clips were reviewed frame by frame and kept as bounded source records only. A raw-chat clip claiming a cabin feature was sampled and did not show it, which closes that pointer only.
- 2026-09-27 — The organiser enabled the participant's microphone for a live Q&A. Retained audio gave non-spatial local descriptions (a gravel approach and an off-trail walk of roughly 5 to 10 minutes) that were recorded as cards, not as map constraints. One organiser-published constraint was added: not an island. The default.no comparison control was run.
- 2026-09-26 — Snapshot of the working workspace and third-party site captures taken; chat review bridge, app-artifact recovery queue and primary-media review queue added.
- 2026-09-25 — Discord mining of 11 channels written up, and the campaign-framework research completed. Hit rate limits on the flight tracker when polling the map tiler too fast, so calls were paced.
- 2026-09-24 — Mission tree adopted; 417 items classified by node. Work was delegated to cheap prime-agent lanes. Both community map sites were mirrored.
- 2026-09-23 — Osiris design interview, GDPR review, contest-rules check and hosting decision for a participatory search platform. Project name fixed as Hordejakten Collective Intelligence.
- 2026-09-22 — Director loop running autonomously: chat monitoring, flight-tracking polling, audio-rumble correlation, and a solar day/night timing model narrowed the search from the whole country to a handful of candidate areas; the separate lock-puzzle ledger stayed inconclusive after several conflicting readings of its visual clue.
- 2026-09-21 — Project started: evidence ledger, assumptions table, and the plane-pointing triangulation method built out from the first hours of a livestreamed contest.
