
Mode breakdown
The problem is not spotting a weird clip once
Matt Wolfe’s video starts with a problem a lot of people already recognize: AI-generated video is getting harder to separate from real footage, and the burden often lands on the person who has to decide whether to trust what they just saw. His idea is simple enough to explain in one sentence. Build an AI slop detector where someone can paste a link from YouTube, TikTok, Instagram, or X, then get a useful verdict before forwarding the clip to everyone else.
That direct answer is also the reason the experiment is worth watching. Wolfe is not trying to prove that AI video detection is impossible. He is trying to see whether a practical consumer-facing tool can exist yet, and what it would take to make one feel dependable. The answer, by the end, is not a clean yes or no. It is closer to: yes, but only if you accept a lot of tradeoffs that are easy to overlook when you see a polished demo.
Why the first detector breaks down so quickly
Wolfe begins by testing the idea inside ChatGPT, then moves into a coding workflow with CodeX and a project folder on his computer. That part sounds straightforward, but the first version exposes a core problem: the model can produce a nice-looking interface and still make bad judgment calls.
His early test case is the kind of clip that feels obviously synthetic to a human viewer. The detector, however, labels it as not AI, then explains its reasoning in a way that sounds confident but does not match the creator’s own judgment. Later, another round of prompting improves parts of the logic, but not the trustworthiness of the overall result. The app keeps landing on “inconclusive” even when Wolfe thinks the clip is clearly AI-generated or clearly real.
That is the most useful lesson in the whole build. A detector does not fail only when it misses obvious AI. It also fails when it cannot explain itself in a way a user can understand. If every questionable clip gets pushed into the same middle category, the tool stops being a shortcut and becomes another place where people second-guess the answer.
The transcript also shows another practical issue: confidence labels can be misleading. Wolfe finds that the app is assigning high confidence in ways that do not appear grounded in strong evidence. In a consumer tool, that would be a bigger problem than simply getting a few clips wrong. Users tend to trust the label before they inspect the logic.
The Site Engine workaround gets closer, but not cleanly
After hours of prompting and repeated stalls, Wolfe pivots to Site Engine, a third-party API that can analyze media for signs of AI generation or manipulation. This is the point where the project starts feeling less like a coding experiment and more like a real product decision.
The API route gets the detector closer to something usable. On several clips, the Site Engine side seems to do the heavy lifting: it flags likely AI content, while Gemini’s review often returns no clear indicators. When both systems disagree, the app initially falls back to inconclusive. That is safer than pretending certainty, but it is also frustrating if your goal is to give someone a fast yes-or-no answer.
Wolfe keeps refining the logic until the app begins to behave more usefully on the kinds of clips he tests. Real footage is recognized as non-AI. Synthetic-looking clips are flagged as AI. The report still is not especially elegant, and it does not always explain the evidence in a way a casual user would want. Still, it reaches the point where the tool can separate some obvious examples with better consistency.
That progress comes with a price. Wolfe checks usage and realizes how fast the operations add up. This is where the idea of a public, free-to-use detector starts to break apart. A tool that works only when the creator is paying for uncertain API consumption is not automatically ready for open access. The technical win and the product viability are two different problems.
What the experiment says about AI, AGI, and confidence
Wolfe uses the project to push back on the idea that current public models amount to AGI in any meaningful everyday sense. His argument is narrow, but it lands: if these systems are supposed to be broadly human-level, why are they still shaky at something a person can often do quickly by eye?
That does not mean AI detection should be trivial. Wolfe acknowledges the cat-and-mouse dynamic: as detectors improve, generation tools can improve too. But the video’s broader point is that public-facing models still struggle with a task that feels basic in practice. They can summarize, analyze, and describe video content, yet still hesitate when asked to identify whether the clip itself is synthetic.
There is also a clear editorial distinction here. Wolfe is not claiming that AI can never get better at detection. He is showing that the current public stack still has blind spots, especially when the task requires more than general scene understanding. A model can appear intelligent and still be weak at a narrow verification task.
For viewers trying to interpret the AGI conversation, that matters. It suggests that flashy capability demos and reliable real-world judgments are still separated by a gap. That gap is exactly where tools like an AI slop detector have to live.
Why the unfinished tool still has value
By the end of the video, Wolfe does not ship a polished public product. Instead, he ends with a working prototype that he can use himself, plus code he can share for others who want to run it locally with their own API keys. That outcome may sound smaller than the original goal, but it is actually the most realistic one.
A public detector for social clips has to solve three problems at once: accuracy, explainability, and cost. Wolfe gets some of the way there on accuracy, less so on explainability, and runs straight into cost control. Those constraints are not side notes. They define whether the tool can leave the prototype stage.
What makes the video useful is that it shows the friction instead of pretending it away. If you are hoping to build or use an AI video detection tool, this is the part to pay attention to. A model can be smart enough to impress you in one test and still be too inconsistent, too expensive, or too opaque for everyday trust.
That is why the project feels honest even though it never fully solves the problem. The tool exists, but the gap between “kind of working” and “reliably usable” is still large. For now, Wolfe’s detector is less a finished product than a case study in what AI video verification actually requires.
Recommended next
Products & tools
One Mode Digital Media product and one relevant affiliate recommendation selected for this page.

This Meeting Could Have Been an Email
A clean corporate-style design for the universal workplace thought: “This Meeting Could Have Been an Email.” Styled like a calendar invite, it’s perfect for remote workers, office teams, marketers, developers, and anyone who values fewer meetings and more actual work.
Cerri
Project and work-management software for teams that need clearer planning, ownership and delivery workflows.