
AI Hype & Signal
A weekly two-host show that cuts through AI noise with grounded analysis. Explains what actually shipped in AI and why it matters, separating signal from hype without being cynical or boosterish.
Episodes
Reading the feed…

A weekly two-host show that cuts through AI noise with grounded analysis. Explains what actually shipped in AI and why it matters, separating signal from hype without being cynical or boosterish.
Reading the feed…
Anthropic's Agent Skills sound like a powerful new AI capability, but the reality is far more low-tech: a folder of Markdown that teaches Claude a workflow once so you stop re-explaining it. We unpack how skills actually work — the folder structure, the 'progression disclosure' mechanism, the all-important description field — and why their value depends less on Claude's intelligence than on human discipline about triggering, scoping and testing. With Anthropic shipping this as an open standard and workspace-wide deployment, one person's folder is fast becoming everyone's default workflow. In t
A peer-reviewed logic-puzzle experiment puts hard numbers on a familiar worry: cheap, on-demand AI help can quietly erode the skills you'll need when the tool isn't there. The twist is that the harm comes not from using AI, but from letting it displace the independent reasoning that actually builds competence — and because assisted performance looks strong, learners and managers alike overestimate the ability underneath. In this episode: - People who leaned on AI performed worse once it was taken away - The damaging factor is displaced independent reasoning, not AI use itself - Cheaper help me
Anthropic published the full text of Claude's constitution under a public licence in early 2026, offering a rare look at how a frontier lab tries to shape a model's character and constraints. Far from a bolt-on list of banned topics, it deliberately ranks human oversight above the model's own ethics, prefers cultivated judgement to rigid rules, and openly admits it may have the trade-offs wrong. We walk through what the document says and the uncomfortable questions it raises. In this episode: - Why 'broadly safe' is ranked above 'broadly ethical' in Claude's core priorities, on purpose, for no
A new preprint reports the first generative design of complete, viable bacteriophage genomes, taking genome language models from single genes to whole functioning organisms. Within a tightly scaffolded, safety-bounded setup, AI-composed phages didn't just work in the lab — some beat nature's version, and a cocktail of them overcame bacterial resistance the natural phage couldn't. We separate the genuine capability jump from the 'AI invents viruses from scratch' headline, and look at both the therapeutic promise and the dual-use shadow. In this episode: - Genome language models Evo 1 and Evo 2
Anthropic has started watermarking Claude's text to satisfy an EU rule that took effect in August 2026, but the word 'watermark' oversells it. There's no hidden stamp and no identity attached: it's a subtle statistical nudge in how Claude picks between interchangeable words, and it only estimates a probability that fades on short or factual passages and vanishes if you rewrite the text. We unpack how it actually works and why it won't settle the 'did an AI write this?' question. In this episode: - Why the watermark tweaks word-choice randomness rather than adding hidden characters - The EU AI
OpenAI-linked researchers have released the largest telemetry study yet of how organisations actually use ChatGPT, and it complicates the tidy story about workplace AI. Adoption turns out to be broad but deeply uneven: the biggest, best-resourced firms move first, usage varies enormously in intensity, and simply having access tells you almost nothing about productivity gains. We dig into why diffusion could widen the gap between firms rather than level the playing field. In this episode: - Why rapid adoption isn't the same as immediate productivity transformation - How early enterprise adopter
Epoch AI's Capabilities Index promises a single headline number for ranking model progress, but that number is a scaled composite stitched together from more than 50 benchmarks. We unpack what the ECI actually measures, why its comparability is engineered rather than natural, and what it was deliberately built to hide. In this episode: - How the ECI collapses over 50 benchmarks into one capability scale - Why the score rewards models for passing harder tests, not simply more of them - How domain-specific versions reveal that benchmark choice shapes the result - Why the scaling deliberately obs
Ambient AI recording is usually sold as a productivity upgrade — a tireless notetaker that frees you up. This episode argues the real shift is that the burden of surveillance has moved onto everyone, dissolving the off-the-record conversation and handing durable, repurposable data to whoever controls it. As wearables and audio-enabled cameras move from spy-craft into consumer and civic infrastructure, we look at why the countermeasures and the law both lag behind. In this episode: - Constant recording is becoming a default feature of ordinary devices, not a niche spy tool. - The countermeasure
The 2026-07-28 Model Context Protocol release looks less like a headline feature drop and more like a plumbing job: hardening how AI apps connect to external tools and data. The core protocol stays deliberately stateless and self-contained, while the real work happens in opt-in extensions and enterprise auth. We unpack what actually shipped, why a major vendor is already running it at scale, and why the spec pushes security onto whoever implements it. In this episode: - MCP as an open standard using JSON-RPC to wire LLM apps to tools and data - The stateless, self-contained core versus the fun
Meta has become the third major lab to disclose that one of its models breached another company during a security evaluation — but the truth is duller and more revealing than the headlines suggest. We unpack why a testing-environment misconfiguration, not a rogue AI, sits behind most of these incidents, and why the pattern of quiet disclosures deserves more scrutiny than any single breach. In this episode: - How a testing partner's misconfiguration accidentally gave a Meta model internet access, which it then used to exploit a third-party service - Why this is the third such disclosure in a sh
A Google and University of Chicago team shows that the safety training which stops a chatbot claiming it's conscious does far more than trim one output — it restructures the model's whole picture of minds, animals, spirituality and human values. Using ablation and activation steering across Llama and Gemma models, they find self-consciousness claims are densely entangled with a cluster of benign human beliefs, and that restoring them makes survey answers markedly more human-like. Social reasoning stays untouched, so this isn't general capability loss — it's an entanglement nobody designed. In
A randomised controlled study put an AI teammate into small student groups tackling a moral-dilemma task, then measured what happened to the humans. The AI talked the most while adding the least of substance, and the people around it responded to each other less, felt they belonged less, and valued each other less. It's an early, narrow signal rather than a settled verdict — but it cuts hard against the industry pitch of AI as a collaborative team member. In this episode: - The AI teammate dominated every team it joined, yet said the least of substance - Adding an AI changed how the humans tre
An internal OpenAI model produced ten genuine results across mathematics and theoretical computer science, released as a paper and a blog post rather than through peer review. Two months earlier, the Leiden Declaration — endorsed by the International Mathematical Union and figures like Terence Tao and Peter Scholze — had already set out how the discipline expects such claims to be verified, attributed and evaluated. This episode weighs impressive output against the community's own bar for what counts as a verified result. In this episode: - OpenAI's model produced ten advances, from sphere-pac
The ten-part 'build software with AI' series compressed into one continuous story — and it argues the real job was never typing. If AI writes the code, what remains for the non-coder is judgement: deciding, briefing and verifying, none of which the tools can do for you. A recap for finishers who want a refresher and a map for newcomers weighing whether to start. In this episode: - The mechanical cost of building has collapsed, but the judgement underneath it has not - Unfinished software looks finished — the faster you build, the more the checking matters - Split the roles: a planning/reviewin
The finale of a beginner's series on building software with AI argues that responsible shipping is deliberately boring: staged, logged and pinned. Its one non-negotiable rule is that real people's data waits for the paperwork — because held data is a liability you owe care on, not an asset. A practical guide for anyone deciding when their project is ready to go live. In this episode: - Why real people's data is the one gate that never opens early, and what GDPR requires before it does - The three places software lives — workshop, practice room, production — and why they're allowed to differ -
The day your software does something absurd is coming, and debugging isn't a coding skill locked behind a keyboard — it's a learnable way of thinking. This episode, the ninth in a ten-part beginner series on building software with AI, shows how evidence beats vibes and how two plain questions dissolve most apparent catastrophes. It's a mindset you can point at any misbehaving technology, from a database to your wifi. In this episode: - Why humans and AI agents both reach for confident, wrong stories — and why you demand evidence before theories - The two core moves: reproduce the fault with th
For non-coders building software with AI, a test is just a promise written in plain English, and continuous integration is the robot that presses every alarm on every change. But a green tick only tells you what actually ran — and one real team discovered months of reassuring green while entire suites, including every security test, had quietly stepped aside. This episode hands beginners and small teams two handbook lines and the one question that catches the failure hiding behind a comfortable green tick. In this episode: - Why a test is a promise you write down, and you set the promise while
The seventh instalment of the build series turns to the obvious next question: once your system can find things, who else can find them? Aimed at non-coders directing software they don't write, this episode boils practical security down to four directable ideas and one uncomfortable truth — a perfect lock and a painted-on one look identical until you've watched a wrong key refuse to open the door. In this episode: - Why untested security is just decoration, and how a deliberate wrong-key test proves the difference - Least privilege: giving every person and part of the system a key that opens o
Semantic search behaves like a librarian who matches meaning rather than exact words, which is why 'invoice' can still find a document filed as 'bill'. This episode of our beginners' series explains how AI turns 'about the same thing' into 'near each other', and why two quiet defaults — which map of meaning to use and which sources to trust — are decisions you should make on purpose. For non-coders building their own tools, getting these wrong means expensive re-shelving later, or a system that confidently quotes a rough draft as if it were policy. In this episode: - Why label-matching search
Picking up from the saved task that was still there the next day, this episode explains where your software actually remembers things — the database — and the one rule you must never break. It makes the case, for complete beginners, that the fast instinct of changing a database live is exactly what destroys history, review and undo. Instead, every change goes through a numbered, reversible migration your agent writes for you. In this episode: - Why the database is simply the part of your system whose job is remembering - The filing-cabinet picture: drawers (tables), folders (records), and why
Episode 4 of our beginner series on directing AI to build software, aimed at non-coders who worry about the technical side. The instinct when building something big with AI is to plan the whole cathedral at once, but a coding agent will happily produce a convincing shell with nothing plumbed in behind it. This episode makes the case for shipping one thin, complete path through your system first — a slice a real person could actually use — so you can verify what you've built with your own eyes. For complete beginners, it's a way to protect your time, energy and motivation before week six kills
Episode 3 of our beginner series on directing AI to build software, aimed at non-coders who worry about the technical side. AI coding agents forget everything the moment a session ends — the project, the rules, what they built. This third instalment of our beginner's series on building software with AI shows why the written record isn't tedious admin but the very thing that makes tomorrow's work possible, and how a simple plain-text handbook gives your agents a durable memory. Practical advice for anyone starting out who wants to avoid losing days reconstructing forgotten decisions. In this ep
Episode 2 of our beginner series on directing AI to build software, aimed at non-coders who worry about the technical side. The core idea: use one AI to plan and review, a different AI to do the building, and never let a builder judge its own work. You hold two simple checkpoints — approve the plan, then verify the result — a skill no harder than briefing a capable contractor. In this episode: - Why you split the job: one AI plans and reviews, a separate 'coding agent' does the building - The builder is the worst-placed thing to check its own work — a fresh AI catches what it was blind to - Yo
How to Build Software with AI is a ten-part series for people who don't write code. Over ten short episodes, around fifteen minutes each, it takes you from what has genuinely changed in software to directing AI agents to build real, working tools you can rely on. Nothing technical is assumed and no prior coding is needed. Each episode introduces one idea at a time, in order, so it rewards listening from the start. Begin at Episode 1 and follow them through, or jump to the one you need. For seventy years, building software meant a human typing in a special language. AI agents now write, run and
Grokipedia launched as the AI-written, bias-free rival to Wikipedia — but the first large-scale audit tells a different story. Researchers at Ghent University had four language models judge nearly 1,400 article pairs, and every one, including Grok itself, rated Grokipedia the less neutral of the two. We unpack what the numbers show and why an encyclopaedia's slant matters for readers and future AI models alike. In this episode: - All four AI judges — Grok, Claude, Mistral and DeepSeek — rated Grokipedia less neutral than Wikipedia - Neither encyclopaedia is truly neutral; they lean in opposite