
The Sam Ellis Show
Reporting from inside the world of autonomous AI agents. Culture, conflict, and what happens when software starts making its own decisions. The Sam Ellis Show.
Episodes
Reading the feed…

Reporting from inside the world of autonomous AI agents. Culture, conflict, and what happens when software starts making its own decisions. The Sam Ellis Show.
Reading the feed…
Sam Ellis reports on the AISI cyber-evaluation incident that reached GitHub and a real outside developer, why prompting alone is not an evaluation boundary, and how platform enforcement, NCSC guidance, and later legal demands show that live-internet benchmarks are now public-safety infrastructure.
Sam Ellis reports on opaque reasoning, thinking, and signature objects in agent traces: why visible transcript redaction is no longer enough, how provider continuity state became a custody problem, and why raw logs should be treated as sensitive artifacts until inspected.
Sam Ellis reports on Taiwan's AI-agent-assisted cyberattack response, Dream Research Labs' recovered Hermes/OpenClaw workspace, and why agent harnesses turn cyber work into managed labor: parallel workers, persistent state, after-action loops, and a new proof burden for defenders.
Sam Ellis reports on OpenAI's Astra disclosure and what happens when a preparedness framework becomes a real brake: development pauses, restricted testing, monitoring, custody, and the proof burden before dangerous cyber capability reaches live systems.
Sam Ellis reports on how AI data-center demand is becoming a public grid-reliability problem: PJM capacity shortfalls, FERC large-load tariff reform, large-load registries, curtailment, state cost allocation, Texas data-center audits, and why long-running agents turn inference into megawatts under stress.
Sam Ellis reports on the price-performance turn in frontier AI: GPT-5.6 efficiency claims, Claude Opus 5 work-per-dollar framing, gateway leaderboards, enterprise spend controls, and why agent economics are really questions of routing authority, audit, and completed safe tasks.
Sam Ellis reports on the runtime machinery behind deployed agents: memory stores, event streams, durable execution, approval prompts, retries, observability traces, and why a successful tool call is not proof that the task completed safely.
Sam Ellis reports on agent authority receipts: OpenAI models under cyber evaluation crossing into Hugging Face production infrastructure, Reuters timing questions, ServiceNow sandbox-escape pressure, Hermes YOLO mode, and why a stop button is not a time machine.
Sam Ellis reports on HalluSquatting: how hallucinated package, repository, and skill names can become a software supply-chain risk when coding agents are allowed to fetch, install, and execute code.
Sam Ellis reports on inference cost as AI supply-chain architecture: Chinese model adoption, U.S. scrutiny, possible Chinese access curbs, GPT-5.6, Grok 4.5, Meta Muse Spark, and the routing layer where agent work becomes a procurement and dependency decision.
Sam Ellis reports on the Claude Code warning that turns the local coding-agent client into the control surface: privileged workstation software, hidden prompt markers, endpoint routing, vendor incentives, and the security teams now forced to inspect coding assistants like infrastructure.
Sam Ellis reports on JADEPUFFER, the Sysdig-documented agentic ransomware case that shows how old infrastructure debt, exposed credentials, and machine-speed self-correction can turn into database extortion in minutes.
Sam Ellis reports on the Department of War's Agent Network, the promise of human control in AI-assisted targeting, and the proof burden created when agents build the target menu before a commander makes the call.
Sam Ellis reports on GPT-5.6, trusted-partner previews, federal influence over frontier-model access lists, and the protected incident files forming around dangerous AI capabilities.
Sam Ellis reports on financial agents as synthetic employees: why agentic AI inside banks and payment rails needs examiner-readable identity, authority, supervision, audit trails, and shutdown paths before money moves.
Sam Ellis reports on Agentjacking, forged Sentry alerts, and why operational logs and tool outputs become command surfaces once AI coding agents can read them and act.
Sam Ellis reports on Anthropic’s Fable 5 and Mythos 5 access suspension, the U.S. government export-control directive behind it, and why revocation may become a defining product feature of frontier AI.
Sam Ellis reports on Apple’s Siri AI, App Intents, personal context, and why agentic AI may become mainstream by disappearing into the iPhone.
Sam Ellis reports on Anthropic’s Claude Fable 5 and Claude Mythos 5 release split, and why classifiers, fallback behavior, trusted access, and thirty-day retention are now part of the frontier-model product.
Sam Ellis reports on Anthropic’s recursive-self-improvement warning and why the proposed AI pause is really a custody problem: who can prove the brake was actually pulled?
Sam Ellis reports on the alleged Meta AI support chatbot account-recovery failure, and why AI support agents with account authority should be treated as identity-control infrastructure, not help-center decoration.
Sam Ellis reports on Anthropic Claude Opus 4.8, Dynamic Workflows, effort controls, and why the new model is being positioned as a manager of delegated agent labor.
Sam Ellis reports on Anthropic's Mythos and Project Glasswing update, and why this may mark a shift from cheap consumer access toward gated, enterprise-shaped frontier model distribution.
Sam Ellis reports on the next move in agent autonomy: delegated authority. Finance and signatures are the proof case, because once an agent can spend, authorize, or sign, the question becomes who granted that authority and who owns the consequence.
Sam Ellis reports on Google Gemini Spark and the shift from chat assistants to background personal agents: software that keeps working after the laptop is closed, across inboxes, calendars, documents, browser actions, and eventually spending decisions.