
AI Paper Trails
Dr. Robert Li and Saadia Carbis trace cutting-edge AI research back to the real world.
Episodes
Reading the feed…

Dr. Robert Li and Saadia Carbis trace cutting-edge AI research back to the real world.
Reading the feed…
What are agent skills good for? What are the best (and worst) ways to use them? We explore Demystifying Agent Skills: Why They Work-Until They Don’t (Jiang, et al., 2026) to find out. https://doi.org/10.48550/arXiv.2608.14036
Is it practical to run the most efficient 3T class models? What are the implications of having 3T class model that’s open-source? We explore Kimi K3: Open Frontier Intelligence (Kimi Team, 2026) to find out. https://doi.org/10.48550/arXiv.2607.24653
Can AI write fiction like a human? We explore StoryScope: Investigating idiosyncrasies in AI fiction (Russell, et al., 2026) to find out. https://doi.org/10.48550/arXiv.2604.03136
Do Anthropic’s models contain a global workspace? If so, what are the implications? We explore Verbalizable Representations Form a Global Workspace in Language Models (Gurnee, et al., 2026) to find out. Verbalizable Representations Form a Global Workspace in Language Models
What are the pathways from here to AGI and ASI? Can anyone agree on a definition for AGI? We explore From AGI to ASI (Genewein, et al., 2026) to find out. https://doi.org/10.48550/arXiv.2606.12683
How can a self-evolving harness benefit a model’s capabilities? We explore Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents (Lin, et al., 2026) to find out. https://doi.org/10.48550/arXiv.2605.30621
Is Mythos all its hyped up to be? We explore System Card: Claude Mythos Preview (Anthropic, 2026) to find out. System Card: Claude Mythos Preview
We’re joined by Gareth O’Shea, to ask: Could it be possible to run an MRI on an LLM while it’s thinking? We explore Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models (Qwen Team, 2026) to find out. https://doi.org/10.48550/arXiv.2605.11887
Should we align models based on human preference or human behaviour? We explore Alignment Makes Language Models Normative, Not Descriptive (Shapira, et al., 2026) to find out. https://doi.org/10.48550/arXiv.2603.17218
How does extra domain specific pre-training contribute to code specific models? We explore Qwen3-Coder-Next Technical Report (Qwen Team, 2026) to find out. https://doi.org/10.48550/arXiv.2603.00729
Where do inconsistencies show up within the current state of world models? We explore The Trinity of Consistency as a Defining Principle for General World Models (Wei, et al., 2026) to find out. https://doi.org/10.48550/arXiv.2602.23152
Is RSI through multi-agent systems doomed to degenerate? We explore The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies (Wang, et al., 2026) to find out. https://doi.org/10.48550/arXiv.2602.09877
What’s the best way to teach AI competency in K–12 settings? We explore Unleashing human potential: An artificial intelligence competency framework for K–12 education (Kong & Hu, 2026) to find out. https://doi.org/10.1016/j.caeai.2026.100556
Can GenAI continue to grow beyond the finite limits of human data by generating its own? We explore Welcome to the Era of Experience (Silver & Sutton, 2025) to find out. https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf
Can a party of LLMs play D&D? We explore Setting the DC: Tool-Grounded D&D Simulations to Test LLM Agents (Zeng et al., 2025) to find out. https://openreview.net/forum?id=3Op7kJOvaD
Can a simple Engram lookup transform a model’s search for meaning? We explore Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models (Cheng et al., 2026) to find out. https://doi.org/10.48550/arXiv.2601.07372
Can small, local AI models transform the way we play? We explore Web World Models (Feng et al., 2025) to find out. https://doi.org/10.48550/arXiv.2512.23676
Have DeepSeek discovered a new technique to unlock training at massive scale? We explore mHC: Manifold-Constrained Hyper-Connections (Xie et al., 2025) to find out. https://doi.org/10.48550/arXiv.2512.24880
Should AI agents be optimised for reasoning, or is reliable tool use a more viable path? We explore Adaptation of Agentic AI (Jiang et al., 2025) to find out. https://doi.org/10.48550/arXiv.2512.16301
Welcome to AI Paper Trails! In this first episode we introduce ourselves (hi!), and discuss our very first paper: Agent READMEs: An Empirical Study of Context Files for Agentic Coding (Chatlatanagulchai et al., 2025). https://doi.org/10.48550/arXiv.2511.12884