
Unzip
Your guide to the latest AI and machine learning research. We unpack complex papers into actionable insights for practitioners and enthusiasts alike.
Episodes
Reading the feed…

Your guide to the latest AI and machine learning research. We unpack complex papers into actionable insights for practitioners and enthusiasts alike.
Reading the feed…
## Episode Summary In this episode, we cover: - **Procedura: Agentic 3D Modeling with Procedural Control** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.26238) - **Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.15763) - **Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit** (arXiv) - [Read more](http://arxiv.org/abs/2608.27427v1) - **Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal
## Episode Summary In this episode, we cover: - **CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.25500) - **Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.26809) - **TTPO: Test-Time Policy Optimization** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.27448) - **UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City** (arXiv) - [Read more](http://arxiv
## Episode Summary In this episode, we cover: - **GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.21832) - **The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.24358) - **SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.23564) - **Is Next-Chunk Reasoning RL Really Better than SFT?
## Episode Summary In this episode, we cover: - **When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.24569) - **WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.24053) - **Automata from Agent Traces: Failure and Next-Step Prediction** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.23670) - **SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation** (Hugging Face Daily) - [Read mo
## Episode Summary In this episode, we cover: - **SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?** (arXiv) - [Read more](http://arxiv.org/abs/2608.23564v1) - **ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.22510) - **What AstroPT knows about galaxies, and what that can teach us about LLMs** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.22614) - **Towards a Densing Law for User Representati
## Episode Summary In this episode, we cover: - **Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.18484) - **ParaTempo: Efficient Parallel Reasoning via Temporal Confidence** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.16425) - **Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.16647) - **Hadit
## Episode Summary In this episode, we cover: - **4DAnyone: Create Anyone in 4D from a Casual Monocular Video** (arXiv) - [Read more](http://arxiv.org/abs/2608.20335v1) - **FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.21596) - **Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.08466) - **The Embedder's Dilemma: LLMs Are Better, but at What Cost?** (Hugging Face Daily) -
## Episode Summary In this episode, we cover: - **NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.13210) - **Towards Quantifying Benchmark Optimization in ASR Models** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.19936) - **FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.19758) - **Chain-of-Experience for Continual LLM Improvement** (Hugging
## Episode Summary In this episode, we cover: - **Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.17744) - **ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models** (arXiv) - [Read more](http://arxiv.org/abs/2608.20338v1) - **MidTool: Mid-training Data Synthesis for Agentic Tool Use** (arXiv) - [Read more](http://arxiv.org/abs/2608.20314v1) - **AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement** (arXiv) - [Read
## Episode Summary In this episode, we cover: - **Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.16590) - **SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.17426) - **SPADE: Self-Play in Adaptive Synthetic Executable Environments** (arXiv) - [Read more](http://arxiv.org/abs/2608.19197v1) - **Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion** (Hu
## Episode Summary In this episode, we cover: - **HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.17597) - **PixRestore: Unified Image Restoration via Pixel Diffusion Transformer** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.16793) - **Demystifying Agent Skills: Why They Work-Until They Don't** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.14036) - **EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing** (Hugging Face Daily) -
## Episode Summary In this episode, we cover: - **Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory** (arXiv) - [Read more](http://arxiv.org/abs/2608.16889v1) - **How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.14905) - **StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.1508
## Episode Summary In this episode, we cover: - **DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.13517) - **Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.03744) - **A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.14075) - *
## Episode Summary In this episode, we cover: - **The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity** (arXiv) - [Read more](http://arxiv.org/abs/2608.13520v1) - **An AI4AI Framework for Visual Token Pruning** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.07193) - **Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.11660) - **LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure** (arXiv) - [Read more]
## Episode Summary In this episode, we cover: - **SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.10538) - **Thought-Level Beam Search for Reasoning** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.08020) - **Maglev: Sliding Recurrent Memory** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.02870) - **HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark** (arXiv) - [Read more](http://arxiv.org/abs/2
## Episode Summary In this episode, we cover: - **Mitigating Gender Bias in English to Romanian Machine Translation** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.08606) - **Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.29211) - **QuoteBench: How Matched Scores Can Hide Command-Path Failures** (arXiv) - [Read more](http://arxiv.org/abs/2608.13547v1) - **Vero: Can AI Agents Build Formally Verified Software Repositories?** (arXiv) - [Read more](http://arxiv.org/abs/2608
## Episode Summary In this episode, we cover: - **Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.12036) - **Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.09926) - **SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.05604) - **Convergent Detour Hijacking: Task-P
## Episode Summary In this episode, we cover: - **Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.08160) - **Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams** (arXiv) - [Read more](http://arxiv.org/abs/2608.12262v1) - **The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.06270) - **Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Co
## Episode Summary In this episode, we cover: - **ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.04956) - **MASS: Multiplayer World Models with Authoritative Shared State** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.06257) - **KVAE: Family of Tokenizers for Multimodal Generative Models** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.05798) - **FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for
## Episode Summary In this episode, we cover: - **Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.01481) - **TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories** (arXiv) - [Read more](http://arxiv.org/abs/2608.06346v1) - **Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.05784) - **EffectLearner: World-A
## Episode Summary In this episode, we cover: - **DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.03451) - **Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.05138) - **Learning When to Trust via Selective Context Preference Optimization** (arXiv) - [Read more](http://arxiv.org/abs/2608.06377v1) - **Tracing the Heart: An Evidence-
## Episode Summary In this episode, we cover: - **AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.24821) - **GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.03764) - **ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.05102) - **FinanceHarness: Autonomous Financial Deep Research Framework** (Hugging Face
## Episode Summary In this episode, we cover: - **WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament** (arXiv) - [Read more](http://arxiv.org/abs/2608.04008v1) - **ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.02703) - **When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2608.03506) - **CURV: Enhancing Chart Understanding Throu
## Episode Summary In this episode, we cover: - **Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.28802) - **GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.27042) - **Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.26326) - **Zero-Mem: Zero-Token Memory Operations for LLM Agents** (Hugging Face Daily) - [Read more](
## Episode Summary In this episode, we cover: - **SAF-OPD: Stable Advantage Fusion for On-Policy Distillation** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.29209) - **EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.28229) - **Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning** (Hugging Face Daily) - [Read more](https://huggingface.co/papers/2607.28478) - **CodeShrink: Adaptive Visual Compression for Efficient Multimodal Cod