Daily briefing

Papers fetched on 2026-07-21

Executive Signal

2026-07-21 is led by EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in, Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning, and RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model, with the strongest papers skewing toward production-minded advances that pair novelty with implementation value.

Top Papers

97/100Read

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

Published 2026-07-19 · Fetched 2026-07-21

Innovation Summary

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World: This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds.

Executive Summary

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World: This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Why it matters: Overall signal 97/100 driven by novelty 100 and practical impact 100. It maps to cross-cutting AI systems work even without explicit category metadata. Community signal includes 65 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 97/100 driven by novelty 100 and practical impact 100.
  • It maps to cross-cutting AI systems work even without explicit category metadata.
  • Community signal includes 65 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 97/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

N/AJSON
96/100Read

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Published 2026-07-15 · Fetched 2026-07-21

Innovation Summary

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning: We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training.

Executive Summary

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning: We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Why it matters: Overall signal 96/100 driven by novelty 97 and practical impact 100. It maps to cross-cutting AI systems work even without explicit category metadata. Community signal includes 15 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 91/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 96/100 driven by novelty 97 and practical impact 100.
  • It maps to cross-cutting AI systems work even without explicit category metadata.
  • Community signal includes 15 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 91/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 96/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

N/AJSON
96/100Read

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Published 2026-07-20 · Fetched 2026-07-21

Innovation Summary

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model: We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales.

Executive Summary

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model: We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Why it matters: Overall signal 96/100 driven by novelty 100 and practical impact 100. It maps to cross-cutting AI systems work even without explicit category metadata. Community signal includes 31 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 77/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 96/100 driven by novelty 100 and practical impact 100.
  • It maps to cross-cutting AI systems work even without explicit category metadata.
  • Community signal includes 31 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 77/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 96/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

N/AJSON
96/100Read

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Published 2026-07-20 · Fetched 2026-07-21

Innovation Summary

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune: Based on this finding, we propose SWE-Pruner Pro, which prunes tool outputs directly inside the agent.

Executive Summary

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune: Based on this finding, we propose SWE-Pruner Pro, which prunes tool outputs directly inside the agent. Why it matters: Overall signal 96/100 driven by novelty 95 and practical impact 100. It maps to cross-cutting AI systems work even without explicit category metadata. Community signal includes 64 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 96/100 driven by novelty 95 and practical impact 100.
  • It maps to cross-cutting AI systems work even without explicit category metadata.
  • Community signal includes 64 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 96/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

N/AJSON
96/100Read

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Published 2026-07-19 · Fetched 2026-07-21

Innovation Summary

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs: We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints.

Executive Summary

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs: We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Why it matters: Overall signal 96/100 driven by novelty 100 and practical impact 82. It maps to cross-cutting AI systems work even without explicit category metadata. Community signal includes 104 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 99/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 96/100 driven by novelty 100 and practical impact 82.
  • It maps to cross-cutting AI systems work even without explicit category metadata.
  • Community signal includes 104 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 99/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 96/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

N/AJSON

Additional Papers

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

Published 2026-07-17 · Fetched 2026-07-21

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models: To address these challenges, we present JoyNexus, a unified service for multi-tenant VLA supervised fine-tuning, reinforcement learning, and evaluation.

93/100Read

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?

Published 2026-07-20 · Fetched 2026-07-21

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?: Formally, we characterize a four-axis attack space (Target, Mechanism, Granularity, Temporal); investigate the structural limits of prevention, detection, and recovery; and introduce a workload-conditioned view of.

93/100Read

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Published 2026-07-20 · Fetched 2026-07-21

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications: We present FlashRT, an agent harness that guides coding agents to lift simple developer-written reference implementations into optimized multi-GPU deployments that flexibly weigh target metrics like.

90/100Read

Group Entropy-Controlled Policy Optimization

Published 2026-07-18 · Fetched 2026-07-21

Group Entropy-Controlled Policy Optimization: To address this issue, we propose Group Entropy-Controlled Policy Optimization (GEPO), a lightweight extension to GRPO that uses group entropy, estimated from existing grouped samples to.

90/100Read

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation

Published 2026-07-15 · Fetched 2026-07-21

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation: To address this, we propose Differentiable Geometry Image (DiffGI), an end-to-end 3D-to-2D mapping framework that seamlessly integrates surface representation and geometric optimization.

89/100Read

ShotPlan: Cinematic Video Generation with Learnable Planning Token

Published 2026-07-20 · Fetched 2026-07-21

ShotPlan: Cinematic Video Generation with Learnable Planning Token: To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model.

87/100Read

Distilled Reinforcement Learning for LLM Post-training

Published 2026-07-19 · Fetched 2026-07-21

Distilled Reinforcement Learning for LLM Post-training: Through a concise and interpretable case study, we demonstrate that Distilled RL can effectively transfer previously unavailable knowledge from a teacher model to a student model.

85/100Read

DiFA: Inference-Time Forward-Process Alignment for Diffusion Models

Published 2026-07-20 · Fetched 2026-07-21

DiFA: Inference-Time Forward-Process Alignment for Diffusion Models: In this work, we propose Forward-Process Aligned Diffusion prediction (DiFA), a training-free framework that reframes inference-time data prediction refinement as a sequential state estimation problem.

83/100Read

Can Multimodal Large Language Models Understand OCT?

Published 2026-07-18 · Fetched 2026-07-21

Can Multimodal Large Language Models Understand OCT?: To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding.

82/100Read

OpenLongTail: Generative Scaling of Long-Tail Driving Data

Published 2026-07-10 · Fetched 2026-07-21

OpenLongTail: Generative Scaling of Long-Tail Driving Data: To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline.

82/100Read

Environment-free Synthetic Data Generation for API-Calling Agents

Published 2026-07-18 · Fetched 2026-07-21

Environment-free Synthetic Data Generation for API-Calling Agents: To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models.

66/100Worth Watching

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Published 2026-07-08 · Fetched 2026-07-21

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment: We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools.

61/100Worth Watching

GigaChat Audio: Time-aware Large Audio Language Model

Published 2026-07-11 · Fetched 2026-07-21

GigaChat Audio: Time-aware Large Audio Language Model: We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input.

60/100Worth Watching

Watchlist

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Published 2026-07-11 · Fetched 2026-07-21

GigaAM Multilingual: Foundation Model for Underrepresented Languages: Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance.

55/100Skip

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

Published 2026-07-20 · Fetched 2026-07-21

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry: We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate.

52/100Skip

Archive

Daily record count: 31. Persistent paper JSON lives under public data.

  1. EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldPublished 2026-07-19 · 97/100 · Read
  2. Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied LearningPublished 2026-07-15 · 96/100 · Read
  3. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation ModelPublished 2026-07-20 · 96/100 · Read
  4. SWE-Pruner Pro: The Coder LLM Already Knows What to PrunePublished 2026-07-20 · 96/100 · Read
  5. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMsPublished 2026-07-19 · 96/100 · Read
  6. Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical IntelligencePublished 2026-07-17 · 95/100 · Read
  7. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent EnchancementPublished 2026-07-20 · 94/100 · Read
  8. JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA ModelsPublished 2026-07-17 · 93/100 · Read
  9. Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?Published 2026-07-20 · 93/100 · Read
  10. HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction SynthesisPublished 2026-07-19 · 92/100 · Read
  11. FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationsPublished 2026-07-20 · 90/100 · Read
  12. Group Entropy-Controlled Policy OptimizationPublished 2026-07-18 · 90/100 · Read
  13. Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy DistillationPublished 2026-07-15 · 89/100 · Read
  14. DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D GenerationPublished 2026-07-15 · 89/100 · Read
  15. The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer ArchitecturePublished 2026-07-19 · 89/100 · Read
  16. LLM-as-a-Coach: Experiential Learning for Non-Verifiable TasksPublished 2026-07-20 · 88/100 · Read
  17. ShotPlan: Cinematic Video Generation with Learnable Planning TokenPublished 2026-07-20 · 87/100 · Read
  18. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingPublished 2026-07-20 · 87/100 · Read
  19. Distilled Reinforcement Learning for LLM Post-trainingPublished 2026-07-19 · 85/100 · Read
  20. ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric VideoPublished 2026-07-20 · 85/100 · Read
  21. DiFA: Inference-Time Forward-Process Alignment for Diffusion ModelsPublished 2026-07-20 · 83/100 · Read
  22. Can Multimodal Large Language Models Understand OCT?Published 2026-07-18 · 82/100 · Read
  23. OpenLongTail: Generative Scaling of Long-Tail Driving DataPublished 2026-07-10 · 82/100 · Read
  24. Token-Level Off-Policy Learning for Faithful Generation Under Distribution ShiftPublished 2026-07-20 · 76/100 · Worth Watching
  25. ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video StreamsPublished 2026-07-14 · 67/100 · Worth Watching
  26. Environment-free Synthetic Data Generation for API-Calling AgentsPublished 2026-07-18 · 66/100 · Worth Watching
  27. Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial ConstraintsPublished 2026-07-20 · 65/100 · Worth Watching
  28. DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable EnvironmentPublished 2026-07-08 · 61/100 · Worth Watching
  29. GigaChat Audio: Time-aware Large Audio Language ModelPublished 2026-07-11 · 60/100 · Worth Watching
  30. GigaAM Multilingual: Foundation Model for Underrepresented LanguagesPublished 2026-07-11 · 55/100 · Skip
  31. FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality MimicryPublished 2026-07-20 · 52/100 · Skip