Daily briefing

Papers fetched on 2026-07-01

Executive Signal

2026-07-01 is led by large language models, diffusion models, and 3D Gaussians, with the strongest papers skewing toward production-minded advances that pair novelty with implementation value.

Top Papers

99/100Read

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

Published 2026-06-27 · Fetched 2026-07-01

Innovation Summary

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks: To address this, we introduce Evolution Fine-Tuning (EFT), a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision.

Executive Summary

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks: To address this, we introduce Evolution Fine-Tuning (EFT), a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision. Why it matters: Overall signal 99/100 driven by novelty 100 and practical impact 100. Primary categories: cross-task generalization, evolutionary fine-tuning, evolutionary search, large language models, mathematical conjectures, optimization tasks. Community signal includes 18 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 99/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 95/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 99/100 driven by novelty 100 and practical impact 100.
  • Primary categories: cross-task generalization, evolutionary fine-tuning, evolutionary search, large language models, mathematical conjectures, optimization tasks.
  • Community signal includes 18 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 99/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 95/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 99/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

cross-task generalization, evolutionary fine-tuning, evolutionary search, large language models, mathematical conjectures, optimization tasksJSON
98/100Read

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

Published 2026-06-30 · Fetched 2026-07-01

Innovation Summary

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding: In this paper, we show that this assumption is suboptimal, as the optimal block size varies across samples and plays a critical role in speculative decoding.

Executive Summary

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding: In this paper, we show that this assumption is suboptimal, as the optimal block size varies across samples and plays a critical role in speculative decoding. Why it matters: Overall signal 98/100 driven by novelty 100 and practical impact 100. Primary categories: block-level diffusion, diffusion-based speculative decoding, draft model, inference block size, instance-adaptive decision mechanism, policy learning. Community signal includes 64 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 91/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 98/100 driven by novelty 100 and practical impact 100.
  • Primary categories: block-level diffusion, diffusion-based speculative decoding, draft model, inference block size, instance-adaptive decision mechanism, policy learning.
  • Community signal includes 64 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 91/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 98/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

block-level diffusion, diffusion-based speculative decoding, draft model, inference block size, instance-adaptive decision mechanism, policy learningJSON
98/100Read

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

Published 2026-06-25 · Fetched 2026-07-01

Innovation Summary

RedVox: Safety and Fairness Gaps in Speech Models Across Languages: To address this gap, we introduce RedVox, a multilingual safety and fairness benchmark for audio and speech built on real voices, covering unsafe and unfair stereotypical.

Executive Summary

RedVox: Safety and Fairness Gaps in Speech Models Across Languages: To address this gap, we introduce RedVox, a multilingual safety and fairness benchmark for audio and speech built on real voices, covering unsafe and unfair stereotypical. Why it matters: Overall signal 98/100 driven by novelty 100 and practical impact 100. Primary categories: audio, fairness benchmark, multilingual safety, naturalistic conditions, speech models, speech-capable models. Community signal includes 11 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 100/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 98/100 driven by novelty 100 and practical impact 100.
  • Primary categories: audio, fairness benchmark, multilingual safety, naturalistic conditions, speech models, speech-capable models.
  • Community signal includes 11 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 100/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 98/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

audio, fairness benchmark, multilingual safety, naturalistic conditions, speech models, speech-capable modelsJSON
95/100Read

Orca: The World is in Your Mind

Published 2026-06-29 · Fetched 2026-07-01

Innovation Summary

Orca: The World is in Your Mind: We introduce Orca, an initial instantiation of a general world foundation model.

Executive Summary

Orca: The World is in Your Mind: We introduce Orca, an initial instantiation of a general world foundation model. Why it matters: Overall signal 95/100 driven by novelty 100 and practical impact 100. Primary categories: conscious learning, downstream readouts, embodied action generation, modality-specific decoders, multimodal readout interfaces, next-state-prediction modeling. Community signal includes 164 upvote(s) and 5 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 65/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 95/100 driven by novelty 100 and practical impact 100.
  • Primary categories: conscious learning, downstream readouts, embodied action generation, modality-specific decoders, multimodal readout interfaces, next-state-prediction modeling.
  • Community signal includes 164 upvote(s) and 5 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 65/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 95/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

conscious learning, downstream readouts, embodied action generation, modality-specific decoders, multimodal readout interfaces, next-state-prediction modelingJSON
94/100Read

Dockerless: Environment-Free Program Verifier for Coding Agents

Published 2026-06-26 · Fetched 2026-07-01

Innovation Summary

Dockerless: Environment-Free Program Verifier for Coding Agents: We propose Dockerless, an environment-free agentic patch verifier that evaluates generated code patches without executing them.

Executive Summary

Dockerless: Environment-Free Program Verifier for Coding Agents: We propose Dockerless, an environment-free agentic patch verifier that evaluates generated code patches without executing them. Why it matters: Overall signal 94/100 driven by novelty 95 and practical impact 100. Primary categories: Dockerless, Multilingual, Pro, SWE-bench Verified, agentic patch verifier, environment-free. Community signal includes 78 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 87/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 94/100 driven by novelty 95 and practical impact 100.
  • Primary categories: Dockerless, Multilingual, Pro, SWE-bench Verified, agentic patch verifier, environment-free.
  • Community signal includes 78 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 87/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 94/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

Dockerless, Multilingual, Pro, SWE-bench Verified, agentic patch verifier, environment-freeJSON

Additional Papers

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation

Published 2026-06-22 · Fetched 2026-07-01

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation: We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks,.

94/100Read

Multi-Block Diffusion Language Models

Published 2026-06-30 · Fetched 2026-07-01

Multi-Block Diffusion Language Models: To bridge this gap, we propose Multi-Block Diffusion Language Models (MBD-LMs), obtained by post-training BD-LMs with Multi-block Teacher Forcing (MultiTF).

94/100Read

DOPD: Dual On-policy Distillation

Published 2026-06-29 · Fetched 2026-07-01

DOPD: Dual On-policy Distillation: To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their.

93/100Read

Little Brains, Big Feats: Exploring Compact Language Models

Published 2026-06-29 · Fetched 2026-07-01

Little Brains, Big Feats: Exploring Compact Language Models: In this study, we investigate how smaller language models perform during the generation stage within a Retrieval-Augmented Generation (RAG) system.

90/100Read

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation

Published 2026-06-25 · Fetched 2026-07-01

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation: Built upon this novel continuous mesh representation, we present PolyFlow, a Transformer-based flow-matching framework that achieves fully parallel vertex state denoising conditioned on extracted point-cloud features.

90/100Read

GEAR: Guided End-to-End AutoRegression for Image Synthesis

Published 2026-06-30 · Fetched 2026-07-01

GEAR: Guided End-to-End AutoRegression for Image Synthesis: We present GEAR (Guided End-to-end AutoRegression), which trains a vector-quantized (VQ) tokenizer and an autoregressive (AR) generator jointly and end-to-end, guided by representation alignment.

89/100Read

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Published 2026-06-29 · Fetched 2026-07-01

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation: Inspired by recent advancements in one-dimensional visual tokenization, we present AVTok, a novel unified tokenizer designated for holistic audio-video generation.

88/100Read

MuSViT: A Foundation Vision Model for Sheet Music Representation

Published 2026-06-30 · Fetched 2026-07-01

MuSViT: A Foundation Vision Model for Sheet Music Representation: We introduce MuSViT (Music Score Vision Transformer): the first foundation vision model for sheet music representation -- a ViT encoder pre-trained via Masked Autoencoders on 9.

82/100Read

Xiaomi-GUI-0 Technical Report

Published 2026-06-30 · Fetched 2026-07-01

Xiaomi-GUI-0 Technical Report: To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop.

73/100Worth Watching

Watchlist

No papers in this section.

Archive

Daily record count: 23. Persistent paper JSON lives under public data.

  1. Evolution Fine-Tuning: Learning to Discover Across 371 Optimization TasksPublished 2026-06-27 · 99/100 · Read
  2. BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative DecodingPublished 2026-06-30 · 98/100 · Read
  3. RedVox: Safety and Fairness Gaps in Speech Models Across LanguagesPublished 2026-06-25 · 98/100 · Read
  4. Orca: The World is in Your MindPublished 2026-06-29 · 95/100 · Read
  5. Dockerless: Environment-Free Program Verifier for Coding AgentsPublished 2026-06-26 · 94/100 · Read
  6. Managing Procedural Memory in LLM Agents: Control, Adaptation, and EvaluationPublished 2026-06-22 · 94/100 · Read
  7. MemLearner: Learning to Query Context memory for Video World ModelsPublished 2026-06-30 · 94/100 · Read
  8. Multi-Block Diffusion Language ModelsPublished 2026-06-30 · 94/100 · Read
  9. DOPD: Dual On-policy DistillationPublished 2026-06-29 · 93/100 · Read
  10. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMsPublished 2026-06-30 · 93/100 · Read
  11. SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision HistoryPublished 2026-06-23 · 92/100 · Read
  12. DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image GenerationPublished 2026-06-30 · 91/100 · Read
  13. LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI AgentsPublished 2026-06-29 · 90/100 · Read
  14. Little Brains, Big Feats: Exploring Compact Language ModelsPublished 2026-06-29 · 90/100 · Read
  15. PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh GenerationPublished 2026-06-25 · 90/100 · Read
  16. Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed ViewsPublished 2026-06-28 · 90/100 · Read
  17. BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and LanguagePublished 2026-06-29 · 89/100 · Read
  18. GEAR: Guided End-to-End AutoRegression for Image SynthesisPublished 2026-06-30 · 89/100 · Read
  19. AVTok: 1D Unified Tokenization for Holistic Audio-Video GenerationPublished 2026-06-29 · 88/100 · Read
  20. PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled DenoisingPublished 2026-06-29 · 85/100 · Read
  21. TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial PrimitivePublished 2026-06-30 · 85/100 · Read
  22. MuSViT: A Foundation Vision Model for Sheet Music RepresentationPublished 2026-06-30 · 82/100 · Read
  23. Xiaomi-GUI-0 Technical ReportPublished 2026-06-30 · 73/100 · Worth Watching