Daily briefing

Papers fetched on 2026-06-30

Executive Signal

2026-06-30 is led by GUI agents, Mixture-of-Experts, and VAE encoder, with the strongest papers skewing toward production-minded advances that pair novelty with implementation value.

Top Papers

100/100Read

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

Published 2026-06-27 · Fetched 2026-06-30

Innovation Summary

Agentic Abstention: Do Agents Know When to Stop Instead of Act: We study this problem across web shopping, terminal environments, and question answering, evaluating 13 LLM-as-agent systems and 2 agent scaffolds on more than 28,000 tasks.

Executive Summary

Agentic Abstention: Do Agents Know When to Stop Instead of Act: We study this problem across web shopping, terminal environments, and question answering, evaluating 13 LLM-as-agent systems and 2 agent scaffolds on more than 28,000 tasks. Why it matters: Overall signal 100/100 driven by novelty 100 and practical impact 100. Primary categories: CONVOLVE, LLM-as-agent systems, agentic abstention, context engineering, question answering, sequential decision problem. Community signal includes 54 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 99/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 100/100 driven by novelty 100 and practical impact 100.
  • Primary categories: CONVOLVE, LLM-as-agent systems, agentic abstention, context engineering, question answering, sequential decision problem.
  • Community signal includes 54 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 99/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 100/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

CONVOLVE, LLM-as-agent systems, agentic abstention, context engineering, question answering, sequential decision problemJSON
100/100Read

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Published 2026-06-29 · Fetched 2026-06-30

Innovation Summary

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon.

Executive Summary

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. Why it matters: Overall signal 100/100 driven by novelty 100 and practical impact 100. Primary categories: Mixture-of-Experts, agent horizon, agentic model, agentic trajectories, domain-level teacher models, domain-routed training. Community signal includes 57 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 100/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 100/100 driven by novelty 100 and practical impact 100.
  • Primary categories: Mixture-of-Experts, agent horizon, agentic model, agentic trajectories, domain-level teacher models, domain-routed training.
  • Community signal includes 57 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 100/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 100/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

Mixture-of-Experts, agent horizon, agentic model, agentic trajectories, domain-level teacher models, domain-routed trainingJSON
99/100Read

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

Published 2026-06-26 · Fetched 2026-06-30

Innovation Summary

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents: We introduce TUA-Bench, a general-purpose benchmark for terminal-use agents.

Executive Summary

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents: We introduce TUA-Bench, a general-purpose benchmark for terminal-use agents. Why it matters: Overall signal 99/100 driven by novelty 100 and practical impact 100. Primary categories: benchmark evaluation, computer-use tasks, digital activities, execution-based scoring protocol, general-purpose agents, graphical user interfaces. Community signal includes 37 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 100/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 95/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 99/100 driven by novelty 100 and practical impact 100.
  • Primary categories: benchmark evaluation, computer-use tasks, digital activities, execution-based scoring protocol, general-purpose agents, graphical user interfaces.
  • Community signal includes 37 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 100/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 95/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 99/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

benchmark evaluation, computer-use tasks, digital activities, execution-based scoring protocol, general-purpose agents, graphical user interfacesJSON
97/100Read

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Published 2026-06-28 · Fetched 2026-06-30

Innovation Summary

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction: To address this gap, we introduce VG-GUIBench (Video-Guided GUI Benchmark), a new benchmark designed to evaluate whether MLLM-based GUI agents can follow video tutorials to complete.

Executive Summary

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction: To address this gap, we introduce VG-GUIBench (Video-Guided GUI Benchmark), a new benchmark designed to evaluate whether MLLM-based GUI agents can follow video tutorials to complete. Why it matters: Overall signal 97/100 driven by novelty 100 and practical impact 100. Primary categories: GUI agents, Multimodal Large Language Models, Video Question Answering, keyframe extraction, scene dynamics, task relevance. Community signal includes 13 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 97/100 driven by novelty 100 and practical impact 100.
  • Primary categories: GUI agents, Multimodal Large Language Models, Video Question Answering, keyframe extraction, scene dynamics, task relevance.
  • Community signal includes 13 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 97/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

GUI agents, Multimodal Large Language Models, Video Question Answering, keyframe extraction, scene dynamics, task relevanceJSON
96/100Read

ReFreeKV: Towards Threshold-Free KV Cache Compression

Published 2026-06-26 · Fetched 2026-06-30

Innovation Summary

ReFreeKV: Towards Threshold-Free KV Cache Compression: In this work, we propose a new objective that lifts the threshold constraints for robust KV compression, advocating for "threshold-free" methods that adaptively adjust budget allocation.

Executive Summary

ReFreeKV: Towards Threshold-Free KV Cache Compression: In this work, we propose a new objective that lifts the threshold constraints for robust KV compression, advocating for "threshold-free" methods that adaptively adjust budget allocation. Why it matters: Overall signal 96/100 driven by novelty 99 and practical impact 96. Primary categories: KV cache compression, KV cache pruning, LLM inference, adaptive budget allocation, threshold-free methods. Community signal includes 19 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 96/100 driven by novelty 99 and practical impact 96.
  • Primary categories: KV cache compression, KV cache pruning, LLM inference, adaptive budget allocation, threshold-free methods.
  • Community signal includes 19 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 96/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

KV cache compression, KV cache pruning, LLM inference, adaptive budget allocation, threshold-free methodsJSON

Additional Papers

TACO: Tool-Augmented Credit Optimization for Agentic Tool Use

Published 2026-06-29 · Fetched 2026-06-30

TACO: Tool-Augmented Credit Optimization for Agentic Tool Use: To address this, we introduce Tool-Augmented Credit Optimization (TACO), a GRPO variant for code-tool agents built on two coupled advantage channels.

95/100Read

Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting

Published 2026-06-29 · Fetched 2026-06-30

Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting: In this paper, we present Flux-GS, a real-time Gaussian Splatting method designed to achieve high-fidelity rendering with significantly reduced overhead for resource-constrained mobile platforms.

94/100Read

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Published 2026-06-29 · Fetched 2026-06-30

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark: We develop a new baseline method identifying the critical components to solve the NMO tasks, including a novel representation for modeling structural constraints and a domain-agnostic.

93/100Read

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Published 2026-06-29 · Fetched 2026-06-30

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots: As an attempt to address data challenge in GUI agents, we propose GUICrafter, a weakly-supervised GUI agent leveraging massive unannotated screenshots to substantially reduce the reliance.

93/100Read

AsyncOPD: How Stale Can On-Policy Distillation Be?

Published 2026-06-23 · Fetched 2026-06-30

AsyncOPD: How Stale Can On-Policy Distillation Be: We present the first systematic study of staleness in asynchronous OPD, focusing on a practical setting where teacher feedback is implemented through local KL losses and.

92/100Read

TheoremGraph: Bridging Formal and Informal Mathematics

Published 2026-06-24 · Fetched 2026-06-30

TheoremGraph: Bridging Formal and Informal Mathematics: We introduce TheoremGraph, a unified statement-level dependency graph spanning both informal and formal mathematics.

91/100Read

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

Published 2026-06-26 · Fetched 2026-06-30

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning: To isolate this capability, we introduce Video-MME-Logical, a controlled benchmark organized around five temporal-logical operations: state tracking, sequential counting, temporal ordering, dynamic spatiality, and structural composition.

91/100Read

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

Published 2026-06-29 · Fetched 2026-06-30

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation: In this paper, we introduce ILLUME-X, an advanced unified multimodal paradigm that enables high-quality, free-form interleaved text-image generation by improving multimodal data efficiency and stabilizing the.

90/100Read

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Published 2026-06-29 · Fetched 2026-06-30

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing: To systematically evaluate this capability, we introduce SafePyramid, a safety benchmark comprising 1,000 multi-turn conversations across 10 domains and 3,000 corresponding application-specific policies, which together contain.

90/100Read

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval

Published 2026-06-26 · Fetched 2026-06-30

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval: We present ZooClaw-FashionSigLIP2, a fashion-specialized SigLIP2-base model that resolves this tradeoff with a simple recipe -- full fine-tuning with knowledge distillation on curated in-domain data, followed.

89/100Read

How Good Can Linear Models Be for Time-Series Forecasting?

Published 2026-06-25 · Fetched 2026-06-30

How Good Can Linear Models Be for Time-Series Forecasting: The resulting models beat prior linear forecasters on most dataset-horizon entries and exceed Transformer, MLP, and CNN baselines on six of eight benchmarks.

88/100Read

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

Published 2026-06-25 · Fetched 2026-06-30

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing: In this work, we present a novel streaming video editing framework that performs causal, frame-by-frame editing with strong content preservation and real-time responsiveness.

88/100Read

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

Published 2026-06-29 · Fetched 2026-06-30

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding: To retain the accuracy benefits of two-pass zooming without this extra cost, we propose InnerZoom, a single-forward framework for cross-layer evidence bridging.

88/100Read

Trimming the Long-Tail of Visual World Modeling Evaluation

Published 2026-06-23 · Fetched 2026-06-30

Trimming the Long-Tail of Visual World Modeling Evaluation: In this work, we introduce Tailor-Bench, a benchmark that challenges world models to simulate irregular physical interactions.

88/100Read

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation

Published 2026-06-22 · Fetched 2026-06-30

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation: To address these challenges, we propose RaysUp, an ultra-lightweight, task-agnostic, and VFM-agnostic feature upsampling framework that reconstructs high-resolution feature maps at arbitrary resolutions.

83/100Read

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

Published 2026-06-25 · Fetched 2026-06-30

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE: To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for.

80/100Read

Interleaved Speech Language Models Latently Work In Text

Published 2026-06-21 · Fetched 2026-06-30

Interleaved Speech Language Models Latently Work In Text: Speech language models (SLMs) have been extensively studied, with the common paradigm incorporating text data and pre-trained text LMs.

80/100Read

Beyond IID: How General Are Tabular Foundation Models, Really?

Published 2026-06-29 · Fetched 2026-06-30

Beyond IID: How General Are Tabular Foundation Models, Really: To enable unified benchmarking beyond standard benchmarks, we introduce Data Foundry, a Python framework and metadata schema for curating tabular datasets for predictive machine learning.

62/100Worth Watching

Watchlist

Archive

Daily record count: 36. Persistent paper JSON lives under public data.

  1. Agentic Abstention: Do Agents Know When to Stop Instead of Act?Published 2026-06-27 · 100/100 · Read
  2. Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentPublished 2026-06-29 · 100/100 · Read
  3. TUA-Bench: A Benchmark for General-Purpose Terminal-Use AgentsPublished 2026-06-26 · 99/100 · Read
  4. Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe ExtractionPublished 2026-06-28 · 97/100 · Read
  5. ReFreeKV: Towards Threshold-Free KV Cache CompressionPublished 2026-06-26 · 96/100 · Read
  6. OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World TasksPublished 2026-06-28 · 95/100 · Read
  7. TACO: Tool-Augmented Credit Optimization for Agentic Tool UsePublished 2026-06-29 · 95/100 · Read
  8. Monte Carlo Energy Aggregation for Mobile 3D Gaussian SplattingPublished 2026-06-29 · 94/100 · Read
  9. Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkPublished 2026-06-29 · 93/100 · Read
  10. GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated ScreenshotsPublished 2026-06-29 · 93/100 · Read
  11. AsyncOPD: How Stale Can On-Policy Distillation Be?Published 2026-06-23 · 92/100 · Read
  12. DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World ModelPublished 2026-06-29 · 91/100 · Read
  13. TheoremGraph: Bridging Formal and Informal MathematicsPublished 2026-06-24 · 91/100 · Read
  14. Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical ReasoningPublished 2026-06-26 · 91/100 · Read
  15. Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image GenerationPublished 2026-06-29 · 90/100 · Read
  16. Large-Scale Tunnel Air-Ground Collaboration With FLISP: Fast LiDAR-IMU Synchronized Path PlannerPublished 2026-06-25 · 90/100 · Read
  17. PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentsPublished 2026-06-28 · 90/100 · Read
  18. SafePyramid: A Hierarchical Benchmark for In-context Policy GuardrailingPublished 2026-06-29 · 90/100 · Read
  19. ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion RetrievalPublished 2026-06-26 · 89/100 · Read
  20. How Good Can Linear Models Be for Time-Series Forecasting?Published 2026-06-25 · 88/100 · Read
  21. LiveEdit: Towards Real-Time Diffusion-Based Streaming Video EditingPublished 2026-06-25 · 88/100 · Read
  22. One Forward Beats Two: InnerZoom for Accurate and Efficient GUI GroundingPublished 2026-06-29 · 88/100 · Read
  23. Trimming the Long-Tail of Visual World Modeling EvaluationPublished 2026-06-23 · 88/100 · Read
  24. Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image SynthesisPublished 2026-06-29 · 87/100 · Read
  25. Walking in the Implicit: Interactive World Exploration via Neural Scene RepresentationPublished 2026-06-29 · 86/100 · Read
  26. MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image GenerationPublished 2026-06-24 · 85/100 · Read
  27. RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray RepresentationPublished 2026-06-22 · 83/100 · Read
  28. Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoEPublished 2026-06-25 · 80/100 · Read
  29. Interleaved Speech Language Models Latently Work In TextPublished 2026-06-21 · 80/100 · Read
  30. The Surprising Effectiveness of Video Diffusion Models for Hand Motion ReconstructionPublished 2026-06-29 · 80/100 · Read
  31. Learning Transferable Dynamics Priors from Action to World ModelingPublished 2026-06-28 · 75/100 · Worth Watching
  32. PoseShield: Neural Collision Fields for Human Self-Collision ResolutionPublished 2026-06-29 · 74/100 · Worth Watching
  33. Geometric Stability of Neural Population Codes: Regional Variation, Behavioral Relevance, and Circuit DependencePublished 2026-06-28 · 68/100 · Worth Watching
  34. Beyond IID: How General Are Tabular Foundation Models, Really?Published 2026-06-29 · 62/100 · Worth Watching
  35. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty PredictionPublished 2026-06-26 · 62/100 · Worth Watching
  36. ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning ModelsPublished 2026-06-22 · 49/100 · Skip