Daily briefing

Papers fetched on 2026-07-02

Executive Signal

2026-07-02 is led by reinforcement learning, 3D scene graph generation, and Agentic Cross-Verification, with the strongest papers skewing toward production-minded advances that pair novelty with implementation value.

Top Papers

98/100Read

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Published 2026-07-01 · Fetched 2026-07-02

Innovation Summary

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory: To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems.

Executive Summary

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory: To bridge this gap, we propose MemSyco-Bench, a comprehensive benchmark for evaluating memory-induced sycophancy in agent systems. Why it matters: Overall signal 98/100 driven by novelty 100 and practical impact 100. Primary categories: LLM-based agents, MemSyco-Bench, decision-making, downstream reasoning, factual accuracy, memory. Community signal includes 17 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 98/100 driven by novelty 100 and practical impact 100.
  • Primary categories: LLM-based agents, MemSyco-Bench, decision-making, downstream reasoning, factual accuracy, memory.
  • Community signal includes 17 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 98/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

LLM-based agents, MemSyco-Bench, decision-making, downstream reasoning, factual accuracy, memoryJSON
97/100Read

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

Published 2026-07-01 · Fetched 2026-07-02

Innovation Summary

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving: We present ELDR, an expert-locality-aware decode router for PD-disaggregated MoE serving.

Executive Summary

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving: We present ELDR, an expert-locality-aware decode router for PD-disaggregated MoE serving. Why it matters: Overall signal 97/100 driven by novelty 99 and practical impact 100. Primary categories: K-means, KV cache, TPOT, decode router, disaggregated LLM serving, expert-locality-aware. Community signal includes 16 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 97/100 driven by novelty 99 and practical impact 100.
  • Primary categories: K-means, KV cache, TPOT, decode router, disaggregated LLM serving, expert-locality-aware.
  • Community signal includes 16 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 83/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 97/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

K-means, KV cache, TPOT, decode router, disaggregated LLM serving, expert-locality-awareJSON
96/100Read

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

Published 2026-06-26 · Fetched 2026-07-02

Innovation Summary

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception: We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness.

Executive Summary

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception: We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Why it matters: Overall signal 96/100 driven by novelty 100 and practical impact 100. Primary categories: Circular Peer-Review consensus, Easy-Wrong, Must-Right, Open-Closed Stratification, Reliability Gap, atomic auditing. Community signal includes 26 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 87/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 87/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 96/100 driven by novelty 100 and practical impact 100.
  • Primary categories: Circular Peer-Review consensus, Easy-Wrong, Must-Right, Open-Closed Stratification, Reliability Gap, atomic auditing.
  • Community signal includes 26 upvote(s) and 2 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 87/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 87/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 96/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

Circular Peer-Review consensus, Easy-Wrong, Must-Right, Open-Closed Stratification, Reliability Gap, atomic auditingJSON
95/100Read

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

Published 2026-07-01 · Fetched 2026-07-02

Innovation Summary

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts: To reduce the burden of data curation and training, we propose an analogy-based method that adapts VLA models under environmental shifts through weight vector arithmetic with.

Executive Summary

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts: To reduce the burden of data curation and training, we propose an analogy-based method that adapts VLA models under environmental shifts through weight vector arithmetic with. Why it matters: Overall signal 95/100 driven by novelty 100 and practical impact 100. Primary categories: Vision-Language-Action models, domain-specific information, embodiment shifts, environmental shifts, one-shot adaptation, subspace alignment. Community signal includes 15 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 73/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Why It Matters

  • Overall signal 95/100 driven by novelty 100 and practical impact 100.
  • Primary categories: Vision-Language-Action models, domain-specific information, embodiment shifts, environmental shifts, one-shot adaptation, subspace alignment.
  • Community signal includes 15 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 73/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

No linked implementation is available yet, which raises integration cost and lowers reproducibility confidence.

Estimated Reading Priority

High - 95/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

Vision-Language-Action models, domain-specific information, embodiment shifts, environmental shifts, one-shot adaptation, subspace alignmentJSON
95/100Read

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Published 2026-07-01 · Fetched 2026-07-02

Innovation Summary

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning: In this paper, we propose Perceive-to-Reason (P2R), a unified framework that formulates fine-grained visual reasoning as a two-stage process: the model first localizes question-relevant evidence as.

Executive Summary

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning: In this paper, we propose Perceive-to-Reason (P2R), a unified framework that formulates fine-grained visual reasoning as a two-stage process: the model first localizes question-relevant evidence as. Why it matters: Overall signal 95/100 driven by novelty 100 and practical impact 100. Primary categories: Perceiver, Perception-Reasoning Alternating GRPO, Reasoner, fine-grained visual reasoning, multimodal reasoning, reinforcement learning. Community signal includes 10 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity. Implementation angle: Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows. No linked repository is present, so expect more translation work before the ideas are production-ready. Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work. Caveat: Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Why It Matters

  • Overall signal 95/100 driven by novelty 100 and practical impact 100.
  • Primary categories: Perceiver, Perception-Reasoning Alternating GRPO, Reasoner, fine-grained visual reasoning, multimodal reasoning, reinforcement learning.
  • Community signal includes 10 upvote(s) and 1 comment(s), which helps separate durable interest from title-only curiosity.

Implementation Angle

  • Implementation potential scores 89/100; prioritize adaptation paths for internal agent, evaluation, or platform workflows.
  • No linked repository is present, so expect more translation work before the ideas are production-ready.
  • Technical depth scores 100/100, so a quick skim should focus on architecture, data, and evaluation sections before full adoption work.

Caveat

Evidence appears benchmark-centric, so verify transfer to production workloads before acting on the claims.

Estimated Reading Priority

High - 95/100 signal; read before acting on adjacent agent, evaluation, inference, or ML systems work.

Links

Perceiver, Perception-Reasoning Alternating GRPO, Reasoner, fine-grained visual reasoning, multimodal reasoning, reinforcement learningJSON

Additional Papers

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Published 2026-07-01 · Fetched 2026-07-02

ABot-M05: Unified Mobility-and-Manipulation World Action Model: To align inference conditions, we propose the dream-forcing training strategy that progressively trains inverse dynamics on model-predicted videos, improving train-test alignment and robustness during autoregressive prediction.

94/100Read

ASPIRE: Agentic /Skills Discovery for Robotics

Published 2026-06-30 · Fetched 2026-07-02

ASPIRE: Agentic /Skills Discovery for Robotics: We introduce ASPIRE (Agentic Skill Programming through Iterative Robot Exploration), a continual learning system that autonomously writes and refines robot control programs in a code-as-policy paradigm.

93/100Read

NoPA: Non-Parametric Online 3D Scene Graph Generation

Published 2026-07-01 · Fetched 2026-07-02

NoPA: Non-Parametric Online 3D Scene Graph Generation: To address these issues, we propose NoPA, which represents each object as a separate non-parametric distribution.

89/100Read

Autonomous Scientific Discovery via Iterative Meta-Reflection

Published 2026-07-01 · Fetched 2026-07-02

Autonomous Scientific Discovery via Iterative Meta-Reflection: We introduce DiscoPER, an autonomous large language model-powered framework that conducts open-ended research by dynamically generating and executing code to explore datasets without pre-specified research objectives.

88/100Read

The State-Prediction Separation Hypothesis

Published 2026-07-01 · Fetched 2026-07-02

The State-Prediction Separation Hypothesis: We formulate the state-prediction separation hypothesis: disentangling the two roles yields better language modeling performance.

88/100Read

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Published 2026-07-01 · Fetched 2026-07-02

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination: We develop Graph-PRefLexOR, a family of graph-native reasoning models fine-tuned with Group Relative Policy Optimization (GRPO) to organize reasoning into explicit phases for mechanism exploration, graph.

87/100Read

Valdi: Value Diffusion World Models

Published 2026-07-01 · Fetched 2026-07-02

Valdi: Value Diffusion World Models: In preliminary experiments on the CarRacing environment, we show that Valdi, using a single diffusion step at both training and inference, matches a deterministic MLP baseline.

82/100Read

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

Published 2026-06-30 · Fetched 2026-07-02

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously: We present AutoTrainess, a LM agent that exposes these operations as a repository of agent-computer interfaces for planning, data preparation, training, evaluation, and logging.

66/100Worth Watching

Watchlist

Archive

Daily record count: 24. Persistent paper JSON lives under public data.

  1. MemSyco-Bench: Benchmarking Sycophancy in Agent MemoryPublished 2026-07-01 · 98/100 · Read
  2. ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE ServingPublished 2026-07-01 · 97/100 · Read
  3. PerceptionRubrics: Calibrating Multimodal Evaluation to Human PerceptionPublished 2026-06-26 · 96/100 · Read
  4. Domain Arithmetic: One-Shot VLA Adaptation under Environmental ShiftsPublished 2026-07-01 · 95/100 · Read
  5. Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual ReasoningPublished 2026-07-01 · 95/100 · Read
  6. ABot-M0.5: Unified Mobility-and-Manipulation World Action ModelPublished 2026-07-01 · 94/100 · Read
  7. ASPIRE: Agentic /Skills Discovery for RoboticsPublished 2026-06-30 · 93/100 · Read
  8. Multimodal Continuous Reasoning via Asymmetric Mutual Variational LearningPublished 2026-07-01 · 93/100 · Read
  9. TurboServe: Serving Streaming Video Generation Efficiently and EconomicallyPublished 2026-06-17 · 92/100 · Read
  10. PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence SeekingPublished 2026-06-30 · 91/100 · Read
  11. Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT NetworksPublished 2026-07-01 · 89/100 · Read
  12. NoPA: Non-Parametric Online 3D Scene Graph GenerationPublished 2026-07-01 · 89/100 · Read
  13. AI translation of literary texts is "fine", but readers still prefer human translationsPublished 2026-06-24 · 88/100 · Read
  14. Autonomous Scientific Discovery via Iterative Meta-ReflectionPublished 2026-07-01 · 88/100 · Read
  15. The State-Prediction Separation HypothesisPublished 2026-07-01 · 88/100 · Read
  16. AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report EvaluationPublished 2026-06-30 · 87/100 · Read
  17. Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual RecombinationPublished 2026-07-01 · 87/100 · Read
  18. Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural DenoisingPublished 2026-07-01 · 87/100 · Read
  19. BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge DiscoveryPublished 2026-06-19 · 86/100 · Read
  20. CausalMix: Data Mixture as Causal Inference for Language Model TrainingPublished 2026-07-01 · 86/100 · Read
  21. When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing ErrorsPublished 2026-06-30 · 85/100 · Read
  22. Valdi: Value Diffusion World ModelsPublished 2026-07-01 · 82/100 · Read
  23. AutoTrainess: Teaching Language Models to Improve Language Models AutonomouslyPublished 2026-06-30 · 66/100 · Worth Watching
  24. Seed2.0 Model Card: Towards Intelligence Frontier for Real-World ComplexityPublished 2026-06-30 · 54/100 · Skip